Welcome to the beginning of a series of posts meant to serve as my notes on Amazon’s Builder’s Library, a series of white papers and videos detailing some of Amazon’s findings and lessons learned while building distributed systems. This series will cover Builder’s Library’s 23 pieces in its catalog prior to ReInvent 2020. Perhaps my biggest takeaway from this undertaking that I need to get better at the learning process as reading through the papers, note taking and summarizing has taken a lot more time than expected.
This first article focuses on five pieces that focused more on the continuous delivery process and getting code to customers, with papers focusing on Amazon’s history and adoption of continuous delivery, best practices and high level overviews of their process and deeper dives into important deployment processes.
Going faster with continuous delivery by Mark Mansour
The first paper is a short history of Amazon’s adoption of continuous delivery and is one of the quicker reads of the series at five pages compared to a usual eleven.
Amazon used to take 16 days to get code to production with 14 being deployments and tests. They wanted to save work and reduce delays by automating the process. They already had an internal build system called Brazil and a deployment system called Apollo, which both had to be started manually passing code or an artifact. Pipelines were implemented to automate both, looking something like AWS’s pipeline service today.
Amazon’s Pipelines continuous delivery tool reduced time for teams that used it by 90 percent. All release steps had to be defined such as building the artifact and deploying in a way that built confidence that no defects were present. At first, pipelines could only model one release process per application which helped with simplicity and consistency. As more teams started using the tool, the team started evaluating pipelines realizing testing was often missed. Determining what testing was needed depended on the team and needed to be customizable.
An important balance is that between increasing execution speed and having sufficient testing. The improved release process shouldn’t hinder the business process. Also since teams weren’t learning from each other a better learning tool was needed beyond messaging. Customizable checks were added so that best practices could be recommended in the tool and warn the team if not broken but the team could configure them as they learned. Priorities for the deployment process were availability, then speed and then being engineer friendly, with issues being identified as quickly as possible.
Testing ensures that the deployed artifact can start and respond to work properly. Lifecycle hooks in Code Deploy can be used for triggering scripts at different stages. They should test for the capacity for serving customer traffic and that the minimum healthy hosts are met and be able to rollback if a failure is detected.
For pre-production testing, unit, integration and pre-production environment deployments should be automated in the pipeline. Load and security testing should be included in there and unit tests should include style and code complexity checks. Integration tests are all off box such as browser testing and failure injection. Artifact build functions should be verified with testing as well and using smaller tests allows for quicker feedback. Pre-production environments should be an exact simulation of production.
The process of rolling out to prod in waves is described, adding that synthetic traffic should test all public APIs. Blockers can stop deployment and can be based on content such as a git package. and be overridden if needed. Typically deployment hours are based around business hours.
This paper contains a lot of the same information on deployments as “Automating safe, hands-off deployments” with a lot more time spent on the story of how Amazon moved to using continuous delivery while Automating safe, hands-off deployments goes into more detail on the process itself, briefly touching on Amazon’s history.
For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/going-faster-with-continuous-delivery/
Automating safe, hands-off deployments by Clare Liguori
This second paper discusses best practices for safe, hands off deployments using pipelines and goes much more in depth on the deployment process than the previous. Amazon deploy changes multiple times a day with fully automated pipelines and the last time a developer touches code is when merging it into the “mainline” branch. Pipelines use four phases for continuous delivery; source, build, test, and prod. Source can refer to source code for an app, tool, test or IAC deployment along with static assets, dependency libraries, patches and configurations. Build provides an appropriate action for the type of source, like compiling code, testing, reviewing or packaging and storing an artifact. Source changes are then validated and deployed to prod. All source changes are stored in repositories and version controlled with dependencies updated weekly.
Including various source types with a pipeline for each helps to ensure that no changes go untestested. A microservice could have an app, infrastructure, and operating system patching pipeline, etc. Failure in one pipeline would cause an automatic rollback and wouldn’t affect another.
Before being merged, all changes are code reviewed with the pipeline blocking unreviewed commits from going through. Code is manually reviewed for correctness, safety and quality of tests and monitoring, often using a checklist. After the source is passed to the pipeline, the code is compiled and tested with based on the teams’ preferences of frameworks and linters. The type of code also determines the tools, with unit tests being appropriate for apps and linters for IAC templates. Builds run without network access to isolate them and use simulated dependency calls with live dependencies tested in integration tests. When the build is complete, the compiled code is packaged and signed.
Test deployments are done in pre-production with the pipeline deploying to multiple environments. Alpha and Beta validate code with functional APIs and end to end integration tests, while Gamma simulates prod’s configuration and uses the same monitoring, alarms and testing as prod, including multiple regions. Before being rolled out to Gamma, a “one box” or “canary” deployment is deployed to the smallest deployment unit (single VM) and tested before the rest is deployed after a specified “bake time” with no errors. This also ensures that code is backward compatible and won’t create an issue with rollback.
Integration tests use real APIs and infrastructure and can help catch unexpected behavior and validate unit tests assumptions for how dependencies should act. Invalid inputs and errors are also tested along with fuzz and load tests. Microservices owned by a team call pre-production deployments for the same team’s services and prod for other team’s. Some teams also use a Zeta stage when each microservice calls only production endpoints testing for backwards compatibility.
When doing production deployments, the main objective is preventing negative impact on multiple AZs and regions, ideally limiting further by deploying to shards and AZs in an increasing rollout. For instance, the first wave could deploy to a one region at a time for the first couple waves, with one box tests for each. Then as the team becomes confident in the deployment, they can increase to 3 regions, 12 regions, and so on as the deployment continues, deploying one AZ at a time. Additionally, different deployments can be in different waves running at the same time. While one at a time might be safest, it could be too slow, with waves being a good balance. With each wave starting with a one box, typically 10 percent of requests will start being served by new code. Afterwards, a rolling deployment with a maximum of 33 percent of servers at a time to ensure 66 percent availability is maintained. To limit impact further some teams will rollout five percent at a time and rollback 33% at a time.
Metrics monitored by an alarm and deployment system are used to rollback automatically and notify an on-call engineer. Alarms and metrics vary by service and upstream and downstream alarms are monitored as well, along with canary tests. Teams create high severity aggregate alarms to combine metrics into one alarm. One box should have its own alarms and metrics as well and anomalies in web service metrics such as no requests or high latency can be used to roll back.
To catch problems that don’t show up immediately or if load is low, additional changes aren’t rolled out until a bake time is up with the opportunity for rollback based on a high severity aggregate alarm until then. Earlier waves have higher bake times and a deployment can potentially last days to ensure buggy code isn’t rolled out to multiple regions. Typically there’s one hour between each one box and the rest of the region and twelve before starting on a new region, though in rare cases bake time can be lowered with principal engineer’s review.
The pipeline will prevent automated deployments if there’s a high risk of negative impact. Blockers help to evaluate risk with the pipeline checking alarms to potentially stop from moving forward. The pipeline can also limit deployments to time windows, though too small can cause changes to pile up with large prolonging failed events. Limiting changes to within business hours can help ensure that there’s support available if a problem arises. Blockers can be overridden by developers if needed to fix an issue.
Lastly, the pipelines themselves are maintained as code to help them be reused with settings for regions and availability zones, allowing for inheritance and customization based on service and need.
For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/automating-safe-hands-off-deployments/
Amazon’s approach to high-availability deployment by Peter Ramensky
This is a reinvent talk about highly available deployments and Amazon’s process for improvement. Peter talks about identifying safety practices and working them back into tools across an organization. A deployment is described as any change to the system whether its a software config update or IaC code, data migration, failover, adding a table, certificate change, etc. Failures suggest a customer was impacted so stopping a bug before a customer is harmed is considered a success.
Three important factors when debugging an issue are the duration, latency effects and overall impact of the problem. Correction of errors are held where service teams give summarized response to service events including what happened, data and metrics, the impact to the customer, lessons learned and corrective actions that will be taken. The loop this time is shown as service team operations lead to customer impacting event leading to correction of error and action items, and going back into operations. Some corrective actions can be specific to a service while others can help improve all teams. Tools and best practices should be updated before completing the loop as well to help grow over time. Improvements are tagged with “prevent,” stopping an issue entirely, “shorter,” reducing error time to prevent an issue, and “smaller,” reducing an error size to prevent an issue.
Amazon uses two pizza (8 to 10 engineer) teams with decentralized ownerships and local process decisions. Consistent operational standards form across a team with identical tools and platform used for deployments. Limiting tools helps to solidify best practices. The pipeline can audit that rules are being followed while the rules enforce the pipeline creating a loop. Each tag has one rule for it.
Deployments are done in waves with testing varying by environment. An auditing process models the configuration of a pipeline passing it to rules along with any infrastructure. These rules are written as Lambdas and launched into organizations when ready. Service teams can make their own rules which will have the ability to stop deployments when launched. Deployments that are running before a rule is launched will report where they would have failed and alert a team of what needs to change. Service teams have the option to ignore if needed, which will notify leadership and rules will be enforced again after a day. Teams can also request exemption. The overall classification of a pipeline will be whether a customer impact would occur or not.
Best practices for these pipelines are to use integration tests with a beta stage and dependency, pre-production testing with production dependencies and using a one box and rolling deployment to production. Rollback alarms and synthetic traffic will help test and revert changes if needed during deployment. Some new approaches are doing fractional deployments in waves and traffic shifting to a canary. Instances should always be able to test 50 percent more load than needed and an additional load balancer can be used to point to a canary instance and run pre-production tests against it with production dependencies. The canary can be added back to the prod fleet afterwards.
Anomaly detection is a useful tool that should be attached to various metrics and used to trigger rollback alarms. Any metrics correlated with customer impact should have this such as service metrics like faults and errors, instance metrics like CPU and memory and runtime metrics such as heap, garbage and thread usage. A detector should be trained on data before a deployment to know what normal is and rollback can be based on a single metric or need multiple. Anomaly detection can be applied through organizational tools automatically.
For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/amazon-approach-to-high-availability-deployment/
Amazon’s approach to security during deployment by Colm MacCarthaigh
This is a talk about security at Amazon, it’s importance and best best practices. Security is the the top priority at Amazon followed by durability, availability and then speed. Availability can be sacrificed if there is a security issue that needs fixing. It’s important that the team and culture promote and reward security in a company as it can often go overlooked. Security should be built into services and plans should exist for handling new threats and handling incidents. As Peter Drucker says, “culture eats strategy for breakfast.”
At Amazon, security is important to leadership and the CISO and CEO are aware of team’s security issues and there are challenges for talking to security leadership are established and can be escalated if needed with weekly meetings held normally.
Leaders are owners of products and should insist on the higher standards to make sure products work as well as possible and not sacrifice their long term value for short term gain. It’s important to deep dive on an issue and understand it in detail and keeping systems simple can help keep focused and make security better with less angles to attack from. Code should be clear and readable as typically the least readable code bases have the most issues. Defenses should be in place including ones to prevent human error. Amazon teams aim for technical fearlessness and good security requires humidity and accepting reality, being willing to admit faults and not blaming.
Three pillars of security are policy, processes and tools. Tools save users and systems every day, building security into systems. Policies include training, encryption standards, PII hardening standards, compliance requirements. Processes at Amazon include reviews for new features and service starting early and continuing through the development. Penetration testing, fuzzing and vulnerability scanning are useful for catching issues and formal verification by building formal models and applying verification techniques can help to find vulnerabilities before production.
For building security into tools, Amazon starts building a service a Octane, an internal tool for building new services that includes IAM, CloudTrail, Config, TLS/SSL implementation and more for including security automatically. Every team also has tools to manage operations securely with services for authentication, access control, patch management and auditing with vulnerability notifications. In order to react to new and emerging threats its important to hire and develop the best. They should find gaps in new encryption mores and protocols and collaborate with security researchers to help fix potential vulnerabilities.
One example of a security issue that hit Amazon and the rest of the world was Heartbleed which was a bug in OpenSSL which disclosed memory. The attacker could send a corrupt heartbeat to steal data. Amazon froze deployments and prepared a fix deployment including a patch to fix open ssl, now using S2n (their implementation). Millions of deployments happened within the day and had to help customers update their side including ones using legacy applications. The average team lost three days during the attack.
For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/amazon-approach-to-security-during-development/
Ensuring rollback safety during deployments by Sandeep Pokkunuri
This article is all about deploying changes that are safe to rollback or forward. Rollback errors are tough to catch with pre-production environments alone and the action of rolling forward and backward should be tested to ensure all changes are backwards compatible and won’t disrupt functionality. Distributed systems provide an additional challenge related to this. In a standalone system, run as one process on a device, two versions of the same code can’t run at the same and backward compatibility is met by old and new versions being able to read each other’s data. In a distributed system, changes are incremental and both versions can exist and once needing to be able to communicate with each other and the system.
The most common reason for rollback failure is a protocol change, such as a change compressing data in a way the old version won’t know how to read or a change in heartbeat period causing an instance to be terminated. Explicitly testing that rolling is safe and that different server versions can exist at once. If a problem exists, it can be done with a two phase deployment, split into a prepare phase and activate phase. The prepare phase updates servers to be able to handle both the old and new data formats and is run first. Once all servers are prepared, activate will update each server to write into the new data format. Rollback on prepare causes no issues since its not taking away read functionality and rollback on activate goes back to prepare which can read activate stage data as well. For this example, prepare could give a lenient heartbeat period while activate changes heartbeat frequency. If going this route make sure all servers are successfully updated in the prepare phase before moving forward, as some report success based on percent. This pattern was used to update DynamoDB’s protocol at Amazon. Both stages can’t be rolled back at once and time should be left between rollbacks. Lastly, any left over data should have its format updated with a backfiling process.
Best practices when serializing data involve using tested formats as opposed to developing your own which can handle existence checking and escaping. Structures that can’t be controlled shouldn’t be serialized such as classes and languages. Versions should be written and unknown attributes should be allowed so old versions can ignore new version’s original attributes without overwriting. These can help avoid the need for the two phase deployment in the first place.
Lastly, a step by step process for rollback testing is to set up an environment to model production without the update. Rollout the update it the same way as prod, updating reader then writing going forward and the reverse going backward. Simulate update as closely as possible to be the most confident that the changes are safe to deploy. This helps avoid relying on manual analysis and if a change isn’t safe, it can generally be split into two changes.
For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/ensuring-rollback-safety-during-deployments/




Leave a Reply