Welcome to the third post in my series on Builder’s Library notes. The last posts focused on continuous delivery and monitoring respectively with this entry focusing on techniques for building resiliency. Pieces this time include two high level ReInvent talks focused on building resilient services along with deep dives into ways to handle large traffic and avoid overloading your system. Methods deep dived include timeouts, backoffs and retries, que services, load shedding and static stability.

Amazon’s approach to building resilient services by Marc Brooker

In order to build resilient services, changes are needed for both tech and culture and this 2019 ReInvent talk focuses on two aspects of each which are ownership, operational, safety, service stability and jitter. Ownernership is needed to help close the loop between “dev” and “ops”. The devops loop goes build, fail, alayze, change practices, repeat, and is all about spinning the loop faster with better signal to improve the build process. Teams generally end up failing to change their practices before building again and make the same mistakes. Because of this, it’s important for everyone to be involved in the process, as errors in the project can cause delays, failures to meet goals and outages in the worst case scenario. To help prevent these issues, it’s important to prepare for events that haven’t happened and that’s something that requires cultural support.

At AWS, teams run what they build, being on call, operating in prod and are responsible for their product. Principal engineers are builders and are hands on and familiar with the specifics of their service. Everyone has operational responsibility and knows how to operate their services, with teams meeting to review graphs and metrics. Correction of error reviews help connect builders and operators from across the company and code architecture should help support operators and the review process. It’s important for code to be deployed often and quickly and to be able to deep dive any issues and incorporate learnings back into the implementation. Good intentions aren’t important and what matters is adding an enforceable mechanism to prevent the same error next time. Reviewing the loop can help with this process. Regarding the loop, a bad practice often leads to a fail and analysis causing the operators to learn but not the developers building the product.

A mechanism needs to incorporate the operator’s learning back into the development process and everyone should learn from issues even if they didn’t cause them. To help with this, AWS encourages kind systems that help users learn from usage and mistakes in a manner more similar to a game. A system should help a user build a good mental model and be intuitive to operate. A correction of error shouldn’t settle for operator error as the root cause of in issue as tooling should help solve the issues. Reviewing tooling is important and look for ways to fix the problem instead of blaming people to avoid taking responsibility.

Distributed systems have a problem illustrated with the comical “dog on the roof” example, where the dog usually doesn’t fall off, but sometimes it does and that’s bad. In distributed systems, an issue built in may take time to show itself and can potentially get bad quickly. For instance, an overload can increase latency, causing more failures and retries, adding even more load. To stop this, AWS limits queue sizes and retries and uses backoff with jitter as well as back pressure. Jitter is always used for backoff and period work at AWS and possible for other work.

For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/amazon-approach-to-building-resilient-services/

Architecting and operating resilient serverless systems at scale by David Yanacek

This next piece is a 2019 ReInvent talk by David Yanacek focused on resiliency with load shedding, dependency isolation, avoiding queue backlogs and good operational procedures. The talk goes into detail on the relationship between latency and throughput. Latency can grow at a steady rate as throughput increases, but once the fastest response time is slower than the timeout, total brownout occurs and all work is wasted. While throughput is all requests made, goodput refers to successful requests made in time. Goodput drops at the threshold while throughput can still increase.

Brownout can be avoided in a variety of ways. Scaling can add resources and load testing can let you know what can work. Load shedding can remove work to avoid overload. A timeout can also be set on the server side if the client’s timeout is known to avoid extra work and check-pointing can be used to save state with smaller requests. Some services can have fixed resources per unit of work such as Lambda’s Firecracker VM to avoid concurrency issues along with a placement function that can help avoid cold starts. Isolating unrelated APIs and services can help with this which AWS helps with in a lot of ways but queue backlogs can still create problems.

Queue backlogs can be solved via priority queues and Dead Letter Queues and be managed and retried by lambda. However, the easiest solution is not to have a backup at all by using back pressure or throttling to tell a customer they’ve failed via API Gateway or other services. If durability is more important than immediate feedback surge queues can help after letting enough messages into a warm queue. Shuffle sharding is also used by Lambda to help map workloads to queues if one gets backed up.

Most of the methods here have their own Builder’s Library posts with an article on queue backlogs, resource isolation, shuffle sharding and more.

For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/architecting-and-operating-resilient-serverless-systems-at-scale

Timeouts, backoffs and retries with jitter by Mark Brooker

This talk is about handling failures when one service calls another and designing systems to tolerate and reduce failures and stop them from causing an outage. Specifically, it talks about the three methods in the title. Timeouts avoid wasting time and resources by letting a service know when to quit, retries can fix random failures by making another request. Since traffic isn’t consistent, backoffs with jitter can add random time to avoid multiple requests overwhelming a service. To use these properly, APIs should be designed to be idempotent and allow for retries.

Some best practices for timeouts include to always set timeouts and to choose carefully as too high can reduce value and too little can waste work and add unnecessary retries. Base these off latency metrics for downstream services and know the acceptable false timeout rate. Factor in network latency and global distance and you can implement your own timeouts for better control. Adding a timeout to the request message can help let the server know how long you’re willing to wait to save work there. Lastly, you can establish a connection before making the request for a more consistent timeout.

Retries are useful for random failures though tradeoff taking more server time for a higher chance of success which can overload if there are load failures. Another issue is that distributed systems have multiple layers and create a chain of services retrying on a single failure. Using a single point in a system to retry can avoid multiplying retries and a token bucket can help as well, as used by the SDK. API’s should be idempotent to allow for safe retries when possible and HTTP can tell you when an issue is the client’s fault or the server’s letting you know not to retry if it’s a server error. Overall retries can be useful but require judgment.

Backoffs with jitter is AWS’s preferred solution to mitigating the issues with timeouts and retries. Backing off an exponential amount each time with a cap can help prevent server from being overloaded. If there’s a problem and a bunch of requests fail at once, normally there’s a good chance they’d all retry at the same time. Because of this, adding jitter spreads out requests with additional randomness to avoid spikes. Keeping the number for a host consistent can help debug and these are commonly used by Amazon where they can be used.

For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/

Using load shedding to avoid overload by David Yanacek

This article and David’s talk, Architecting and operating resilient serverless architectures, have a lot of overlap regarding the relationship between latency, throughput and goodput, but this adds a lot of detail regarding load shedding. Both mention the idea of the Universal scalability law and Amdah’s law which suggest that you can increase a system’s throughput with parallelization until it is limited by the throughput of tasks that can’t be parallelized, or points of serialization. When working on the “Services Frameworks” team, he helped integrate functionalities into services via configuration and faced the challenge of figuring out timeouts and max connections that the pre-ELB load balancers could take on. Without clear enough signals, load shedding was one of the techniques used to handle issues.

When a system is constrained on capacity, some of the worst things to do are to waste work or increase retries. Shedding load is a lot cheaper of an action than doing the work itself, but too many rejections could still slow down the process as they do take some work. A service shouldn’t reject any traffic that it can serve, but false positives should be kept to zero. Metrics on rejected traffic should also be separated from handled requests as it can make response time look faster than it is. It’s also advised to load test and know the failing behavior, making sure that the system still has enough capacity to failover by throttling before that point.

In rare cases, dropping requests can be worse than holding on to them, which is safer when not holding the up the applications thread. The server also needs to be able to receive a ping from it’s load balancer and respond to a health check to avoid being deregistered. Because of this, requests should be prioritized by using priority queues. Health checks are the most important task, and human requests are likely more important than automated crawlers which can likely be delayed. Clients can also use timeout hints to tell the server how long they’re willing to wait and allow it to determine if it can do the work or fail without wasting work.

Other methods such as using pagination and bounded work with multiple smaller entries can also help work to get done, and also prioritize earlier page requests. Queues can be configured to handle more recent requests depending on what’s needed with Classic Load Balancers using surge queues for excess traffic and newer, Application Load Balancers spilling over and rejecting excess traffic. Either way, surge queues and spillover metrics should be monitored and used to improve the process.

Service generally provide features to protect the service (such as max_connections). These should be used as last resorts for prioritizing traffic. Amazon sets max connections high on load balancers and lets servers shed locally while ensuring there’s room for a health check if overloaded. Iptables can also put bounds on connections and reject cheaply with sophisticated controls, usable with Network Load Balancers as they preserve the IP address of a caller. Overall, the core idea of the paper is that excess load can be shed to maintain consistent performance when facing overload.

This paper is overall very similar to the talk, with the paper being a bit more detailed while the talk can give you a lot of the same details if you’re in the mood to watch a video. The talk is given by the same author. A lot of papers and talks have this relationship.

For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/using-load-shedding-to-avoid-overload

Static stability using availability zones by Mike Furr and Becky Weiss

This article is a deep dive into statically stable design where systems continue to work even when their dependencies become impaired. AWS has a lot of ways to handle failure. Availability zones are physically separate and logically isolated sections of a region with fiber optic networking for a quick failover to another zone, with many services taking advantage of this. Scaling in response to an impairment may seem natural, but is less effective than static stability as it requires reacting to a failure instead of building the service to be prepared in advance.

An example of a statically stable design is EC2 which splits its processes into a control and data plane. The control plane makes changes to the system, such as launching instances, finding space and allocating network interfaces and related storage, IAM permissions, security group rules. The data plane keeps the EC2 instances running as expected, routes packets according to the VPCs route tables, reads and writes EBS data along with other tasks to keep the system functioning as it had been. The goal is that if the control plane is impaired, the data plane can keep running existing services even if new instances can’t be launched and configuration changes may be stalled. EC2 instances’ physical machines have needed routing information to keep communicating as expected.

Generally, the data plane is more important to keep available and runs at a higher volume and should keep running without contact from the control plane. The control plane typically is more complicated with more moving parts and more likely to break or have issues scaling due to that. Each plane should have it’s own separate scaling rules. Typically, services should be separated across these planes.

A couple of patterns for running these services are active-active such as load balanced services and active-standby such as relational databases. Amazon’s active-active services are composed of horizontally scalable stateless fleets or instances or containers run in autoscaling groups across three or more availability zones allowing for the loss of an availability zone. With the load balancer, traffic can be shifted away from a failing zone requiring no change to keep functioning. Elastic Load Balancer is also a service designed with this principle, though load balancers aren’t mandatory. The same idea can be run with a group of instances getting work from SQS instead.

Active-standby patterns can be helpful with stateful services requiring a leader node to coordinate work. A highly available setup of a database has a primary instance you write to and a standby candidate as well as potentially additional read replicas and a warm standby in another zone to fail over to. Similarly, this requires no additional infrastructure for failover, just a DNS change. RDS requires multiple availability zones and handles failover and the DNS change for you. Other services with leader nodes can failover by electing a node in another zone.

Apps can use EC2 for high availability with multiple availability zones and potential region management. Network traffic is kept within a zone and a regional highly available system is a consumer of the zones. Allowing zone independence is a foundation for high availability. An example given compares an architecture where one highly available service running across zones calling a similar service to an architecture where a service is deployed once in each zone. The benefit to the services running separately in each zone is that an issue in one zone has a lower chance of impacting a call when the whole process is limited to one zone instead of having a chance to call different zones at different parts of the process.

Because of this, NAT Gateway is a zonal service with one per zone. This is a part of the data plane and can survive an availability zone loss by keeping data plane resources within an availability zone. For helping with zone independent services, data can be backed up across zones, such as with S3.

A couple of simpler service patterns are regional calling regional and regional/zonal calling zonal. Regional calling regional involves regional services calling each other, which can be useful for external and internal facing services and provides high availability even when a zone is impaired. Regional calling zonal or zonal calling zonal can help provide zone independence for data plane components and keep all network traffic within a zone.

For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/static-stability-using-availability-zones/

Avoiding insurmountable queue backlogs by David Yanacek

When queues are short, they’re simple but people tend to only start thinking about their behavior when things go wrong. This article talks about design approaches for draining queues, prioritizing workloads and preventing backlog build up. Queues are helpful for building resilient asynchronous systems. They can accept persisted messages from another system and store until processed, surviving outages for dependent systems. The queue will redrive the messages until they’re successfully processed. Overall, queues increase durability at the cost of occasional latency increases with retries.

Amazon uses queues for physical and virtual processing. It helps increase availability but if processing stops while messages keep arriving, a backlog can occur and delays can make work complete too late to be useful causing an availability hit. A queue based system has two modes. No backlog and low latency and the other case where more messages are arriving than being processed and latency and work required increase. An example of a queue based system is Lambda’s which has a synchronous mode where output is returned in HTTP response or asynchronously with the HTTP response telling the user that the function is running behind the scenes. Lambda needs to ensure functions run even with servers and has a durable queue to be able to redrive a request when it fails. IOT core is another service that makes use of queues. Devices and apps connect to and subscribe to pub sub topics, they publish messages and subscribed apps can receive the message. This needs to be asynchronous since many devices are offline and need to get messages when back, requiring persistence behind the scenes. SQS is a durable and scalable queue implementation that’s used behind the scenes for these two services. In these systems, a component produces messages with consumers processing and deleting them.

When failures occur, re-queues and retries can be done, but recovering from an outage can require double capacity for the same time. Elastic services can help scale servers to handle the backlog. If other services are needed and aren’t ready for the increased load, delays can occur. Synchronous services usually drop excess requests and recover quickly, but asynchronous will build up backlogs and have a delayed recovery time. This a hidden risk of asynchronous processes.

Another important question is how to measure availability. Dead letter queues are a good measurement that’s caught too late, but can still help to set alarms on. Latency can help by measuring time stamps of message in queue and producing metrics. In some cases such as IOT, backlogs are expected and it’s important to distinguish devices offline from unexpected backlog. For instance, IOT uses the age of the first attempt for the first subscribers with other metrics tracking the flow. Separating first attempts from retries is a commonly used strategies since the retries are likely for offline devices. X-ray and request IDs are commonly used for tracing messages across systems.

Backlogs in multi tenant are a problem as customers expect the same performance as a single tenant system. Internal queues are managed by Amazon with API’s for services being exposed to customers, adding the caller’s info to messages. Many APIs return a response while processing the request asynchronously. It’s important to use fairness to make sure customers have a good experience within their limits. Queue specific mitigation strategies involve each component protecting itself and ensuring fairness. Additional queues can help shape traffic as it’s hard to isolate workloads with only one queue. LIFO behavior is typically desired during a backlog.

Amazon’s strategies for handling backlogs involve giving each customer its own queue if only a few customers, or to use other techniques if tens of or hundreds of customers. Shuffle sharding can help assigning customers to queue. Additionally, excess traffic can be put in spillover queue to work on later and live traffic can be prioritized. If time sensitive, older message can also be dropped. Threads and resources can be limited per workload. IOT rules engine uses non block I/O to avoid exhaustion and can use a semaphore to measure and limit concurrency for a workload overall, fairness should be used to help this. The rules engine is queue based for IOT and buffers services and devices.

Across Amazon, separate thread pools are used per workload and atomic integer limits concurrency. Rate based throttling can help for rate based resource isolation. Back pressure can also be used to start rejecting work if certain workloads are a problem and it’s acceptable. Delay and dead letter queues can help avoid too many inflight messages as well if reprocessing is better. Additional buffers should be added to polling threads per workload to avoid running out of buffer to absorb surge in traffic. A last tip is to heart beat long running messages and let SQS know work is still being done or stop wasting work after its timeout.

For the full article these notes are taken on, check out https://aws.amazon.com/builders-library/avoiding-insurmountable-queue-backlogs/

Leave a Reply

Trending

Discover more from NikCreate

Subscribe now to keep reading and get access to the full archive.

Continue reading