SLO from Methodology to Practice: Part1 – Establishing Effective SLOs¶
In recent years, organizations have increasingly adopted Service Level Objectives (SLOs) as a fundamental part of their Site Reliability Engineering (SRE) practices. Google pioneered the best practices around SLOs—the Google SRE books provide an excellent introduction to this concept. At its core, SLOs are rooted in the idea that service reliability and user happiness go hand in hand. Setting specific, measurable reliability targets helps organizations strike the right balance between product development and operational work, ultimately leading to a positive end-user experience.
We will learn about SLOs in three parts:
SLO from Methodology to Practice: Part1 – Establishing Effective SLOs
SLO from Methodology to Practice: Part2 – SLO Tool Selection
SLO from Methodology to Practice: Part3 – Best Practices for Managing SLOs with Guance
Key Terminology¶
Before we proceed, let's first break down some key terms that will be used throughout the series:
- Service Level Indicator (SLI) – A metric used to measure the level of service delivered to end users (e.g., availability, latency, throughput).
- Service Level Objective (SLO) – The target service level, as measured by an SLI. SLOs are typically expressed as a percentage over a period of time.
- Service Level Agreement (SLA) – A contractual agreement that outlines the level of service end users can expect from a service provider. Failure to meet these commitments may result in significant consequences, often financial (e.g., service credits, subscription extensions).
- Error Budget – The acceptable level of unreliability a service can have before failing to meet its SLO. In simple terms, it is the difference between 100% reliability and the SLO target. You can think of an error budget like a financial budget—except in this case, developers spend the budget on building new features, redesigning system architecture, or any other product development work.
Who Cares About Service Level Objectives?¶
For SLOs to be adopted by key stakeholders across the organization, you need them to agree on realistic reliability targets that consider business priorities and the projects they wish to undertake. In this section, we will examine what end users, developers, and operations engineers care about, and how we should consider their goals and priorities when setting SLOs.
End Users¶
Regardless of the product, end users have expectations about the quality of service they receive. They expect the application to be accessible at any given time, load quickly, and return correct data. While you can measure customer dissatisfaction through support tickets or incident pages, you should not rely solely on these to make product decisions, as they do not fully capture your end-user experience. For example, resolving all tickets does not necessarily mean you have met the service level expected by end users.
In reality, achieving 100% reliability is impossible. SLOs help you find the right balance between product innovation (which delivers greater value to end users but risks breaking things) and reliability (which keeps those users satisfied). Your error budget determines how much unreliability development work can tolerate before end users experience degraded service quality.
Developers and Operations Engineers¶
Traditionally, friction between developers and operations engineers stems from their opposing goals and responsibilities: developers aim to add more features to their services, while operations engineers are responsible for maintaining the stability of those services. SLOs not only drive positive business outcomes but also foster a cultural shift, giving development and operations teams a shared sense of responsibility for application reliability.
With SLOs and their accompanying error budgets, teams can objectively decide which projects or initiatives to prioritize. As long as there is remaining error budget, developers can release new features to improve overall product quality, while operations engineers can focus more on long-term reliability projects such as database maintenance and process automation. However, when the error budget starts to deplete, developers need to slow down or freeze feature work and collaborate closely with the operations team to stabilize the system before any SLA or SLO is violated. In short, the error budget is a quantifiable method to align the work and goals of developers and operations engineers.
From SLI to SLO¶
Now that we have defined some key concepts related to SLOs, it is time to start thinking about how to create them. Deeply understanding how your users experience your product—and which user journeys matter most—is the first and most important step in building useful SLOs. Here are a few questions you should consider:
- How do your users interact with your application?
- What is their journey through the application?
- Which parts of your infrastructure do these journeys interact with?
- What are their expectations of your system, and what do they hope to accomplish?
In this series, assume you work for an e-commerce business and consider how such a business would set SLOs. You need to figure out how your customers interact with the website—and their path from first entry to exit. At a basic level, your customers need to be able to log in, search for products, view details of individual products, add items to a cart, and checkout. Such key user journeys are directly related to user experience, so it is important to set SLOs for them.
Once you complete this exercise, you can move on to selecting metrics, or SLIs, that quantify the level of service you deliver in these critical user journeys.
Choosing Good SLIs¶
As your infrastructure grows more complex, setting external SLOs for every database, message queue, and load balancer becomes increasingly cumbersome. Instead, we recommend organizing your system components into a few major categories (e.g., request/response, storage, data pipeline) and specifying SLIs within each category.
When you start picking SLIs, keep this short but important saying in mind: “All SLIs are metrics, but not all metrics are good SLIs.” This means that while you may track hundreds or thousands of metrics, you should focus on the most important ones: those that best capture user experience.
You can use the table below (from the Google SRE books) as a reference.
Now, imagine your shopper is stuck on the checkout page, waiting for a slow payment endpoint to return a response. The longer they wait, the more likely they are to form a negative impression of your business. Beyond reputation damage, abandoned carts can have costly consequences. In fact, some of the largest and most successful organizations have found that every second of delay correlates with a significant drop in revenue. From this example, we can see that response latency is a particularly important SLI for online retailers to track, ensuring their customers can complete critical business transactions quickly.
Contrast this with a metric that almost certainly would never be a good SLI: CPU utilization. Even if your servers are experiencing a spike in CPU usage—and your infrastructure team is receiving alerts about this high usage more frequently—your end users may still be able to checkout seamlessly. The point here is that no matter how important a metric is to your internal teams, if its value does not directly impact user satisfaction, it is not useful as an SLI.
Once you have identified good SLIs, you need to measure them using data from your monitoring system. Again, we recommend pulling data from the component closest to the user. For example, you could use the payment API to accept and authorize credit card transactions as part of the checkout service. While many other internal components may make up this service (e.g., servers, background job processors), they are typically abstracted from the user's view. Since SLIs are used to quantify your end-user experience, collecting data from the payment endpoint alone is sufficient because it exposes the functionality to the user.
Converting SLIs into SLOs¶
Finally, you need to set target values (or ranges of values) for your SLIs to turn them into SLOs. You should specify what your best- and worst-case criteria are—and over what time period the condition must remain valid. For example, an SLO tracking request latency might be: “Over 30 days, 99% of authentication service requests will have a latency of less than 250 milliseconds.”
When you start creating SLOs, keep the following points in mind.
Be Realistic No matter how tempting it is to set an SLO at 100%, it is practically impossible to achieve. Without accounting for an error budget, your development team may become overly cautious when trying new features, which can stifle your product's growth. Typical industry standards are to set SLO targets at multiple nines (e.g., 99.9% is referred to as “three nines”, 99.95% as “three and a half nines”).
As a general rule of thumb, your SLOs should be stricter than what you detail in your SLAs. It is always better to err on the side of caution to ensure you meet your SLAs, rather than consistently underdelivering.
Experiment There are no hard and fast rules for refining SLOs. Each organization's SLOs will vary based on the nature of the product, the priorities of the teams managing them, and the expectations of end users. Remember that you can always continue to optimize the targets until you find the right values. For example, if your team consistently exceeds the target by a wide margin, you might want to tighten those values or leverage the unused error budget by investing more in product development. Conversely, if your team consistently misses the target, it may be wise to lower it to a more achievable level or dedicate more time to stabilizing the product.
Keep It Simple Last but not least, when defining SLO targets, resist the temptation to set too many SLOs or make SLI aggregation overly complex. Instead of setting separate SLIs for every cluster, host, or component that makes up a critical journey, try to aggregate them meaningfully into a single SLI. In general, you should limit SLOs and SLIs to only those that are critical to your end-user experience. This helps eliminate noise and allows you to focus on what truly matters.
Now You Know Your SLOs¶
In this blog post, we explored how selecting the right SLIs and converting them into well-defined SLOs can set your organization on a path to success. By using SLIs to measure the service level you deliver to users—and tracking your performance against actual SLOs—you will be better equipped to make decisions that improve both feature velocity and system reliability. We have summarized this guide as a simple checklist that you can refer to when you start creating SLOs and bring more team members on board.
Continue reading the next part of this series to learn how technical and business teams can use Guance to collaborate more effectively by managing SLOs alongside their other monitoring data.
