The Reliability Top 10, First Edition
Ten of the thirty-eight causal factors, ranked by the number of public incidents each appears in, across 1,287 incidents from 139 organizations. A reference for deciding where reliability effort goes, with the counting rules and the known biases stated alongside it.
When teams think about reliability, they usually think about how a lack of it wakes them up in the middle of the night. Many teams stay in that mode, examining their own incident history but unable to move toward a more proactive reliability posture. They have a list of how things have gone wrong for them in the past, but not an evidence-based set of factors to watch for. This list is a second reference: the ten conditions that show up most often across 1,287 public incidents from 139 organizations. Use it to look for the gaps your own history can’t show you. If you only have a few minutes, read the table and then skip to What to do with it.
This is the third post in a series. Building a Causal Factor Enumeration for Reliability covers the motivation and the method: why reliability has no shared names for what goes wrong, and how the population behind these numbers was built. The causal factor enumeration is the catalog of all 38 contributing conditions. This post ranks ten of them by the number of distinct public incidents each appears in.
This is not a ranking of how often these conditions occur. It’s a measure of how often they get described in public by organizations that publish incident reports. Sentences that claim “X causes N percent of outages” based on this ranking are false. “X appears in N percent of the incidents in this corpus” is the correct form.
The Reliability Top 10
1,287 incidents across 139 organizations, 3,850 ranked conditions.
| Slot | Entry | Incidents | % | Orgs |
|---|---|---|---|---|
| RT01:2026 | CF-0033 An external provider the service depends on became unavailable or degraded | 367 | 29% | 111 |
| RT02:2026 | CF-0028 An internal dependency failed and the failure propagated to its callers | 238 | 18% | 63 |
| RT03:2026 | CF-0002 A software release introduced a defect into production | 232 | 18% | 59 |
| RT04:2026 | CF-0001 A configuration change altered behavior beyond its intended scope | 216 | 17% | 56 |
| RT05:2026 | CF-0015 Demand exceeded the capacity the service was provisioned for | 214 | 17% | 60 |
| RT06:2026 | CF-0030 A network path lost connectivity or capacity | 150 | 12% | 53 |
| RT07:2026 | CF-0026 A protective control engaged and its action became the outage | 87 | 7% | 33 |
| RT08:2026 | CF-0003 A planned change was carried out as intended and had an effect nobody anticipated | 86 | 7% | 37 |
| RT09:2026 | CF-0025 Recovery took longer or needed more manual work than planned | 76 | 6% | 37 |
| RT10:2026 | CF-0012 A rollback, failover or restart left the service degraded | 71 | 6% | 34 |
How to read the table
Two identifiers. RT01:2026 is a ranked slot for this edition because the ranking is expected to change. CF-0033 is the entry that slot points at, and it will never change. Cite the CF identifier when you mean the condition, and the RT slot when you mean its position in a particular year. The catalog numbers are assignment order and carry no ranking at all, which is why an entry numbered 33 can sit first here.
Rows overlap, so no column adds up to a total. A single incident usually involves several of these conditions, and it’s counted in every row it appears in. One where a vendor outage led to a failed rollback sits in the CF-0033 row and the CF-0012 row both. Across the full enumeration that averages two entries per incident, so the ten incident counts add to more than the 1,287 in the population, the organization counts do the same, and the percentages add past 100. Read each row on its own, as a share of the same 1,287.
All ten hold a top-ten place under every counting method we tried: by incident, by organization, by collapsed shared event, and by document. Whatever you think of the exact order, the membership of the set isn’t an artifact of one choice.
Nothing in the top four is going to surprise anyone who has carried a pager, and that’s roughly the outcome I wanted. A list whose leaders nobody recognized would be measuring something other than production.
Like the OWASP Top 10, this is an awareness document and a starting point. It’s directional and it isn’t a study. The ranking exists so a team can check its own reliability work against something besides its own incident history, which is the only evidence most teams have and is both small and selected by whatever they happened to survive. A list drawn from 139 organizations is a second reference, with different blind spots. The counting rules and the known biases are below so you can decide how far to trust any single row.
How the counting works
The unit is a distinct incident, keyed as one organization on one day. The corpus often holds the same incident twice, once from a company’s status feed and once from its status page, and the key folds those copies into one count. Two separate events at the same company on different days stay separate. The key is crude in both directions. When the two copies of one incident carry dates a day apart, the incident counts twice, which adds 29 incidents, or about 2 percent. When a company has separate incidents on the same day they count once, which happens in 72 of the 1,287 keys and hides about 87 incidents. Net, the 1,287 is roughly 4 percent low.
The extraction returns nothing when a document doesn’t say why a failure happened. 898 of the 2,619 analyzed documents produced no conditions at all, about 34 percent. Those had all passed the causal-language filter, so that share is partly the filter’s own false positives and partly the extractor not finding a cause in the text.
Entries are not ranked by the number of extracted conditions. That would measure how thoroughly an organization documents its causal conditions, not how often they actually occur. A detailed analysis from a large provider yields around ten conditions per document where a status update yields one or two.
Disclosure depth is a property of the incident which splits three ways. An incident yielding one or two conditions is short, one yielding eight or more is detailed, and a large middle group sits in between. 67 of the 1,287 incidents are detailed accounts of distinct incidents. A detailed incident may carry many conditions and appear under several entries at once.
The organization column counts organizations, not publisher names. Publisher names are assigned per feed, so one company can have several publishers. Examples are Atlassian Jira, and Atlassian Confluence, which are two distinct publishers, but one company. Atlassian actually has eight publishers in the corpus. We’ve mapped publisher names to organizations. Publishers that didn’t operate the system they describe are removed before anything is counted. That leaves 139 organizations against 210 publisher names. A publisher whose documents produced no conditions contributes no counted data, so 139 is the number of organizations that described a cause the extraction could classify.
The top entry is exposed to fan-out, so we measured it. When a large provider goes down, every customer that depends on it posts a notice, and one event can look like forty. CF-0033’s 442 documents fall across 211 distinct days, and only a handful of those days carry five or more organizations reporting together, which is the signature of a shared event. Collapsing those into single events takes the entry from 367 incidents to 288, a reduction of about a fifth, and it still ranks first by a wide margin.
Matching keywords produces very different results. Run analysis on the same corpus using causal keywords and processing backlog CF-0024 comes out on top. Classified against the enumeration, that condition sits thirteenth, in 61 incidents. Keyword matching finds the failure modes that have settled language in incident writing, which is not the same property as how often a condition gets described.
Limitations of publicly disclosed incident data
No analyst familiar with publicly disclosed incident data will tell you it’s perfect. A third of the documents that passed the filter don’t say why anything failed. Only 5 percent of incidents come with a detailed analysis. The 1,602 documents behind the ranking turn out to be 1,382 distinct incident records, because 218 of those records were collected two or three times. The incident count is off by a few percent in each direction. The top entry loses about a fifth of its incidents when shared provider outages are collapsed.
That doesn’t make the list or its data useless. Every one of the known defects was measured and the results survived examination from multiple angles. The same ten entries hold a top-ten place under four different counting methods. The fan-out correction removes 79 incidents from the first entry and it stays first by a wide margin. The keying errors are a few percent of the population. I haven’t re-ranked with incidents keyed by record, so I can’t promise that no slot would move, but the ten items currently ranking would be in the set regardless.
It’s important to understand that the data supports coarse conclusions, not narrow claims. This ranking is meant to be directional. The corpus can tell you that provider failure, release defects and configuration changes belong near the top of anyone’s list. It can’t tell you that the fourth entry (216 incidents) matters more than the fifth (214 incidents). Understand it within your specific context and at the resolution intended.
What to do with it
Here are four potential uses, in rough order of how quickly they might pay dividends.
Run a coverage check. Take the ten into a room with your senior engineers and ask two questions of each: could this happen here, and would we detect it inside our target. The list is short enough to get through in an hour. What you’re looking for is the entry where the answer is “yes and probably not”, because that gap won’t surface from reading your own incident writeups, which only contain the things you’ve already experienced.
Tag your own incidents with the identifiers. Start marking each incident with the entries that appear in it, working from the CF identifiers in the published catalog. After a couple of quarters your own history becomes countable, and measurable against a public baseline. There’s a compounding effect as you practice this which is why the enumeration matters more than the ranking.
Put the entries into design review. Each one is a single sentence describing what a system did, which makes it usable as a review prompt. CF-0026 works especially well: for every rate limiter, circuit breaker, automated suspension and failover in a design, ask what it looks like to a customer when that control fires correctly under conditions nobody modeled.
Compare the shape of your investment to the shape of the list. The ten fall into four groups. Three are about changes you made on purpose: the release defect, the configuration change, and the planned change with the unanticipated effect. Three are about things you depend on: the external provider, the internal dependency, the network path. Three are about what happened when somebody responded to trouble: the protective control, the recovery that ran long, the failover that left things degraded. One is capacity. Reliability budgets tend to lean heavily into the first group through testing, review and release engineering, and into capacity planning for the fourth, while the middle two get considerably less attention.
Three findings worth acting on
Three of the ten describe what happened after a normal action
- CF-0026, a protective control engaged and its action became the outage.
- CF-0003, a planned change was carried out as intended and had an effect nobody anticipated.
- CF-0012, a rollback, failover or restart left the service degraded.
Rate limiting, load shedding, failover, restart: deliberate actions, carried out correctly, where a normal planned action led to or extended the outage.
Observed in incident reports
| Google 2023-10-29 | “our networking systems correctly triaged the traffic overload and dropped larger, less latency-sensitive traffic in order to preserve smaller latency-sensitive traffic flows” |
| Linode 2026-08-13 | “Akamai engineering teams identified a configuration discrepancy on the secondary caching infrastructure node that prevented it from absorbing the full traffic load after the failover.” |
| Honeycomb 2026-05-01 | “The database failover that caused query issues earlier in the EU has also had knock-on effects on our Activity Log beta feature.” |
None of those describes a defect. For anyone who owns this work, the lesson is that safeguards and recovery paths need the same review attention as the failure paths they attempt to avoid, including an analysis of what their correct operation looks like from outside the system. Most programs review a control for whether it works. Fewer review it for what happens when it does.
This family is also the argument for building an enumeration in the first place. A static analyzer can’t propose any of these categories, because there’s no defect for them to find. A practitioner survey won’t ask, because you’d need to already know to put it on the questionnaire.
Entries about response are likely to be undercounted
Three entries describe what happened when someone or something responded to trouble:
- CF-0026, a protective control engaged and its action became the outage.
- CF-0025, recovery took longer or needed more manual work than planned.
- CF-0012, a rollback, failover or restart left the service degraded.
Each of them draws more support from thorough analyses than from short status updates. Of the incidents where CF-0025 appears, 12 are short and 27 are detailed. For CF-0026 it’s 21 against 26, and for CF-0012, 13 against 18. Every other entry in the top ten leans the other way, most of them heavily. All three counts are published as floors.
Short updates say a service was restored. They don’t say that restoring it took a manual failover, or a path nobody had walked before, or that the failover left the service degraded in some new way. So the response family is both large in this list and systematically under-described, which means its real share is higher than what’s printed. If your incident response metrics look stable but you’re still spending a lot of effort, this is a prompt to check whether your own writeups capture what recovery actually required or simply recorded when you thought it finished.
5% of the corpus carries most of what’s known about how failures spread
In this dataset, detailed analyses account for 67 of the 1,287 incidents, which is approximately 5 percent. Most of what we can say about propagation, amplification and recovery comes from those.
The most extreme case is CF-0037, separate faults presented as one incident or one fault presented as several. It appears in 13 incidents, not one of them short and 8 of them detailed, and it is the only entry in the enumeration with no support from a status update anywhere. Nobody writes “we thought this was one problem and it turned out to be three” on a status page.
Known limitations
The method page carries the limitations of the population itself: the keyword test that selected the documents, the human review of contested classifications, the split across two inference hosts, the share of conditions too vague to classify, and the corpus leaning toward large organizations that keep a public status page.
Security is nearly absent. CF-0038, an adversary obtained access and acted on the system, ranks last at 5 incidents from 5 organizations. The entry exists because the extraction found conditions about adversarial action and they needed a classification. Those cluster into a handful of long breach writeups. Organizations disclose security incidents through a different channel and on a different timetable, so the absence is about publishing practice but says nothing about frequency.
There are no confidence intervals. Most entries have too few supporting incidents for an interval to carry much meaning, so publishing one would only pretend to add precision the data doesn’t support yet.
Entries carry no relationships to each other, so how they co-occur and in what order is a separate line of work. That matters more for a ranking than for a catalog, because a list of ten conditions read on its own invites the assumption that they’re independent, but there’s overlap in the table which indicates they can’t entirely be.
What would help, and what comes next
If you publish incident reports and you’re willing to have them read as part of this corpus, write to hello@revelara.ai. If you think an entry is wrong, a category missing, or one of your own incidents doesn’t fit anywhere in the thirty-eight, say so. Three factors of the thirty-eight exist because a reviewer gave feedback on the draft.
There are some changes in mind for the second edition: a trained classifier replacing the keyword pattern that selects the population, a review sample large enough to support an accuracy figure, identification of documents that describe the same event so one incident counts once, and a single inference host for the whole extraction.
Slots are expected to move between editions. The CF identifiers are the beginning of what I hope will be a conversational anchor that lets one organization’s experience be read against another’s. Allowing the advancement of our knowledge, no matter what order they end up in.