When someone asked what the “OWASP Top 10 for Reliability” was a couple of weeks ago in a forum I frequent, I was instantly intrigued by the question. It raised more questions than answers. What would the list consist of, but also why hasn’t one been created before? Is it because the act of ranking itself is bad (rankings do have a number of inherent problems), or is it something else? I could imagine all sorts of things, from bad data, to limited analysis. But I have a fair chunk of incident data, so I set off to learn something and figure out if the results might be interesting.

A little inspiration

The OWASP Top 10 is what people go to when they want a list of security things that can go wrong. It might not be entirely common knowledge that OWASP doesn’t do the naming. MITRE does, and the Common Weakness Enumeration (CWE) list is a separate artifact that OWASP ranks against.

Four properties make that work.

  • The identifiers are stable. CWE-79 means one thing permanently.
  • The list is versioned and its history is public. It’s currently at 4.20. (v5.0 is in the works as of the time of writing.)
  • Somebody maintains it. MITRE reviews and updates the list on a regular cadence.
  • It’s been adopted widely enough that scanners emit the identifiers.

That last one is what lets tooling and rankings speak the same language, and without a description of a failure nobody else can reference a description or build from it. For reliability, most knowledge is locked in reference materials and blog posts, or shown on conference agendas, but otherwise stays trapped in senior SREs’ heads. Common terms do exist, like “incident” and “postmortem”, but many are borrowed from other disciplines and overloaded to the point where two people using them are often describing different things.

The cost of that shows up in our own data. The contributing-condition fields in the corpus are free text, and across 2,309 entries there are 2,207 distinct values, 97.6 percent of which occur exactly once. Twenty-three years after the first published taxonomy, the field still invents the words every time.

A little history

Two prior attempts stand out.

David Oppenheimer, Archana Ganapathi and David Patterson published Why Do Internet Services Fail, and What Can Be Done About It? at USITS in 2003. Three large internet services and a cause taxonomy drawn from them. It’s the ancestor of everything since.

Haryadi Gunawi and colleagues went further with Why Does the Cloud Stop Computing? at SoCC in 2016. They studied 597 unplanned outages across 32 internet services over seven years, drawn from 1,247 news reports and public incident reports, produced thirteen root cause categories, and released the underlying database as COSDB with 597 outage descriptions and 3,249 metadata tags. They performed a rigorous study, and the data release was generous.

Neither became a vocabulary. Nobody files an incident as a Gunawi category today, and no tooling emits those classifications. The interesting question is why? What does a successful candidate have to do differently?

A different lens

I speculate there were a couple of reasons previous works didn’t take hold. One technical and one structural.

The technical one is about the abstraction. Gunawi’s categories are Unknown, Upgrade, Network, Bugs, Config, Load, Cross, Power, Security, Human, Storage, Server and NatDis. Most of those name a layer of the infrastructure stack or a kind of artifact, so the question they answer is which part of the system was involved.

That’s a useful question but it’s the wrong axis for a vocabulary. Two incidents filed under Network can have nothing in common: a fiber cut is a component failing, and a routing change that partitions a network is a control action taking effect well beyond what anyone modeled. An engineer reading the second one learns nothing from the first, which defeats the purpose of having a shared name in the first place.

Stack-layer categories also age with the stack. Server meant something specific in 2016 that it doesn’t mean in a serverless architecture, and Storage has fractured into half a dozen distinct things since. A vocabulary keyed to infrastructure ends up describing a world that has moved on, which is a hard problem for a list that needs a decade to become standard.

STAMP and CAST (Leveson, MIT) classify on a different axis: where control broke down. We came up with seven top level categories, for example: “The deciding logic was wrong for the situation it met”, and “The model of the current state was wrong.”

Those are relationships between a controller and a process, and hold regardless of the substrate. A rate limiter in 2016 and an autoscaler in 2026 can occupy the same category in that structure, which is what a name needs if it’s going to survive long enough to get adopted.

The structural reason is about who does the work. Both attempts were published as findings, and a finding is complete when it appears in print. A catalog is never complete. It needs a maintainer, a version number, and identifiers stable enough that other people can cite them, and none of that looks like research, so none of it gets funded or staffed. CWE has MITRE behind it.

Security also had pressures that reliability never had: compliance regimes that required reporting in a common format, scanners that had to emit something a buyer could compare, and a vendor market where agreeing on names was worth actual money. Nothing comparable pushed reliability toward converging on a lexicon, so everyone keeps writing each cause from scratch.

The method

The corpus was frozen on 2026-09-16 at 12,507 public incident documents. The analysis ran on 2,619 of them. Here’s what came out at each stage and why.

StepDocuments
Public incident documents12,507
States a cause2,730
Removed, classifier output unreadable9
Removed, not an incident report100
Population for extraction2,621
Removed, extraction failed twice2
Analysis population2,619

Documents that don’t state a cause. Documents that don’t say why something failed can’t supply a contributing condition, so the first test is if it states one. We match case-insensitively against causal language: root cause, caused by, due to, traced to, resulted from, and similar phrases.

An earlier version used a 2,500-character length threshold as a proxy for how thorough a document was, and that was wrong in both directions, discarding short updates that state a cause while keeping long reference material that doesn’t actually describe an incident at all but heavily uses causal language as reference. Selecting on causal language reached 228 distinct publishers against 229 for the broadest length rule we tested, using 27 percent fewer documents. Keeping a reasonable breadth of sources and incident specific documentation, with considerably less noise.

Documents that aren’t incident reports. Plenty of what sits in a corpus like this isn’t an incident at all: books, framework documentation, articles about how to build reliable systems, write-ups of somebody else’s outage, announcements of planned work. Every one of the 12,507 documents was classified twice, independently, by two language models from different families running locally, and they agreed at a Cohen’s kappa of 0.853.

A document was removed if either model objected. Removing only when both agree would have kept ten more documents, but the stricter rule won because the two mistakes have different costs. Dropping a document wrongly loses one piece of evidence out of thousands. Keeping one wrongly puts a text that isn’t an incident document in a competition for a ranked position where it may hold unfair advantage.

Reference contamination is under 4 percent of the population overall. Among documents that a ranking actually counts, it rises to 12.7 percent, and for individual classes it reaches 27.6 percent. It concentrates in the classes described because reference material uses textbook vocabulary, so a chapter explaining cascading failure carries more references to cascading failure than real incidents do in practice. Removing it changes the order of the list.

When a scheduled maintenance becomes an incident. An announcement of planned maintenance and a report of planned maintenance that took production down are different documents. An early version treated those as competing kinds, which made maintenance that went wrong impossible to classify: a maintenance notice by vocabulary and an incident by outcome, landing wherever the model happened to lean. Two automated attempts failed to fix it, so it was resolved by human review, so maintenance that went wrong is now an incident report and the condition became a catalog entry in its own right.

Conditions without a verifiable quotation. Each extracted condition carries a quotation from its source document, checked mechanically after whitespace and case are normalized. If the quote isn’t in the document the condition is flagged as ungrounded and excluded. Nine of 4,370 failed that check, which is 0.21 percent, leaving 4,361 with a verified quotation.

Conditions from publishers who didn’t operate the system. Aggregators, news coverage and reference publishers describe incidents they didn’t experience, and 276 conditions came out on that basis. They’re 2.2 percent of documents but 14.6 percent of publisher values, so leaving them in would distort any count keyed to organizations far more than the document share suggests.

Conditions that name a failure without naming a mechanism. “An upstream provider issue” and “an infrastructure outage” state that something went wrong and stop there. There’s nothing to classify, so those are held out of the ranking and counted separately: 197 conditions, 4.7 percent. A further 29, or 0.7 percent, describe a mechanism that no entry covers.

What’s in the catalog

Thirty-eight contributing conditions, each with a stable identifier of the form CF-NNNN, organized into seven categories adapted from the CAST causal factor structure.

  1. The deciding logic was wrong for the situation it met
  2. The model of the current state was wrong
  3. The condition could not be observed in time
  4. An action was issued and did not take effect
  5. Several controllers acted on the same process
  6. The system drifted outside the conditions its controls assumed
  7. Conditions outside the organization’s authority acted on it

Every entry, with its definition, is in the full catalog.

There are three levels: category, contributing condition, and variant. Only the middle level is ranked, because the categories are too broad to act on and the variants have too few examples to count.

Entries are written as behaviors and never blamefully attributed as absences. “The caller waited on the dependency indefinitely” is an entry; “no timeout was set” isn’t. An absence invites counterfactual reasoning about what the operators should have had in place, which nobody outside the organization can settle from a publicly published incident document, where a behavior states what was observed from the system as it actually worked at the time of the incident.

Entries also never locate a failure in a person. “Separate faults presented as one incident” is an entry, and “the team misdiagnosed the fault” is not. The first says what the system did, which can be checked against a document. The second is normative language about an idea or norm of behavior which can lead to speculation in incident analysis.

Entries describe conditions inside the reporting organization’s control structure, or acting on it from outside, and they intentionally stop where that organization’s account stops. A facility losing power and a provider becoming unavailable each had their own contributing conditions. It’s important to understand that the corpus only holds the customer’s view of them, so those entries can’t decompose further. Those events may have causes, but they aren’t captured in the data we have.

CF-0026 is a good example, it reads “a protective control engaged and its action became the outage,” which covers rate limiting, load shedding, an automated account suspension, and a storage integrity check that blocked traffic, at four different organizations running very different tech stacks. Still, an engineer can read that causal failure definition and know immediately whether they’ve seen it or something like it.

The categorization is CAST-inspired but it isn’t a full CAST analysis. A full analysis requires people who know a system to examine its actual control structure. These entries are candidates for engineering teams to review, not findings about anybody’s system.

Known limitations

The causal-language test is itself a keyword test and carries the usual problems. It admits documents like “delayed due to scheduled maintenance,” which don’t describe failures, and it misses analyses that explain failures without using causal vocabulary. This edition doesn’t measure these cases, and doesn’t try to solve for them. The next edition might replace the pattern matching with a trained classifier and keep the pattern published as a baseline.

A person reviewed twenty contested document classifications instead of a random sample of the population. A random set of twenty had a high likelihood of being mostly routine status updates. This doesn’t support making an accuracy claim or measurement.

Extraction ran across two inference hosts, both running the same model at different quantization. Output volume was indistinguishable, however one host assigned late-chain roles approximately 50 percent more often than the other. The cause is unresolved, so every row records its provenance for future investigation.

The corpus over-represents large organizations, regulated industries and anyone who maintains a public status page. A condition that’s common at ten-person companies can be invisible here for reasons that have nothing to do with how often it happens.

The catalog is incomplete, and the earlier figures are the size of it: 4.7 percent of conditions name a failure with no mechanism, and 0.7 percent describe a mechanism no entry covers. Entries also carry no relationships to one another, so how they co-occur and in what order is a separate piece of work.

What would help

If you publish incident reports and you’re willing to have them read as part of this corpus, please get in touch. If an entry is wrong, you think a category is missing, or one of your own incidents doesn’t fit anywhere in the thirty-eight, let us know. Three of the thirty-eight are in there because a reviewer gave feedback on the draft and it improved the classification.

The catalog is meant to last. A ranking built on it will move from edition to edition, but stable names are what lets one organization’s experience become another’s proactive practice.

The Causal Factor Enumeration Catalog (v1.0)

The Reliability Top 10, First Edition