Causal Factor Enumeration v1: The Full Catalog
All 38 contributing conditions in the first edition of the catalog, grouped into seven categories. Every entry has a stable CF identifier and its own link, so you can cite it.
This page lists every entry in the first edition of the catalog from Building a Causal Factor Enumeration for Reliability. That post covers the prior work, the method and the known limitations. This page is the reference. The Reliability Top 10 ranks ten of these entries by the number of public incidents each appears in.
Edition v1 · curated 2026-09-22 · 38 entries · 7 categories
How to read an entry
Each entry has an identifier, a name and a definition.
The identifier is stable. CF-0026 keeps its meaning, and a retired identifier stays retired. The numbers aren’t a rank.
The name states what the system did. It never names an absence, such as a missing timeout, and it never puts the failure in a person, so you can check it against an incident report.
The definition sets the boundary of the entry: what it covers and where it stops.
Some entries also carry a curation note. The note records a decision about that entry, for example why it stays separate from a broader one.
Each category opens with a short index of its entries. Under each entry, Details holds the rest: the category, any curation note, quotes from public incident reports that the entry covers, and how to cite it. Each quote names the organization that published the report, with the report’s date, and the name links to the report. Publishers sometimes take old reports down, so a few links may no longer resolve. The quotes are verbatim. A quote that ends in an ellipsis stops before the end of its sentence.
The seven categories are adapted from the CAST causal factor structure. Only the entries are ranked, because a category is too broad to act on.
Each entry has its own link. To cite one, give the identifier, for example CF-0026, and link to the entry on this page.
Category 1: The deciding logic was wrong for the situation it met
In this category
- CF-0001 A configuration change altered behavior beyond its intended scope
- CF-0002 A software release introduced a defect into production
- CF-0003 A planned change was carried out as intended and had an effect nobody anticipated
- CF-0004 A data or schema migration removed or altered data that running code still used
- CF-0005 Code processed input outside the range it was written for
- CF-0006 A value reached a boundary the code or schema could not represent
- CF-0007 A cleanup or deletion operation removed data or resources still in use
- CF-0026 A protective control engaged and its action became the outage
- CF-0039 A written procedure was ambiguous or out of date
CF-0001 A configuration change altered behavior beyond its intended scope
A change to configuration, routing or feature flags reached more of the system than intended, or had effects beyond those its authors modeled. Validation passed against the intended target.
Details
Observed in incident reports
| Microsoft 2025-10-30 | “A specific sequence of customer configuration changes, performed across two different control plane build versions, resulted in incompatible customer configuration metadata being generated.” |
| Google 2026-07-14 | “The disruption was traced to a network configuration update that introduced a conflict, causing inter-site communication failures, which triggered VMware Stretched Cluster failover events.” |
| Atlassian 2026-05-02 | “A schema configuration change was deployed to the API gateway layer before the corresponding application-level change had been deployed across all production environments.” |
Cite as
CF-0001 revelara.ai/blog/causal-factor-enumeration/#cf-0001
CF-0002 A software release introduced a defect into production
A new version of code or a dependency reached production carrying a fault.
Details
Observed in incident reports
| GitHub 2026-06-05 | “These issues were triggered by a change to an internal authorization component that did not correctly resolve access for user-to-server tokens against organization-owned repositories.” |
| Duo Security 2026-02-04 | “A recent software update Duo made related to certificate security improvements inadvertently changed the order of operations for user sessions in our legacy enrollment experience.” |
| CircleCI 2021-05-21 | “This upgrade therefore resulted in some messages hitting the acknowledgment timeout, and the RabbitMQ channel associated with the corresponding consumer to be closed.” |
Cite as
CF-0002 revelara.ai/blog/causal-factor-enumeration/#cf-0002
CF-0003 A planned change was carried out as intended and had an effect nobody anticipated
Maintenance or planned work proceeded as designed and still caused unexpected impact. Includes emergency maintenance.
Details
Observed in incident reports
| HiBob 2025-12-04 | “During routine maintenance performed on a different, non-production cluster, a configuration that was shared behind the scenes caused an unexpected and unplanned impact on production.” |
| Google 2026-08-20 | “The disruption originated during scheduled fiber optic maintenance, which unexpectedly compromised network capacity between data centers within the us-west1 region.” |
| Box 2025-12-18 | “the migration process exposed some observability and guardrail gaps in the pipeline that would have prevented this issue from happening and that extended the impact” |
Cite as
CF-0003 revelara.ai/blog/causal-factor-enumeration/#cf-0003
CF-0004 A data or schema migration removed or altered data that running code still used
A change to stored data or its structure left running code reading something that no longer existed in the form it expected.
Details
Observed in incident reports
| Amazon Web Services 2012-12-24 | “After this data was deleted, the ELB control plane began experiencing high latency and error rates for API calls to manage ELB load balancers.” |
| GitHub 2026-08-12 | “During a database migration, two indexes were removed while application settings still referenced them, causing affected requests to fail.” |
| KnowBe4 2026-06-18 | “As a result, phishing campaigns still using legacy phishing categories were unable to load the Phishing Security Test Reports page.” |
Cite as
CF-0004 revelara.ai/blog/causal-factor-enumeration/#cf-0004
CF-0005 Code processed input outside the range it was written for
Malformed, oversized or unexpected input reached a code path that handled it incorrectly.
Details
Observed in incident reports
| HiBob 2026-08-31 | “A regression caused validation to run across all history table entries, including historical rows containing archived values, which caused updates to fail.” |
| PubNub 2025-10-06 | “This issue occurred because our event processing service did not correctly handle a malformed message format, which caused the processing queue to stall.” |
| CircleCI 2026-04-06 | “the animation logic reads open/close state from a specific React context that is only available when modals are opened via a particular component” |
Cite as
CF-0005 revelara.ai/blog/causal-factor-enumeration/#cf-0005
CF-0006 A value reached a boundary the code or schema could not represent
A counter, identifier, clock or quota reached a value outside the range its type or logic represents.
Details
Observed in incident reports
| GitHub 2026-05-06 | “Once the values in the primary table passed the available 32-bit ID space in the lookup table, attempts to create new review threads began failing” |
| Cloudflare 2017-02-18 | “reaching the end of a buffer was checked using the equality operator and a pointer was able to step past the end of the buffer” |
| Linode 2026-04-24 | “triggered the generation of configuration identifiers beyond the supported threshold, resulting in further degradation” |
Cite as
CF-0006 revelara.ai/blog/causal-factor-enumeration/#cf-0006
CF-0007 A cleanup or deletion operation removed data or resources still in use
Routine cleanup, deletion or retirement targeted resources that were still active.
Details
Observed in incident reports
| Front 2025-04-29 | “This was the result of a human error during a planned database maintenance operation that resulted in one database shard becoming unrecoverable” |
| Mixpanel 2026-05-14 | “Our cleanup tooling did not programmatically enforce the safety invariant that the date filter must be strictly before the reference snapshot.” |
| Amazon Web Services 2025-10-19 | “it then invoked the plan clean-up process, which identifies plans that are significantly older than the one it just applied and deletes them” |
Cite as
CF-0007 revelara.ai/blog/causal-factor-enumeration/#cf-0007
CF-0026 A protective control engaged and its action became the outage
Rate limiting, load shedding, lockout, suspension or another safeguard acted as designed, and that action is what users experienced as the failure.
Details
Observed in incident reports
| Sumsub 2025-08-01 | “Temporarily blocking traffic from specific Autonomous System Numbers (ASNs) across the entire sumsub.com domain, impacting services such as the API, Cockpit, and website.” |
| Google 2023-10-29 | “our networking systems correctly triaged the traffic overload and dropped larger, less latency-sensitive traffic in order to preserve smaller latency-sensitive traffic flows” |
| Microsoft 2025-11-05 | “The recovery process prioritizes data integrity and durability over availability, resulting in customer traffic being blocked until all integrity validations was completed” |
Cite as
CF-0026 revelara.ai/blog/causal-factor-enumeration/#cf-0026
CF-0039 A written procedure was ambiguous or out of date
A runbook or documented process gave instructions that did not match the system it described, or left an order or timing open to more than one reading.
Details
Curation note
Added 2026-09-23. A testing-coverage entry was considered and held: three supporting conditions, and the wording available for it described an absence.
Observed in incident reports
| Google 2026-09-01 | “the complete list of transceiver replacements across all routers to be issued to the technician without instructions to sequence the work one router at a time” |
| Mixpanel 2026-05-14 | “It had been recently authored to handle the new file-storage code path and had not gone through a formal review before being used.” |
| CircleCI 2025-04-04 | “did not prioritize investigating WAF configuration expecting that any changes would have gone through our Terraform pipeline” |
Cite as
CF-0039 revelara.ai/blog/causal-factor-enumeration/#cf-0039
Category 2: The model of the current state was wrong
In this category
- CF-0008 Cached or derived state no longer matched its source of truth
- CF-0009 A failover promoted a replica whose state differed from the primary's
- CF-0010 A health signal reported healthy while the component was failing
- CF-0037 Separate faults presented as one incident, or one fault presented as several
CF-0008 Cached or derived state no longer matched its source of truth
Caches, indexes or other derived data served values that had diverged from the authoritative source.
Details
Observed in incident reports
| Slack 2026-05-27 | “Some users continued to experience permissions-related issues, including inability to manage channels and Canvas edits not saving, as a result of residual stale cache data.” |
| LaunchDarkly 2026-07-10 | “SDKs and Relay Proxy instances which attempted to establish a connection to LaunchDarkly on July 10 between 09:11 and 11:42 AM PT were intermittently affected.” |
| Google 2025-07-18 | “an operational topology change while the network control plane was in a failed open state caused our network fabric's topology information to become stale” |
Cite as
CF-0008 revelara.ai/blog/causal-factor-enumeration/#cf-0008
CF-0009 A failover promoted a replica whose state differed from the primary's
Promotion of a follower exposed state (sequences, unreplicated writes) that differed from what the old primary held.
Details
Observed in incident reports
| Cloudflare 2020-11-02 | “There was a defect in our cluster management system that requires a rebuild of all database replicas when a new primary database is promoted.” |
| Razorpay 2019-12-01 | “the standby (new primary) never received this data; which is why we could not find it in the master instance after the failover” |
| GitHub 2018-10-21 | “The database servers in the US East Coast data center contained a brief period of writes that had not been replicated” |
Cite as
CF-0009 revelara.ai/blog/causal-factor-enumeration/#cf-0009
CF-0010 A health signal reported healthy while the component was failing
Health checks or readiness signals kept reporting a failing component as healthy, so traffic and work kept going to it.
Details
Observed in incident reports
| Amazon Web Services 2021-09-02 | “While these devices were not correctly forwarding traffic, they were not being removed from the network through the normal automated processes” |
| PubNub 2026-08-25 | “Because that response appeared successful, the publishing layer received no error, and messages sent to that address were routed incorrectly.” |
| Microsoft 2026-02-02 | “the health assessment defect caused the incorrect targeting state to propagate broadly across public cloud regions in a short time window” |
Cite as
CF-0010 revelara.ai/blog/causal-factor-enumeration/#cf-0010
CF-0037 Separate faults presented as one incident, or one fault presented as several
Concurrent or overlapping faults produced signals that did not distinguish them. One fault's symptoms were read as another's, and recovery followed the wrong one. Includes a mitigation that appeared to work and pointed away from the condition still failing.
Details
Curation note
Added 2026-09-23 after external review. Distinct from CF-0010 (a signal reporting healthy) and CF-0011 (a signal arriving late): this entry concerns the identity and count of faults.
Observed in incident reports
| CircleCI 2025-04-04 | “This occurred just as our teams were concluding another unrelated incident, which initially caused some confusion about whether the issues might be connected.” |
| PubNub 2025-06-30 | “Unfortunately, the errors we initially encountered pointed us in incorrect directions, causing the investigation to take longer than we normally strive for.” |
| Cloudflare 2022-10-25 | “At first glance, the errors looked like they were caused by a different system that had started a release some time before.” |
Cite as
CF-0037 revelara.ai/blog/causal-factor-enumeration/#cf-0037
Category 3: The condition could not be observed in time
In this category
CF-0011 Monitoring reported the failure late, or reported a normal state during it
The failure was not visible to operators in time to act on, because monitoring lagged, was itself degraded, or measured the wrong thing.
Details
Observed in incident reports
| Amazon Web Services 2024-07-30 | “the cell management system incorrectly determined that the healthy hosts were unhealthy and began redistributing shards that those hosts had been processing to other hosts” |
| GitHub 2026-02-02 | “This outage was caused by a loss in telemetry that cascaded to mistakenly applying security policies to backend storage accounts in our underlying compute provider.” |
| Mixpanel 2026-08-26 | “Alerting on storage volumes flagged the growth but was not escalated as critical on a per-server basis, which delayed detection until query failures began.” |
Cite as
CF-0011 revelara.ai/blog/causal-factor-enumeration/#cf-0011
Category 4: An action was issued and did not take effect
In this category
- CF-0012 A rollback, failover or restart left the service degraded
- CF-0025 Recovery took longer or needed more manual work than planned
- CF-0028 An internal dependency failed and the failure propagated to its callers
- CF-0030 A network path lost connectivity or capacity
- CF-0031 Name resolution returned wrong answers or none
- CF-0032 An authentication or identity service failed
- CF-0035 A host, disk or other hardware component failed
CF-0012 A rollback, failover or restart left the service degraded
A corrective action was taken and the service did not return to its prior state.
Details
Observed in incident reports
| Linode 2026-08-13 | “Akamai engineering teams identified a configuration discrepancy on the secondary caching infrastructure node that prevented it from absorbing the full traffic load after the failover.” |
| Honeycomb 2026-05-01 | “The database failover that caused query issues earlier in the EU has also had knock-on effects on our Activity Log beta feature.” |
| PubNub 2025-10-20 | “undefined steps in some of our failover processes and delays accessing some tools due to the provider issue, existing connections for some services remained degraded” |
Cite as
CF-0012 revelara.ai/blog/causal-factor-enumeration/#cf-0012
CF-0025 Recovery took longer or needed more manual work than planned
Restoring service required steps, time or decisions that the plan did not anticipate.
Details
Observed in incident reports
| Microsoft 2025-11-05 | “Recovery of storage scale units was prolonged because some storage servers entered a degraded state, necessitating sequential validation before they could be brought back online.” |
| GitLab 2017-01-31 | “pg_basebackup will silently wait for a master to initiate the replication progress, according to another production engineer this can take up to 10 minutes.” |
| Mixpanel 2026-05-14 | “During this window, fewer than 2% of customers were still affected — specifically, those whose deleted files had not yet been fully restored from backup.” |
Cite as
CF-0025 revelara.ai/blog/causal-factor-enumeration/#cf-0025
CF-0028 An internal dependency failed and the failure propagated to its callers data_store
A service inside the organization failed, and services calling it failed or slowed with it. Includes a database or storage system that stopped serving (tag: data_store); merged from CF-0029 at curation.
Details
Observed in incident reports
| GitHub 2026-04-27 | “This resulted in intermittent failures for services relying on our search data including Issues, Pull Requests, Projects, Repositories, Actions, Package Registry and Dependabot Alerts.” |
| Pusher 2025-10-20 | “When replacement Redis instances in the US3 redis-main cluster were launched, they failed initialization checks and repeatedly restarted, rendering the cluster unavailable.” |
| Snowflake 2026-06-16 | “This triggered a cascade of connectivity issues in backend components supporting the Snowsight platform, resulting in service disruptions and intermittent access problems.” |
Cite as
CF-0028 revelara.ai/blog/causal-factor-enumeration/#cf-0028
CF-0030 A network path lost connectivity or capacity
Component-shaped: packet loss, partitions or reduced network capacity.
Details
Curation note
Kept as its own entry at curation: an infrastructure layer with its own failure behavior, not a generic dependency.
Observed in incident reports
| Google 2026-08-20 | “This underlying network degradation subsequently impacted higher-level service components through severe packet loss, request throttling, and increased latency across inter-campus dependencies.” |
| Microsoft 2026-01-10 | “which caused newly deployed or updated VMs within this Availability Zone (on hosts managed by impacted SLB resources) to have continued intermittent outbound connectivity issues” |
| Snowflake 2026-07-10 | “the load balancing infrastructure responsible for managing connections to Snowsight temporarily entered a degraded state, preventing requests from being processed successfully” |
Cite as
CF-0030 revelara.ai/blog/causal-factor-enumeration/#cf-0030
CF-0031 Name resolution returned wrong answers or none
Component-shaped: DNS or other address resolution failed.
Details
Curation note
Kept as its own entry at curation: an infrastructure layer with its own failure behavior, not a generic dependency.
Observed in incident reports
| GitHub 2026-04-23 | “multiple GitHub services experienced elevated error rates and degraded performance due to DNS resolution failures originating from our DNS infrastructure in our VA3 datacenter” |
| Cloudflare 2023-10-04 | “some of Cloudflare’s resolver systems stopped being able to validate DNSSEC signatures and as a result started sending error responses (SERVFAIL)” |
| Amazon Web Services 2025-10-19 | “The root cause of this issue was a latent race condition in the DynamoDB DNS management system that resulted in an incorrect empty DNS record” |
Cite as
CF-0031 revelara.ai/blog/causal-factor-enumeration/#cf-0031
CF-0032 An authentication or identity service failed
Component-shaped: login, token or identity verification failed for dependent services.
Details
Curation note
Kept as its own entry at curation: an infrastructure layer with its own failure behavior, not a generic dependency.
Observed in incident reports
| Atlassian 2026-07-06 | “the elevated identity latency contributed to a cascading failure in a downstream user search service, degrading user search and user picker experiences across multiple products” |
| Lever 2026-04-21 | “As a result, the calendar service was unable to read newly issued Microsoft access tokens, causing scheduling requests to fail for Office365-connected accounts.” |
| CircleCI 2026-07-14 | “The Bitbucket permissions check is used broadly across our platform, and these permissions failures also affected some of our internal job-processing systems.” |
Cite as
CF-0032 revelara.ai/blog/causal-factor-enumeration/#cf-0032
CF-0035 A host, disk or other hardware component failed
Physical or virtual infrastructure failed underneath a service.
Details
Observed in incident reports
| Google 2025-03-30 | “The UPS system, which relies on batteries to bridge the gap between utility power loss and generator power activation, experienced a critical battery failure.” |
| JFrog 2023-08-30 | “a small number resources are currently being replaced or repaired due to thermal damage in order to bring the remaining impacted storage accounts back online” |
| GitHub 2026-08-26 | “runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing” |
Cite as
CF-0035 revelara.ai/blog/causal-factor-enumeration/#cf-0035
Category 5: Several controllers acted on the same process
In this category
- CF-0013 Two automated processes acted on the same resource and one undid the other
- CF-0014 Scheduling or placement concentrated load onto one node or partition
CF-0013 Two automated processes acted on the same resource and one undid the other
Independent automation, each correct in isolation, made conflicting changes to shared state.
Details
Observed in incident reports
| GitHub 2020-02-19 | “we also encountered a race condition between our process manager and service configurations which slowed our ability to change our file limit to 1048576” |
| Atlassian 2026-05-14 | “a race condition in our internal deployment orchestration platform during a routine rollback operation of a core identity service in the us-east region” |
| Amazon Web Services 2025-10-19 | “the first Enactor (which had been unusually delayed) applied its much older plan to the regional DDB endpoint, overwriting the newer plan” |
Cite as
CF-0013 revelara.ai/blog/causal-factor-enumeration/#cf-0013
CF-0014 Scheduling or placement concentrated load onto one node or partition
A scheduler or placement decision put too much work on one node, shard or partition.
Details
Observed in incident reports
| Google 2026-08-20 | “Automated rerouting mechanisms failed to properly redistribute traffic to alternate capacity, resulting in network congestion as volumes exceeded available bandwidth in the affected area.” |
| Amazon Web Services 2024-07-30 | “The cell management system focuses on balancing work based on throughput and other I/O dimensions that impact the ability of hosts to handle additional shards” |
| Microsoft 2026-03-10 | “traffic routes were temporarily being determined using incomplete data, which led to a disproportionate amount of traffic being directed to a limited set” |
Cite as
CF-0014 revelara.ai/blog/causal-factor-enumeration/#cf-0014
Category 6: The system drifted outside the conditions its controls assumed
In this category
- CF-0015 Demand exceeded the capacity the service was provisioned for
- CF-0016 Requests queued for a shared resource until it stopped serving any of them
- CF-0017 Hosts or instances exhausted their memory
- CF-0018 Expensive or runaway queries saturated a data store
- CF-0019 A provisioned limit on connections, quotas or instances was reached
- CF-0020 Demand grew faster than new capacity came online
- CF-0021 Retries and reconnection traffic multiplied load on a recovering service
- CF-0022 A cache miss or cold cache sent full load to the backing store
- CF-0023 A bulk or background job consumed resources shared with serving traffic
- CF-0024 Work accumulated faster than it drained and kept the service degraded after the trigger cleared
- CF-0027 A certificate or credential reached its expiry while still in use
CF-0015 Demand exceeded the capacity the service was provisioned for
Traffic or workload rose beyond provisioned capacity, including unexpected or abusive sources.
Details
Observed in incident reports
| Atlassian 2026-04-08 | “On a subset of clusters configured for high‑density workloads, the increased reservations exceeded available node capacity interrupting search and related experiences for affected customers.” |
| Microsoft 2026-02-02 | “The resulting surge in resource management operations generated a significant and sudden increase in requests to the backend service supporting Managed Identities for Azure resources.” |
| Airtable 2025-10-20 | “Airtable is operational but certain features may experience degraded performance. We are working with our infrastructure provider to address capacity constraints.” |
Cite as
CF-0015 revelara.ai/blog/causal-factor-enumeration/#cf-0015
CF-0016 Requests queued for a shared resource until it stopped serving any of them
Connections, threads, pools or locks were held by waiting work until the resource granted nothing.
Details
Observed in incident reports
| xMatters 2019-07-28 | “During this brief period, the Integration Builder was accepting notification requests but not processing outbound notifications, which resulted in the delays experienced by some customers.” |
| CircleCI 2019-04-02 | “It turned out that checking the index requires a database level lock which, while only held for a short period, can cause significant contention.” |
| GitHub 2026-02-02 | “This caused replication delays, resulting in a similar cascade of failures and again leading to connection exhaustion in the Git HTTPS proxy.” |
Cite as
CF-0016 revelara.ai/blog/causal-factor-enumeration/#cf-0016
CF-0017 Hosts or instances exhausted their memory
Memory use grew until processes became unresponsive or were killed.
Details
Observed in incident reports
| KnowBe4 2026-02-10 | “a subset of Workspace instances across multiple regions became slow or temporarily unresponsive due to memory exhaustion on their underlying virtual machines” |
| Honeycomb 2024-06-24 | “This stampede resulted in significant memory utilization growth, causing some services to crash loop with out-of-memory exceptions while retrieving schemas” |
| Atlassian 2026-06-12 | “Idle connections accumulated across clusters rather than being closed, creating a connection leak and increasing memory pressure on the database layer.” |
Cite as
CF-0017 revelara.ai/blog/causal-factor-enumeration/#cf-0017
CF-0018 Expensive or runaway queries saturated a data store
Costly, parallel or unbounded queries consumed a data store's capacity.
Details
Observed in incident reports
| KnowBe4 2026-07-01 | “Upon deployment, this specific query pattern bypassed optimal indexing strategies, resulting in full table scans and highly extended execution times on the database reader infrastructure.” |
| Front 2025-11-21 | “In this cell a particular workload caused high database contention that ultimately timed out and then started over, which prevented us from escaping the problem.” |
| 1Password 2025-09-03 | “A poorly performing cache operation was triggered repeatedly in a short period of time across multiple servers, leading directly to greatly delayed responses.” |
Cite as
CF-0018 revelara.ai/blog/causal-factor-enumeration/#cf-0018
CF-0019 A provisioned limit on connections, quotas or instances was reached
A configured ceiling (connection count, quota, instance limit) was hit under ordinary or elevated load.
Details
Observed in incident reports
| Microsoft 2026-05-29 | “In a small number of regions, recovery was further constrained because scale-out options were limited by available compute quota for replacement virtual machine SKUs.” |
| xMatters 2019-11-07 | “This limited the ability of other databases in the cluster to accept new requests, resulting in intermittent access to the web user interface” |
| GitHub 2023-03-16 | “increased load during peak hours on our mysql1 database, causing our database proxying technology to reach its maximum number of connections” |
Cite as
CF-0019 revelara.ai/blog/causal-factor-enumeration/#cf-0019
CF-0020 Demand grew faster than new capacity came online
Scaling or provisioning responded, but more slowly than load arrived.
Details
Observed in incident reports
| Atlassian 2026-06-01 | “the cloud infrastructure was unable to provision additional nodes in that availability zone due to insufficient underlying capacity” |
| Salesloft 2025-10-20 | “As a result, Salesloft users may experience intermittent issues with the Salesloft Dialer while this scaling process continues.” |
| CircleCI 2021-11-08 | “the new nodes took longer than expected to join the fleet, which led to many of them being marked unhealthy and terminating” |
Cite as
CF-0020 revelara.ai/blog/causal-factor-enumeration/#cf-0020
CF-0021 Retries and reconnection traffic multiplied load on a recovering service
Clients retried or reconnected together, adding load at the moment the service was least able to absorb it.
Details
Observed in incident reports
| Microsoft 2026-02-02 | “Automatic retry behaviors in upstream components, designed to provide resiliency under transient failure conditions, amplified traffic against the managed identity backend.” |
| Linode 2026-05-12 | “the accumulated backlog of automatic retries from the telemetry sender's built-in retry mechanism created a sudden, massive spike in traffic against the ingest endpoint” |
| Google 2025-05-29 | “as Service Control tasks restarted, it created a herd effect on the underlying infrastructure it depends on (i.e. that Spanner table), overloading the infrastructure.” |
Cite as
CF-0021 revelara.ai/blog/causal-factor-enumeration/#cf-0021
CF-0022 A cache miss or cold cache sent full load to the backing store
Loss or expiry of cached data sent traffic straight to the slower system behind it.
Details
Observed in incident reports
| Atlassian 2026-09-03 | “As cache requests timed out, the authorization service made more fallback requests to upstream systems, further increasing load and causing circuit breakers and timeouts.” |
| Honeycomb 2026-06-24 | “failed remote cache retrievals, and subsequently, a MySQL stampede due to several services falling back to retrieving schemas from the database” |
| GitHub 2018-09-11 | “The system load generated by the site’s query load on a cold cache soon caused Percona Replication Manager’s health checks to fail again” |
Cite as
CF-0022 revelara.ai/blog/causal-factor-enumeration/#cf-0022
CF-0023 A bulk or background job consumed resources shared with serving traffic
Batch, backfill, replay or deletion work competed with user-facing traffic for the same resources.
Details
Observed in incident reports
| CircleCI 2026-07-02 | “The incident originated from a combination of internal maintenance and customer-initiated project deletions running concurrently, which caused elevated load on a data service.” |
| Wasabi 2026-05-21 | “The host failures occurred under elevated internal system activity associated with platform health-management operations and concurrent storage reclamation workflows” |
| Mixpanel 2026-08-26 | “The replay tooling wrote its files to the query servers' local disks, so runaway replay data exhausted storage capacity the servers need to answer queries.” |
Cite as
CF-0023 revelara.ai/blog/causal-factor-enumeration/#cf-0023
CF-0024 Work accumulated faster than it drained and kept the service degraded after the trigger cleared
A backlog built during the fault and kept the service slow or wrong after the fault itself was gone.
Details
Observed in incident reports
| Microsoft 2026-02-02 | “The re-enablement of the VM extension package storage layer allowed multiple previously blocked workflows to resume concurrently, across infrastructure orchestration systems and dependent services.” |
| Atlassian 2026-05-08 | “events generated during the impact window were queued for replay, and some background services remained delayed until that replay and related validation work completed” |
| Snowflake 2026-08-03 | “some serverless tasks may have experienced residual queuing or failures until the backlog of requests which accumulated during the impact windows was processed” |
Cite as
CF-0024 revelara.ai/blog/causal-factor-enumeration/#cf-0024
CF-0027 A certificate or credential reached its expiry while still in use
A time-limited credential expired in production.
Details
Observed in incident reports
| Lever 2026-07-13 | “Because this was a shared credential used for the Microsoft integration, the expiration affected customers using Office 365 broadly rather than a specific account.” |
| GitHub 2026-07-19 | “The incident was caused by a certificate lifecycle management failure in a subset of internal services, resulting in an SSL certificate expiration” |
| CircleCI 2026-07-14 | “workflow status updates for Bitbucket pipelines belonging to customers with an expired access token began to be dropped” |
Cite as
CF-0027 revelara.ai/blog/causal-factor-enumeration/#cf-0027
Category 7: Conditions outside the organization's authority acted on it
Choosing a provider, and deciding whether to run without a second one, are control actions inside the organization, so entries about those decisions belong in Category 1. This category covers what then acted on the system from outside that authority.
In this category
- CF-0033 An external provider the service depends on became unavailable or degraded
- CF-0034 A facility lost power or cooling, or was physically damaged
- CF-0036 An external authority or policy restricted the service
- CF-0038 An adversary obtained access and acted on the system
CF-0033 An external provider the service depends on became unavailable or degraded
A provider outside the organization failed. Most exposed to one incident being reported by many customers.
Details
Curation note
Kept as its own entry at curation: an infrastructure layer with its own failure behavior, not a generic dependency.
Observed in incident reports
| Linode 2026-05-14 | “The degraded performance observed two days later (Window 2) was the result of a separate, independent infrastructure incident within our telemetry partner's message-queuing service…” |
| Box 2025-09-18 | “Box’s cloud provider experienced an issue that caused global load balancer services to incorrectly return HTTP 502 responses for a small percentage of user requests” |
| Snowflake 2026-06-15 | “Backend infrastructure that supports authentication to the Snowsight service experienced a degradation due to a third-party cloud platform infrastructure issue.” |
Cite as
CF-0033 revelara.ai/blog/causal-factor-enumeration/#cf-0033
CF-0034 A facility lost power or cooling, or was physically damaged
Utility power, backup power, cooling, or a physical event such as severe weather affected a facility.
Details
Observed in incident reports
| Microsoft 2026-05-30 | “As temperatures rose beyond safe operating thresholds, a subset of cloud infrastructure automatically shut down to prevent physical damage and preserve data integrity.” |
| Google 2025-03-30 | “This power outage triggered a cascading failure within the uninterruptible power supply (UPS) system responsible for maintaining power to the zone during such events.” |
| Amazon Web Services 2023-06-13 | “multiple storage nodes became unavailable simultaneously in a single data center (after power was lost to the servers on which these nodes lived)” |
Cite as
CF-0034 revelara.ai/blog/causal-factor-enumeration/#cf-0034
CF-0036 An external authority or policy restricted the service
Licensing, regulation, network providers or reputation systems restricted operation.
Details
Observed in incident reports
| Smartsheet 2025-12-02 | “The account suspension was triggered by our vendor's compliance team due to reported email abuse originating from a compromised customer account on our platform.” |
| KnowBe4 2026-01-21 | “Some messages routed through this restricted address experienced temporary delays while we attempted delivery using other sending addresses” |
| GitHub 2026-02-02 | “Those policies blocked access to critical VM metadata, causing all VM create, delete, reimage, and other operations to fail.” |
Cite as
CF-0036 revelara.ai/blog/causal-factor-enumeration/#cf-0036
CF-0038 An adversary obtained access and acted on the system
A third party gained access it was not granted and took actions that degraded or exposed the service. Includes compromised accounts, stolen credentials or sessions, and malicious code reaching production through a supply chain.
Details
Curation note
Added 2026-09-23. 19 of the 70 unplaced conditions were adversarial; a further 15 had been absorbed into CF-0002, CF-0030 and CF-0005.
Observed in incident reports
| CircleCI 2023-01-04 | “an unauthorized third party leveraged malware deployed to a CircleCI engineer’s laptop in order to steal a valid, 2FA-backed SSO session.” |
| Heroku 2022-05-16 | “the actor accessed and exfiltrated data from the database storing usernames and uniquely hashed and salted passwords” |
| Pusher 2025-10-27 | “unusual activity on the cluster, likely caused by at least one customer experiencing a cyber attack” |
Cite as
CF-0038 revelara.ai/blog/causal-factor-enumeration/#cf-0038
Retired and excluded identifiers
Identifiers are never reused. An entry that is merged keeps its number, marked retired, so a gap in the numbering always has a reason.
- CF-0029 Retired. Merged into CF-0028 (tag data_store).
- CF-X Not an entry. The statement names a failure without a mechanism. Excluded from ranking pending long-tail analysis.
Corrections
If an entry is wrong, or an incident you know doesn’t fit anywhere in the thirty-eight, write to hello@revelara.ai.