Infrabench / Crash Cart
Tool 02 — Fleet failure triage

You cannot pull them all. Which ones this week?

A failure rate says how many drives will die this year, not which ones. Crash Cart turns thirteen years of public fleet data into a short list: the drives worth a service trip this week, with the evidence behind each one. Every constant is in the method.

The console

Fleet profile Redundancy scheme
My fleet is a of drives in groups.
The crew can pull drives a week.

The answer
Pull this week.
Assumptions
percent of drives

MB/s

The work, ranked
By data at risk
Reality check
How it works

Four steps from a base rate to a work order

Each row of the queue is one drive population carried through the same four multiplications. You can do the arithmetic yourself.

STEP 01 / BASE RATE

Start from the published rate

Each model starts at its published lifetime failure rate, carrying the drive days behind it, so a thin record shows up as a wide confidence interval rather than as false certainty.

STEP 02 / AGE

Adjust for age

Failure rates follow a bathtub: a bump in the first months, a floor near year two, a climb that gets serious after year five. Each cohort gets the multiplier for its age, not the fleet average.

STEP 03 / EVIDENCE

Weigh the SMART evidence

In the public fleet, 77 percent of failed drives showed one of five SMART warnings, against 4 percent of healthy ones. A flagged drive is about 18 times likelier to fail; a clean one, a quarter as likely.

STEP 04 / BLAST RADIUS

Weigh what a failure costs

Every failure starts a rebuild that exposes the whole group. Multiply the chance of failure by the data sitting behind it, and the queue sorts itself: risk times blast radius.

The wear curve and a worked example
The wear curve

The bathtub, with your fleet on it

The age multiplier applied at step 01, drawn against drive age. Each dot is one cohort in the profile you picked, sized by drive count. A fleet sitting on the right-hand climb is not a maintenance problem, it is a refresh decision.

Failure rate multiplier against drive age A shallow bathtub curve. The multiplier starts near 1.45 in the first months, falls to about 0.8 at year two, and climbs steadily past year five to roughly 3.9 by year ten. Your selected fleet cohorts are plotted along it.
Worked example

One row of the default queue

600 x Toshiba MG07ACA14TA, 14 TB, 5.1 years old, of which 25 are flagged by SMART

  1. lifetime AFR 1.0 percent per year
  2. times the age multiplier at 5.1 years, 1.50 = 1.51 percent
  3. times the flagged likelihood ratio, 18.3 = 27.5 percent per year
  4. over 30 days = 2.23 percent per drive, across the 25 that are flagged
  5. 0.56 expected failures in the next 30 days
  6. each one opens a 25.9 hour rebuild across a 17+3 group holding 238 TB
  7. 133 TB drops to two parity units in the next 30 days, and that is the sort key
  8. chance any of it is actually lost, about 1 in 12 billion
What moved this row up the queue was not the AFR, which is unremarkable. The SMART flag multiplied the risk eighteen-fold, and a 238 TB group multiplied the consequence.
The drive table

Published lifetime rates

Every model the console can queue, with its published failure rate and the drive days behind it.

The full drive table
Drive models the console can queue, with published lifetime annualized failure rates, 95 percent confidence intervals and drive days observed, from Backblaze Drive Stats
ModelCapacityLifetime AFR95% intervalDrive daysNotes
The method

Every constant, in the open

This is the whole model. If you disagree with a constant, you know exactly which one to argue with.

The model9 equations
Annual ratelifetime AFR × age multiplier × evidence ratio
Evidence ratio18.26 if flagged, 0.24 if clean, normalized to the prevalence you set
P(fail in 30 days)1 − exp(−annual rate × 30 / 365)
Rebuild hoursdrive capacity TB × 10^6 / rebuild MB per s / 3600
Exposure per survivor1 − exp(−cohort rate × rebuild hours / 8760)
P(group loses data)P(at least m of the k+m−1 survivors fail inside that window)
Blast radiusk × drive capacity, the data the group is holding
TB at risk, 30 daysdrives × P(fail in 30 days) × blast radius  the sort key
TB actually lostthe same, × P(group loses data)  the fleet total
The judgement calls5 notes
  • The SMART evidence is Bayes, not a threshold. Flagged drives carry an 18x likelihood ratio, clean ones 0.24x, renormalized to the prevalence you set so total expected failures are preserved.
  • The age multiplier follows the published bathtub shape, normalized so a mixed-age fleet averages 1.0.
  • Rebuild rate is the effective rate on a live system, not the drive's spec sheet; 150 MB/s is an honest default. Unrecoverable read errors are counted only where no parity is left.
  • The queue holds drives with a reason. The bands sit at 1.5 percent a year (watch), 5 (schedule) and 15 (pull). A clean drive is usually fine even at six years.
  • A triage heuristic, not a warranty model. Figures are the published lifetime table, rounded; re-pull the current quarter before you spend money, and nothing here replaces watching your own fleet.
Built by
Adrian Mucha

Product designer for the software most designers avoid: operator consoles where a wrong click costs real money.

Focus
Data centers · AI tooling
Base
Los Angeles, CA
Status
Open to new work
Copied