You cannot pull them all. Which ones this week?
A failure rate says how many drives will die this year, not which ones. Crash Cart turns thirteen years of public fleet data into a short list: the drives worth a service trip this week, with the evidence behind each one. Every constant is in the method.
Assumptions
Four steps from a base rate to a work order
Each row of the queue is one drive population carried through the same four multiplications. You can do the arithmetic yourself.
Start from the published rate
Each model starts at its published lifetime failure rate, carrying the drive days behind it, so a thin record shows up as a wide confidence interval rather than as false certainty.
Adjust for age
Failure rates follow a bathtub: a bump in the first months, a floor near year two, a climb that gets serious after year five. Each cohort gets the multiplier for its age, not the fleet average.
Weigh the SMART evidence
In the public fleet, 77 percent of failed drives showed one of five SMART warnings, against 4 percent of healthy ones. A flagged drive is about 18 times likelier to fail; a clean one, a quarter as likely.
Weigh what a failure costs
Every failure starts a rebuild that exposes the whole group. Multiply the chance of failure by the data sitting behind it, and the queue sorts itself: risk times blast radius.
The wear curve and a worked example
The bathtub, with your fleet on it
The age multiplier applied at step 01, drawn against drive age. Each dot is one cohort in the profile you picked, sized by drive count. A fleet sitting on the right-hand climb is not a maintenance problem, it is a refresh decision.
One row of the default queue
600 x Toshiba MG07ACA14TA, 14 TB, 5.1 years old, of which 25 are flagged by SMART
- lifetime AFR 1.0 percent per year
- times the age multiplier at 5.1 years, 1.50 = 1.51 percent
- times the flagged likelihood ratio, 18.3 = 27.5 percent per year
- over 30 days = 2.23 percent per drive, across the 25 that are flagged
- 0.56 expected failures in the next 30 days
- each one opens a 25.9 hour rebuild across a 17+3 group holding 238 TB
- 133 TB drops to two parity units in the next 30 days, and that is the sort key
- chance any of it is actually lost, about 1 in 12 billion
Published lifetime rates
Every model the console can queue, with its published failure rate and the drive days behind it.
The full drive table
| Model | Capacity | Lifetime AFR | 95% interval | Drive days | Notes |
|---|
Every constant, in the open
This is the whole model. If you disagree with a constant, you know exactly which one to argue with.
- The SMART evidence is Bayes, not a threshold. Flagged drives carry an 18x likelihood ratio, clean ones 0.24x, renormalized to the prevalence you set so total expected failures are preserved.
- The age multiplier follows the published bathtub shape, normalized so a mixed-age fleet averages 1.0.
- Rebuild rate is the effective rate on a live system, not the drive's spec sheet; 150 MB/s is an honest default. Unrecoverable read errors are counted only where no parity is left.
- The queue holds drives with a reason. The bands sit at 1.5 percent a year (watch), 5 (schedule) and 15 (pull). A clean drive is usually fine even at six years.
- A triage heuristic, not a warranty model. Figures are the published lifetime table, rounded; re-pull the current quarter before you spend money, and nothing here replaces watching your own fleet.
Product designer for the software most designers avoid: operator consoles where a wrong click costs real money.
- Focus
- Data centers · AI tooling
- Base
- Los Angeles, CA
- Status
- Open to new work