MTBF - Mean time between failures
MTBF explained: what mean time between failures really measures
Mean time between failures (MTBF) is the average operating time a repairable asset delivers between one failure and the next, found by dividing total uptime by the number of failures in that period. It is the standard measure of reliability, and it applies only to assets you repair and put back into service.
For field service
Know what the job needs before the van rolls
Venta Capture, a product of VentaVid, lets the customer show you the fault first, so the engineer arrives with the right part or does not need to arrive at all.
You will see it written as MTBF, mean time between failure and mean time between failures. Same metric. What is not the same metric is MTTF, and that mix-up is covered below.
How is MTBF calculated?
The formula is MTBF = total operating time divided by the number of failures in the same window. Two rules keep it honest. Operating time means uptime, so repair hours come out of the numerator. And a failure means a loss of the required function, not any stop at all, so a changeover or a planned shutdown does not count against you.
MTBF explained: a worked example
A compressor is scheduled to run continuously through a 30 day month, so 720 scheduled hours. It fails four times and total downtime for those failures is 20 hours, leaving 700 operating hours. MTBF is 700 divided by 4, so 175 hours. The same four events give a mean time to repair of 20 divided by 4, so 5 hours.
Put the pair together and availability is MTBF divided by (MTBF + MTTR), which here is 175 divided by 180, or 97.2 percent. That is roughly 20 hours of lost production a month on one machine, which is a very different conversation from "it broke four times".
Why MTBF is not life expectancy
This is the most expensive misreading of the metric. An MTBF of 175 hours does not mean the compressor dies after 175 hours. It means that, across the period measured, it delivered an average of 175 hours of work between repairable events.
The same trap sits inside vendor datasheets. A component rated at 100,000 hours MTBF is not rated for eleven years of service life. That figure describes the failure rate during the flat useful-life stretch of the curve, before wear-out begins, and it is usually derived from a large population over a short test rather than from one unit running for a decade. As the US Department of Energy's Operations and Maintenance Best Practices Guide describes it, the wear-out period is defined by a rapidly rising failure rate, and MTBF says nothing about when that period starts.
MTBF, MTTF and MTTR: which one applies
- MTBF covers repairable assets. A pump, a conveyor, a chiller, a compressor.
- MTTF, mean time to failure, covers items you replace rather than repair. A bearing, a lamp, a sealed sensor. There is no "between" because there is only one failure per item.
- MTTR, mean time to repair, measures maintainability rather than reliability: how long the fix takes once the failure has happened.
Reliability and maintainability are separate levers and they trade against each other. An asset with poor MTBF and excellent MTTR can be more available than an asset that fails rarely and takes a fortnight to source parts for.
What moves MTBF
Changing MTBF means changing failure causes, which is slower work than shaving repair time. The lever with the strongest published evidence behind it is condition monitoring. Independent surveys reported in the DOE guide put the industrial average effect of a functioning predictive maintenance programme at a 70 to 75 percent elimination of breakdowns, a 35 to 45 percent reduction in downtime and a 25 to 30 percent cut in maintenance costs.
Scheduled work matters too, though the DOE figures are more modest: preventive maintenance is estimated to save 12 to 18 percent over running assets to failure. Beyond that, MTBF responds to root cause analysis, correct installation and alignment, operating within design conditions, and parts quality.
Where MTBF misleads
- Small failure counts. Four failures is not a sample. Month-to-month swings in an MTBF built on a handful of events are mostly noise.
- Averaging a fleet. One chronically bad unit can be masked by nineteen good ones. Calculate per asset first, then roll up.
- Counting every stop as a failure. Blocked or starved time, changeovers and operator stoppages are availability losses, not reliability losses. Mixing them in makes MTBF a production metric by accident.
- Assuming a constant failure rate. MTBF is a fair summary during useful life. Through infant mortality and wear-out, the average hides a trend that matters more than the number.
Reliability engineers get the most out of MTBF when it is paired with the metrics that describe what happens after the failure, particularly first-time fix rate and the avoidable truck roll. A high MTBF with a poor first-time fix rate still means long outages. Better remote diagnostics before dispatch is often what closes the gap between the two.