FMEA RPN and Weibull reliability: ranking risk and reading failure data
How to score failure modes with severity, occurrence and detection, why RPN has limits, and how Weibull shape and scale describe failure behaviour over time.
FMEA and the risk priority number in plain terms
Failure mode and effects analysis (FMEA) lists how a component or process can fail, what happens when it does, and what stops or reveals it. Each failure mode is scored on three ratings from 1 to 10: severity (how bad the effect is), occurrence (how likely the cause is) and detection (how likely current controls are to catch it before the effect, where 10 means hard to detect). The FMEA RPN calculator multiplies them:
- RPN = Severity × Occurrence × Detection
The maximum is 10 × 10 × 10 = 1000. The scales are defined by your own team, and the same scale must be used for every failure mode in one study.
Worked example: ranking three failure modes
A pump team scores four failure modes:
| Failure mode | S | O | D | RPN |
|---|---|---|---|---|
| Seal leak from dry running | 8 | 4 | 5 | 160 |
| Motor winding failure | 9 | 2 | 2 | 36 |
| Coupling wear | 4 | 6 | 7 | 168 |
| Impeller erosion | 5 | 4 | 8 | 160 |
Ranked by RPN, coupling wear (168) is first, with the seal leak and impeller erosion tied at 160. The winding failure sits last at 36, yet its severity is 9. The calculator flags severity of 9 or higher for review whatever the RPN.
After an action is complete, re-rate. Suppose a dry-run protection device lowers the seal leak detection score from 5 to 2: 8 × 4 × 2 = 64. Re-rating confirms the action changed the risk. It does not prove it; the new scores still rest on judgement.
Limits of RPN
- Different combinations give the same number. Severity 8, occurrence 4, detection 5 and severity 5, occurrence 4, detection 8 are both 160, but they are very different risks.
- The ratings are ordinal ranks, not measurements. Multiplying them can reverse a ranking that a sensible risk view would not.
- A severe failure with rare occurrence and good detection can score low and be ignored.
- There is no universal action threshold. The team sets one.
- Some organisations now use action priority tables that weigh severity first. The tool uses the simple product, so treat it as a ranking aid.
Weibull: shape, scale and the bathtub curve
FMEA tells you what could fail. Weibull analysis describes how failure builds with time, using two parameters. The shape β says how the failure rate changes with age. The scale η (characteristic life) is the age by which 63.2 percent of units have failed. The Weibull reliability calculator takes both parameters and an operating time t.
- β below 1: failure rate falls with time, suggesting early-life failures such as installation or commissioning faults.
- β about 1: constant failure rate, random failures. Age-based replacement gives little benefit.
- β above 1: failure rate rises with time, indicating wear-out.
These regions correspond to the three phases of the bathtub curve: early failures, useful life and wear-out. A real item may show different phases at different ages, and one Weibull fit describes only one of them.
Formulas: R(t) = exp(−(t ÷ η)β); F(t) = 1 − R(t); h(t) = (β ÷ η) × (t ÷ η)β−1; MTTF = η × Γ(1 + 1 ÷ β); B10 life = η × (−ln 0.9)1 ÷ β.
Worked examples: wear-out and early-life behaviour
Example A (wear-out). β = 2.5, η = 8,000 h, t = 4,000 h.
- (t ÷ η)β = 0.52.5 = 0.1768, so R = exp(−0.1768) = 83.7967 %
- F = 16.2033 %
- h = (2.5 ÷ 8,000) × 0.51.5 = 1.1049 × 10−4 per hour
- MTTF = 8,000 × Γ(1.4) = 7,098.1105 h
- B10 = 8,000 × (−ln 0.9)0.4 = 3,252.0794 h
Example B (early life). β = 0.7, η = 5,000 h, t = 1,000 h gives R = 72.3155 %, F = 27.6845 %, h = 2.2689 × 10−4 per hour, MTTF = 6,329.1175 h and B10 = 200.8147 h. Hazard falls with age: 2.7934 × 10−4 at 500 h, 2.2689 × 10−4 at 1,000 h and 1.8429 × 10−4 at 2,000 h.
At t = η the reliability is exp(−1) = 36.7879 % for any β. The calculator shows hazard rates with four decimals, so these small values may display as 0.0001 or 0.0002; use the scientific figures above to compare. The tool takes parameters as given and does not fit data. For observed MTBF from a repairable item, the MTBF calculator is the simpler route, valid only when the failure rate is roughly constant.
Common mistakes
- Using FMEA ratings from different scales or teams in one ranking.
- Scoring detection the wrong way round. Here 10 means hard to detect.
- Acting only on high RPN and ignoring severity 9 or 10.
- Treating RPN differences of a few points as meaningful.
- Entering η in the wrong time unit, such as calendar days with t in running hours.
- Entering a guessed β with no data behind it, then applying the result to a replacement decision.
- Fitting data that mixes different failure modes, which can hide the real shape.
- Leaving out suspended units (items still running) when fitting parameters. That biases the fit.
Checks before you trust the result
- Verify that FMEA ratings follow written scale definitions and were agreed by people who know the equipment.
- Rank by severity as well as RPN, and re-rate after actions.
- Make sure β and η come from a Weibull analysis of your own failure data, with time measured the same way as t.
- Remember that a small sample gives wide uncertainty on both parameters.
- Compare predictions with what has happened. If the B10 life is far from actual early failures, the fit or failure mode grouping needs another look.
All outputs are estimates and are not substitutes for supplier data or site procedures.
Frequently asked questions
What RPN is too high?
No universal limit exists. Your team agrees a threshold, and any failure with severity 9 or 10 should be reviewed regardless.
What does η mean in practice?
It is the age at which about 63.2 percent of units have failed, whatever the shape.
What does B10 life tell me?
It is the time by which 10 percent of units are expected to have failed. It is useful for planning inspections or replacement before wear-out begins when β is above 1.
Can I get β from the MTBF?
No. MTBF is one average. Estimating β needs the individual failure ages, plus the ages of units that have not failed.
Recording data for FMEA and Weibull
Both methods depend on records. For every failure, write down the asset, install date, failure date, running hours at failure, failure mode and whether the unit was repaired or replaced. Also record units removed while still working, since their ages matter to a Weibull fit. For actions arising from an FMEA or an RCA, use the RCA action log template, which treats an action as closed only after verification. Downtime and repair hours can be summarised as described in the guide on MTBF, MTTR and downtime.