On what hardware safety metrics measure, and what they do not — assembled from a long conversation between an automotive electronics engineer and an unusually patient interlocutor.
A wire, then a second wire in parallel for redundancy. The system is now safer in every physical sense — and yet the standard rates it lower. How?
Consider the simplest possible safety improvement. A wire that can break is replaced by two wires in parallel. Both must break before the connection is lost. The probability of total failure drops by orders of magnitude. The system is, on any honest physical accounting, better than what it replaced.
But ISO 26262 will not say so directly. The single wire's failure is a single-point fault: detect it, react, you have a defensible ASIL claim. The dual wire arrangement introduces a new category: the latent fault. The first wire to fail does so silently. The redundancy is quietly consumed without anyone noticing. For ASIL B and above, latent faults must be detected — and if you have not added a diagnostic to monitor each wire independently, your redundant system fails the latent fault metric while the non-redundant version did not.
Buying a part as a black box can lead to a system with higher ASIL capability and higher failure rate at the same time. The paperwork wins where the physics does not.— the opening observation
This is the paradox in its first form: a physical improvement that the standard does not credit. The asymmetry sharpens when the wire is replaced by a chip, and the chip can be either developed in-house or purchased as a Safety Element out of Context. Buy the chip, inherit the supplier's safety manual, and the system's ASIL claim is structured around the supplier's coverage declarations. Build the same chip in-house, analyze it with internal conservatism, and the claim comes out lower. The silicon is identical. The metric is not.
Two wires, two independent exponential lifetimes, no diagnostics. Sample, count, compute. The number does not agree with the standard.
The cleanest way to test a closed-form expression is to simulate the physics it claims to describe. For two identical wires with constant failure rate λ, statistically independent, over mission time T, the Monte Carlo procedure is elementary: sample t₁ and t₂ from the exponential distribution, count the trials in which both have failed before T, divide by T to obtain a rate.
The simulation, for any sufficiently large N, converges on the rate λ²T. The formula in ISO 26262-5 Annex F, applied to the same configuration, returns λ²T/2. The two are off by a clean, unambiguous factor of two — in the optimistic direction, with the standard's number being the smaller one.
This is not a question of interpretation. Both computations apply identical physical assumptions. The simulation samples the joint failure distribution; the formula is meant to express that same distribution in analytical form. They should agree. They do not.
The formula was derived for an asymmetric architecture. It is being applied to a symmetric one. The missing factor lives in that mismatch.
The PMHF dual-point latent fault expression in Annex F assumes a designated safety mechanism watching over a designated safety-related element. The mechanism fails latently; the element later fails actively; the safety goal is violated at the second failure. The integral counts trajectories of exactly that shape.
The derivation is sound for the architecture it models. A watchdog protecting a microcontroller is exactly this shape: the watchdog is the mechanism, the microcontroller is the element, one fails first and the other second, and the formula computes the right rate.
For two parallel wires, there is no mechanism and no element. There are two channels, identical in every physical respect, either of which can fail first and either of which can fail second. Both orderings cause the same physical event — connection lost — and both must be counted. The asymmetric formula counts one. The factor of two is precisely the ordering it omits.
Applies the asymmetric formula. Counts trajectories in one ordering only. Underestimates the rate of safety goal violation by a factor of two.
What the Monte Carlo measures. Counts both orderings of failure because both produce the same physical event. The rate the system actually exhibits.
This is not a modeling simplification or a conservative approximation. It is the formula computing the wrong integral for the symmetric case. Every symmetric redundancy in every ISO 26262 safety case in the field is being computed with a formula that is off by a factor of two, in the direction that makes the system look safer than it is.
There is a cleaner way to see the error. The safety goal was misformulated. Component-level instead of system-level. Per-element instead of functional.
The factor of two has a deeper cause. The standard's methodology presses the analyst to express safety goals at the component level — this wire shall not fail — when the physical safety goal is at the system level: the connection shall be maintained. The two formulations describe the same hardware. They produce different metric values.
Decomposes into per-wire obligations. Forces the asymmetric mechanism/element designation. Produces the latent fault category and the factor of two.
A property of the system as a whole. No asymmetry to designate. No latent fault category. The rate is computed symmetrically and matches the simulation.
Civil engineering does not assign a reliability target to each rivet. It assigns one to the bridge. Aerospace does not enumerate failure modes per component in isolation. It allocates an absolute residual risk target across functions. ISO 26262 part 5 inherits a component-level decomposition from IEC 61508, designed for process industries with different organizational shapes, and presses every architecture into that frame — including those whose actual safety goals are functional and emergent.
There is no latent fault category. The question of whether a fault is latent is meaningless once the safety goal is stated symmetrically.— the central reframing
The latent fault concept is genuinely useful in one specific architectural shape: a designated safety mechanism protecting a designated element, distinct from one another. Outside that shape, the concept is an artifact of forcing symmetric systems into asymmetric form. The category does work it should not be asked to do.
A component is a system if you look closely. A system is a component if you look at a distance. A correct safety metric should not depend on where the analyst chooses to stand.
Physical systems do not have a privileged level of decomposition. A transistor is a system of dopants. A logic gate is a system of transistors. A microcontroller is a system of logic blocks. An ECU is a system of microcontrollers. A vehicle is a system of ECUs. At every level, the same object is a component viewed from above and a system viewed from below. The choice of where to put the label is a modeling convention.
A correct safety metric must be invariant under this choice. The probability of harm is a property of the physical system and its environment. It does not depend on whether the analyst describes the system as one component or a hundred. This invariance is called covariance, and it is what separates a measurement from a convention.
The ISO 26262 hardware metrics are not covariant. The same artifact, described at different levels of decomposition, produces different PMHF values, different ASIL claims, different SEooC boundaries. Move the analytical seam — buy instead of make, designate channel A as the mechanism instead of channel B — and the number changes while the silicon does not.
Adjacent engineering disciplines have solved this problem. They do not enumerate failure modes on a spreadsheet. They simulate the physics.
Imagine you are about to drive across a new bridge. You are offered two pieces of evidence about its safety. The first is an FMEDA-style document: each bolt enumerated, each failure mode classified, each diagnostic coverage estimated, the lot summed into a single metric below a threshold. The second is a finite-element model of the structure under realistic load cases, validated against material tests, with a quantified margin against the loads the bridge will see.
No civil engineer would defend the first method. No regulator would accept it. The bridge analogy is not rhetorical: it is exactly what every other safety-critical discipline does, and ISO 26262 part 5 is the outlier.
Semiconductor manufacturers — the very people who supply the chips ISO 26262 is regulating — do not use FMEDA to predict whether their chips will work. They use SPICE Monte Carlo, aging simulation calibrated against JEDEC qualification, and measured failure rates against field returns. The methodology that builds the parts is physics. The methodology that certifies the systems built from those parts is paperwork. These are not the same method, and the second does not measure what the first does.
The exposure parameter is a probability. The PMHF threshold is a probability. The standard maps the first to the second through a table, without ever computing what the two together actually mean.
ASIL is determined from severity, exposure, and controllability. These are sound risk parameters. The product of their underlying quantities is the unmitigated risk; the logarithmic sum S+E+C is a reasonable index of that product. So far the construction is coherent.
The recursion appears in the next step. The ASIL is then mapped — by table lookup — to a fixed allowed failure rate: 10⁻⁸/h for ASIL D, 10⁻⁷/h for ASIL C. The probability of the dangerous situation went into the analysis once, in E. The probability of system failure comes out the other end, in the PMHF threshold. The two are coupled through a table, but they are never combined into a residual risk and compared against an absolute target.
Aerospace closes this loop. Severity sets an absolute target — 10⁻⁹ catastrophic events per flight hour — and the allowed system failure rate is derived from that target after accounting for exposure and controllability. The probabilities multiply, the product is bounded, residual risk is a calculated quantity.
What ISO 26262 should do, in the same logic, is associate each severity class directly with an absolute harm-rate budget:
| Class | Consequence | Allowed harm rate per operating hour |
|---|---|---|
| S3 | Life-threatening or fatal injuries | 10⁻⁹ |
| S2 | Severe injuries, survival probable | 10⁻⁷ |
| S1 | Light or moderate injuries | 10⁻⁵ |
| S0 | No injuries | — (no quantitative target) |
The allowed system failure rate would then be computed: severity budget divided by the product of exposure and controllability probabilities. The result is a continuous quantity, derived from the physics, and the residual risk it produces is explicitly bounded. The recursion disappears because the loop is closed.
The defects are not independent. They are four faces of the same underlying property. Once that property is named, the rest of the critique writes itself.
Across the conversation, four structural problems with the part 5 methodology surfaced. Stated together, they form a pattern:
The PMHF dual-point latent formula is derived for asymmetric architectures and applied to symmetric ones without correction. For any symmetric redundancy, the formula underestimates the rate of safety goal violation by a factor of two, in the optimistic direction.
The same physical hardware produces different metric values depending on where the component-system boundary is placed and how the safety architecture is described. SEooC inheritance, latent fault classification, and the buy-versus-make asymmetry are all consequences of this single defect.
Probability enters the analysis in the exposure parameter and exits in the PMHF threshold, but the two are never combined into a residual risk and compared against an absolute target. The categorical mapping replaces what should be a calculation.
Every other safety-critical discipline measures physics. ISO 26262 part 5 measures the structured output of an analytical procedure applied to imported numbers. The procedure is valuable; the claim that its output measures hardware safety integrity is not supported by the procedure itself.
These are not four separate observations. They are one observation seen from four angles. The methodology operates on categorized structural inputs and prescribed outputs, with fixed mappings between them. It does not operate on calculated physical quantities. Wherever the categorical structure matches the physical structure of the system being analyzed, the method produces reasonable results. Wherever the categorical structure is forced — symmetric redundancies, distributed functions, novel architectures, autonomous systems — the method produces artifacts that the standard cannot self-correct.
The auditing method does not measure safety, which is what the auditing method is supposed to measure.— the load-bearing claim
None of this invalidates what the standard does well. ISO 26262 raises the floor of automotive safety practice. It enforces process discipline. It provides cross-organizational vocabulary. It gives legal defensibility to manufacturers who follow it. These contributions are real and they are valuable. They are also separable from the claim that the part 5 metrics measure physical hardware safety integrity.
The honest position is to keep what works and to name what does not. Part 5 is the part that does not. Its metrics are useful as conventional values produced by an agreed procedure — useful for coordination, useful for certification, useful for legal cover — and they are not measurements of the physical quantities they nominally describe. The factor of two in the symmetric case is the cleanest demonstration of this, because it is verifiable in fifty lines of simulation code, and because it shows the gap without any rhetorical assistance. The Monte Carlo is right. The standard is, in that specific case, wrong. The rest of the critique extends from there.
This document was assembled from a conversation about an industry that has not yet been willing to ask the question out loud.