Fooled by Best Practice, Part IV: The Barbell, Physics at One End and Simulation at the Other
J. McKenney
Paper 4 of five in Fooled by Best Practice. Part I made the survivorship argument in ordinary language. Part II separated what an operator can observe from the distribution those observations were drawn from. Part III gave the method for sampling the branches that did not run. This part is the engineering consequence, and it is the most technical of the five. Part V is the allocation decision.
Licence: CC BY 4.0. 16 September 2026.
| Field | Value |
|---|---|
| Document ID | WG-02-DT-FBP-4 |
| Slug | fooled-by-best-practice-4-barbell |
| Working group | WG-02-DT, Digital Twin |
| Series | Fooled by Best Practice, part 4 of 5 published |
| Author | J. McKenney |
| Published | 16 September 2026 |
| Revision | |
| Status | Published |
| Preceded by | WG-02-DT-FBP-1, WG-02-DT-FBP-2, WG-02-DT-FBP-3 |
| Followed by | WG-02-DT-FBP-5 |
Executive Abstract#
For twenty years the corporate security industry has sold office tooling into industrial plants: endpoint agents, scanners, firewalls and cloud log forwarders, all pitched to protect a substation like a bank's laptops. A substation is not a laptop. An office computer holds information; a control system holds a physical process, and when it stops a turbine comes apart. Information restores from backup; a compressor does not.
Four mechanisms do the damage. An agent that steals a control processor for fifteen milliseconds trips a relay running a ten-millisecond loop. Scan probes lock up equipment designed in the 1990s, costing the operator view and control. Every agent is privileged software; when one fails, as it did worldwide in July 2024, every machine running it fails at once. A firewall can check a command is well formed, not whether obeying it destroys the equipment, a question for thermodynamics, not protocol.
The response is Taleb's barbell from Antifragile, 2012, after The Black Swan, 2007: weight the two ends, empty the middle. One end holds protection with no software at all: a rupture disc, a mechanical governor, a bimetallic switch wired straight to a breaker trip coil, a one-way optical link. Nothing to patch, no network to reach. The other holds all the computation you like, including the Part III simulation, on a supervisory system that reads and never writes; if it crashes, the plant runs on. What leaves is the middle: software able to hurt the process and too complex to fail predictably.
Abstract#
This part makes the engineering case against transplanting enterprise security tooling into operational technology. Four mechanisms of harm are set out: preemption of hard real-time scan cycles by user-space monitoring; connection-state exhaustion in legacy field devices under active scanning; enlarged attack surface and correlated failure from privileged agents and telemetry; and syntactic protocol validation standing in for a physical feasibility check no packet inspector can perform. They follow from position, not implementation, and fall on the unobserved side of Part II's table. In their place, Taleb's barbell, translated to plant architecture: one end deterministic physical grounding, mechanical or analogue layers with no firmware, operating system or network interface; the other a read-only supervisory plane across a unidirectional boundary, with no write path to any actuator; the middle evacuated. The case is stated in the terms of IEC 61508 and IEC 61511: an interlock is a protective layer whose claim rests on independence; software failure is systematic, not random; a shared software component defeats layer independence as a common cause. Allocation goes to Part V.
1. Introduction#
1.1 Where the series has got to#
Part I argued that an incident-free record is a single sample path and not a measurement of defensive quality. Part II separated the observable side of the table from the distribution behind it and showed where the audited drawing and the operating plant come apart. Part III set out what it takes to sample the branches that did not run this time, and what a model has to represent before that sampling means anything.
All three of those parts are about knowing. This one is about building.
If the argument so far is right, then the risk that matters lives in states the plant has not visited, and any component whose failure modes you cannot enumerate is contributing to that risk in a way no audit will show you. That is an architectural statement, and it has consequences that I can be concrete about, because the consequences are the ones I spend my working weeks arguing over on sites.
1.2 The claim I am making#
I am going to argue that the class of security product most commonly sold into operational technology is, as a class and not as a matter of individual product quality, in the wrong place.
I want to be careful about what that does and does not mean. I am not saying the vendors are dishonest, and I am not saying the products are badly made. Several of them are excellent at what they were built for. I am saying that the assumptions under which they were built do not hold in a plant, that the resulting failure modes are physical rather than informational, and that those failure modes are invisible to the compliance instrument that motivated the purchase.
I am also not arguing for less software. The control system is software. The simulation work in Part III is software, and a great deal of it. The argument is about position: what a piece of software is allowed to write to, and what happens to the process when that software is wrong.
1.3 What this part is not about#
Two subjects sit adjacent to this one and belong elsewhere. The first is how much any of this costs and how an operator should divide capital between the two ends of the barbell. That is Part V. The second is the modeling requirement for the supervisory twin, which Part III has already set out; I refer to it here and do not re-derive it.
2. Two different machines#
2.1 What enterprise tooling assumes#
Enterprise security tooling encodes a set of assumptions so consistently that they are rarely written down.
Confidentiality is the first-order concern, and availability is a secondary one that can be traded against it. Compute is abundant, so an agent that takes a few percent of a processor is free. Latency is soft, so a delay of tens or hundreds of milliseconds is beneath notice. Reboots are acceptable, and a machine that has to be restarted to apply a patch can be restarted tonight. The population of endpoints is large, homogeneous and replaceable, so per-host reliability matters less than fleet coverage. Update cadence is fast by design, because signatures and detection content must be current to be useful.
Every one of those assumptions is reasonable for a fleet of office computers. Every one of them is false on the plant floor.
2.2 What a plant is#
In operational technology the priority order is inverted, and not as a matter of preference. Safety comes first, then availability, then integrity of the process data, and confidentiality last. This is not because operators do not care about secrecy. It is because a setpoint is not a secret worth much, and a stopped process can hurt someone.
Compute is fixed at design time and often fixed for the asset's life, which may be twenty or thirty years. A protection relay was specified with the cycles it needs and no surplus. Latency is hard, in the formal sense: a control loop that must complete inside its period does not degrade gracefully when it misses, it fails. Reboots are not available on demand, because restarting a controller means interrupting the process it controls, and the process may be an electrolyser, a furnace, or a section of railway with trains on it. The population of devices is small, heterogeneous, and in many cases no longer manufactured. And update cadence is slow on purpose, because a change to a device that governs a physical process must be validated against that process before it is allowed anywhere near it.
The two environments are not points on a spectrum. They are answers to different problems.
2.3 Why the tooling lands where it does#
The Purdue reference model, in the form most operators use, separates the enterprise levels from the control levels with a demilitarized zone in between. It is a useful model and I use it. What it also does, in practice, is create a natural sales gradient: a vendor who has sold into levels 4 and 5 is already inside the organization, already trusted by the people who hold the security budget, and already able to point at the lower levels and observe correctly that they are unmonitored.
The next step is the mistake. The tooling moves down, level by level, because each individual move looks small. An agent on the engineering workstation. Then on the historian. Then on the human-machine interface, because the HMI is a Windows box and looks like the others. Then a scan of the control network, because the asset inventory came back incomplete. Then a collector that forwards logs upward through the DMZ, because the security operations centre needs the events.
None of those steps is where the argument is won or lost. The argument is about where the tooling ends up, which is in the region between the physical process and the people watching it: close enough to the actuator to affect it, complex enough that nobody can say precisely how it will fail.
That region is the fragile middle, and the rest of this paper is about getting out of it.
3. Four mechanisms of software-induced equipment damage#
The commercial security industry rests on software complexity. In critical infrastructure, software complexity produces four concrete physical failure mechanisms. I take them in order of how often I see them.
3.1 Microsecond jitter and real-time loop preemption#
Programmable logic controllers and remote terminal units run hard real-time operating systems. VxWorks, QNX and FreeRTOS are the ones I meet most. The primary task of such a device is the scan cycle: read the inputs, solve the control equations, write the outputs, repeat. Typical cycle periods in the plant I work on run from about one millisecond to about ten. Motion control and high-speed protection sit at the fast end of that range and sometimes below it.
The word that matters is deterministic. The requirement is not that the cycle completes quickly on average. It is that it completes inside its period every time, with bounded jitter, because the control law and the protection settings were tuned on the assumption that it would.
Now introduce a user-space process on the same host or on a host sharing the same fieldbus segment: an endpoint agent, a logging daemon, an inventory collector. It does not have to be badly written. It competes for the memory bus, for cache lines, for interrupt service time, and on a shared-memory architecture it competes for the very resources whose contention behavior the real-time scheduling analysis assumed away. A scan of the file system, a signature database update, a burst of compression before a log upload: each is a period of sustained contention.
A scheduling delay of fifteen milliseconds inserted into a synchronous loop that was specified at ten is not a performance degradation. It is a missed cycle. Downstream of a missed cycle, a high-speed compressor protection relay sees a stale or absent value where it expected a fresh one, and it does what it was designed to do, which is trip. The facility goes down, uncommanded, for reasons that will be logged nowhere near the agent that caused them.
I have watched the subsequent investigation more than once. The event record shows a protection trip. The protection engineer confirms the relay behaved correctly. The control engineer confirms the setpoints were right. Nobody looks at the security agent, because the security agent is not on the process side of the organization and its logs are held by a different team in a different system with a different retention period.
Two further points about this mechanism, because it is the one most often dismissed.
First, it is intermittent by nature, which is the worst property a fault can have. Contention depends on what else the host is doing, so the failure appears under load, or after an update, or on the day the scheduled scan coincides with a process transient. It does not reproduce on the bench.
Second, and this is the part that connects to Part II, the site that has not yet seen this failure has not demonstrated that it is safe from it. It has demonstrated that the coincidence has not occurred yet. The agent has been installed for eighteen months and the plant has not tripped, and that sentence has exactly the epistemic weight Part I gave it, which is very little.
3.2 Active scanning against fragile legacy stacks#
An active vulnerability scanner discovers what is on a network by sending traffic at it: rapid TCP SYN sequences, malformed and truncated UDP datagrams, SNMP sweeps, version probes, and protocol-specific interrogations designed to elicit a distinguishing response.
Against a modern server this is unremarkable. Against industrial equipment it is an attack, and not a metaphorical one.
Consider what is actually on a control network. Ethernet-to-serial converters whose job is to carry Modbus RTU frames over a network they were retrofitted onto. Protocol gateways translating between fieldbus families. Serial bridges, terminal servers, older safety logic solvers, and drives with network cards added as an option module. A great deal of this equipment was designed in the 1990s or early 2000s by engineers whose network design target was a handful of well-behaved polling masters on an isolated segment.
The consequence is that these devices have small, primitive protocol stacks. Connection state tables sized for four or eight concurrent sessions. No SYN backlog management. No rate limiting. Parsers that assume well-formed input because on the network they were designed for, input was well formed. Watchdogs that cover the application task but not the network stack.
A standard scan exhausts the connection table in seconds. The stack stops accepting new sessions, and in the cases I have seen most often the existing polling session is dropped along with everything else. Some devices recover when the scan stops. Some require a power cycle. Some, where the fault reaches the shared microprocessor rather than being contained in the network task, stop doing their control job as well.
The operational term for the result is precise and it is worth using: loss of view, and loss of control. Loss of view means the control room no longer knows the state of the plant. Loss of control means the control room can no longer change it. A scan that produces both has converted a routine security activity into a denial of service against a physical process, executed by the operator against itself, during a maintenance window, with a change ticket approved.
There is a standard mitigation, which is to scan only during planned outages, or to use passive discovery instead. Passive discovery is the right answer and I recommend it constantly. But notice what has happened to the argument at that point: the tool has been retained and its defining capability has been switched off. What remains is an inventory function that a span port and a protocol decoder would have provided without the risk.
3.3 Expanded attack surface and kernel privilege#
Every agent added to an industrial host is new software that did not previously exist on that host, and it arrives with three properties that matter.
It runs with high privilege, usually including kernel-mode components, because that is what endpoint detection requires in order to see what it is meant to see. It opens network sockets, inbound for management and outbound for telemetry, on a host that may previously have spoken only its control protocol. And it updates frequently, often daily, by pulling content from outside the facility.
Take those in turn.
Privilege means that a defect in the agent is a defect in the machine. A fault in a kernel-mode driver does not produce a failed process that can be restarted. It produces a stopped host. In July 2024 a content update to a widely deployed endpoint sensor did precisely this to a very large number of Windows machines worldwide at effectively the same moment, and the reason the event was so instructive is not that the vendor made a mistake. Vendors make mistakes. It is that the failure was perfectly correlated. Every host running that agent failed for the same reason at the same time, and the redundancy that the affected organizations had designed into their systems did nothing, because redundancy is defence against independent failure and this failure was not independent.
I will return to that point in section 6, because it is the whole argument in miniature.
Sockets mean attack surface. An agent that listens for management connections has added a listening service to a control host. The service is authenticated, no doubt competently. It is still a path that did not exist before, on a machine that was chosen for its role partly because it did so little.
Outbound telemetry means something worse, and it is the one I argue about most. A cloud-tethered log forwarder sitting at Purdue level 2 or 3 needs to reach a collector outside the facility. To make that work, somebody opens an outbound path through the demilitarized zone. It is usually presented as safe on the grounds that it is outbound only and encrypted.
The DMZ was drawn to enforce a direction. Once a component inside the control zone can initiate a session outward and maintain it, the direction is no longer enforced by the architecture; it is enforced by the correctness of that component's configuration and code. An established outbound session is a bidirectional channel by construction, whatever the intent of the party who opened it. And the encryption that makes the path acceptable to the security team is the same encryption that makes its contents unavailable to the inspection the DMZ was supposed to perform.
I have stood in front of a network drawing that showed a clean separation and been told, in the same conversation, about the log forwarder that crosses it. Both statements were made sincerely. This is the reference architecture illusion from Part I, in its most consequential form.
3.4 Syntactic validation standing in for thermodynamic reality#
The fourth mechanism is different from the first three. The first three are ways the tooling breaks the plant. This one is a way the tooling fails to protect it while appearing to.
A next-generation firewall with industrial protocol awareness inspects the syntax and semantics of a transaction. For a Modbus TCP write it can check that the function code is one it permits, that the register address falls inside an allowed range, that the value is within the declared data type, that the source address belongs to an authorized engineering station, and that the session is properly formed. Some products go further and enforce a learned baseline of normal command patterns.
All of that is real work and I do not dismiss it. Here is what it cannot do.
Consider a single write: function code 0x06, write single register, to the register that holds the position command for a cooling water valve, value zero. Every check passes. The function code is permitted. The register is inside the allowed range. The value is a legal position. The source is the engineering workstation, which is authorized. The session is well formed. The firewall permits it, correctly, according to every rule it holds.
Downstream, the cooling water valve closes on an exothermic reaction that is running. The reaction now has no heat removal. What happens next is determined by the reaction enthalpy, the vessel's thermal mass, the time constant of the temperature rise and the relief capacity, and none of those quantities exists anywhere in the firewall.
The general statement is this. A firewall validates conformance to a protocol specification. Harm is produced by conformance to the laws of thermodynamics. Those are different specifications, and no amount of deeper packet inspection closes the gap, because the information required is not in the packet. It is in the state of the process, the physical configuration of the plant, and the constraint set that governs them.
There is a class of device that does address this, and it is worth naming so that the argument is not read as broader than it is: a unidirectional protocol break with a process-aware constraint check, or a control action validated against a physical model before it is allowed. That is a legitimate engineering approach, and the model it needs is the one Part III describes. But notice that the moment the check becomes physical rather than syntactic, the component stops being a firewall and becomes a piece of control engineering, with all the validation burden that implies. Which is the honest position, and it is nothing like what is being bought.
4. Why these are not implementation defects#
The obvious reply to section 3 is that each mechanism has a mitigation. Pin the agent to an isolated core. Use passive discovery. Do not deploy kernel drivers on level 1 and 2 hosts. Put a physical diode in the telemetry path. Add process-aware validation to the firewall.
Every one of those is correct advice, and I give it. But taken together they amount to an admission, and the admission is the point of this section.
4.1 The mitigations remove the capability#
Work through what is left of an endpoint agent after you have denied it the processor share it needs to inspect the file system, after you have refused it kernel-mode visibility, after you have cut its outbound update path so its detection content is frozen at the version you validated, and after you have accepted that you cannot restart the host to apply a fix. You have a process that consumes memory and produces a heartbeat.
Work through what is left of a scanner once it may only run during outages and may not send malformed traffic. You have an inventory tool that runs twice a year, which is a slower and riskier way to obtain what a passive tap gives you continuously.
The mitigations are not adjustments to the products. They are subtractions of the properties that made the products worth buying. When the residue is a compliance artifact that satisfies a control requirement and does no engineering work, the honest conclusion is that the control requirement was satisfied by something else and the product should not be there.
4.2 The failure modes land on the unobserved side#
This is the connection to Part II and it is the reason this part exists inside this series rather than as a standalone complaint about tooling.
Every failure mode in section 3 shares a structure. It is rare, because it needs a coincidence: contention during a transient, a scan against the one device with a four-entry connection table, an update that is wrong on the day, a legal command at the wrong moment in the process. It is severe when it happens, because the consequence is physical. And it is unobserved until it happens, because there is no routine that exercises it.
A rare severe unobserved failure mode is exactly the object Part II located on the left side of the table. It contributes nothing to the audit, nothing to the dashboard, nothing to the incident count. It contributes only to the distribution.
So the tooling's effect on measured risk and its effect on actual risk point in different directions. The measured risk falls, because a control requirement that was open is now closed. The actual risk includes a new term that the measurement does not carry. Whether the net is positive is an empirical question, and the striking thing is that almost nobody asks it. The reduction is assumed and the addition is not counted at all.
4.3 The real problem is unenumerable failure modes#
Underneath all of it is a property I can state plainly. A modern endpoint agent, with its drivers, its update mechanism, its content interpreter and its dependency tree, has a failure mode space nobody can enumerate. Not the operator, not the integrator, and not the vendor, who can enumerate the ones found so far.
For a component in an office that is tolerable, because the consequence of an unenumerated failure is bounded by what information systems can do, which is stop. For a component that can influence an actuator, it is not tolerable, because the consequence is bounded by what the process can do, and in the plants I work in the process can do a great deal.
That is the criterion the barbell is built around, and it is the only criterion it needs. Either a component's failure modes can be enumerated and bounded, or the component must be positioned where its failure cannot reach the physics.
5. The barbell#
5.1 Where it comes from, stated correctly#
The barbell is Taleb's, and the attribution matters because it is routinely got wrong, including by people quoting it approvingly. It is developed at length in Antifragile, published in 2012, where it is the central practical prescription of the book [1]. It appears earlier, in less developed form, in The Black Swan of 2007 [2]. It is not from Fooled by Randomness [3], which is the book that opened this series and which is about a different problem, namely the misreading of track records.
The financial version is simple to state. Rather than hold a portfolio of medium-risk assets, hold a very large fraction of capital in instruments whose downside you are confident about, and a small fraction in positions with large asymmetric upside, and hold nothing in the middle. Taleb's argument for it is not about expected return. It is about the reliability of the estimate. Medium-risk instruments are the ones whose risk is most confidently modeled and most often modeled wrong, and their failures arrive in the tail, correlated, at the moment the model says they should not.
The general form, stripped of finance, is this. When your estimate of a risk is itself unreliable, restructure the exposure so that the unreliable estimate stops mattering: make the downside bounded by something you do not have to estimate, and confine the things you cannot bound to a place where being wrong is survivable.
5.2 Translating it into architecture#
The translation to plant design is not an analogy. It is the same operation applied to a different exposure.
Ask of every component in a facility two questions. What is the worst physical consequence of this component being wrong, whether through fault, defect or adversary action? And how confident am I in my enumeration of the ways it can be wrong?
Those two questions divide the plant into three regions.
Components whose failure modes are few, physically determined and enumerable, and whose worst consequence is bounded by construction. A rupture disc can fail to burst or burst early. That is the list. It is a short list because the device is made of metal and geometry.
Components whose failure modes are many and not enumerable, but whose worst consequence is bounded by position, because they hold no authority over anything physical. A supervisory simulation that reads and never writes can be wrong in an unlimited number of ways, and the worst it can do is tell a human engineer something false, which a human engineer is positioned to reject.
And the middle: components whose failure modes are not enumerable and whose worst consequence is not bounded. Anything with write authority over a process and a software surface too large to characterize. That region is where the enterprise tooling of section 3 sits, and it is also where a great deal of ordinary control system practice sits once the connectivity has accumulated.
The barbell says to weight the first two and evacuate the third. Nothing remains in the middle that can fail in a way the physics cannot absorb.
5.3 End one: deterministic physical grounding#
At the actuator boundary, protection is achieved by physical law and by nothing else.
Mechanical protection is the first category. A rupture disc specified to burst at a stated pressure, say eighteen bar, and to relieve into a containment path sized for the flow. A mechanical overspeed governor that closes the fuel or steam valve when a flyweight moves against a spring at a set rotational speed. A spring-loaded pressure relief valve. A counterweight damper that falls closed when the holding force is lost. A mechanical interlock on a switchgear truck that physically prevents racking a breaker in under load.
Analogue electrical protection is the second. A bimetallic thermal switch that opens a contact when the metal reaches its set temperature, say eighty-five degrees Celsius, wired directly into a breaker trip coil with no logic between the contact and the coil. A float switch on a hardwired permissive. A pressure switch in series with a solenoid. A thermocouple into a discrete trip amplifier whose output is a relay contact.
The one-way boundary is the third, and it is what makes end two possible. A hardware optical data diode is a transmitter with an emitter and no receiver, facing a receiver with a detector and no emitter. Information leaves the control zone as light and there is no physical path for it to return, because the reverse component does not exist. This is a different claim from a firewall rule that permits traffic in one direction. A rule is a configuration, and configurations can be changed, misapplied or subverted. The absence of a receiver is a fact about the hardware.
What these devices have in common is the property that matters. They contain no firmware, so there is no code to be wrong. They run no operating system, so there is no scheduler to be preempted and no memory to leak. They have no network interface, so there is no path by which an adversary who is present on the network can reach them at all. A published vulnerability in a protocol stack is not relevant to a bimetallic strip. A supply chain compromise in a software dependency does not reach a spring.
They are also, and I want to be honest about this rather than let it be discovered as a gotcha, limited. They are slow to change, expensive to retrofit, incapable of nuance, and they fail in their own ways. Section 7 deals with that.
5.4 End two#
supervisory computation that does not actuate
All the analytical complexity goes to the other end, and it is allowed to be as complex as it needs to be, because it has been placed where complexity is affordable.
The supervisory plane ingests telemetry across the unidirectional boundary. It builds and maintains the model of the facility that Part III specifies: the cyber and physical layers together, the dependency structure, the process constraints. It samples counterfactual histories rather than reporting the one that happened, which is the whole method of Part III and the reason that part exists. It runs graph reasoning over the model to find paths that no human enumeration would reach. It produces advice: this dependency is carrying more of the facility's risk than anyone thinks, this configuration change closes more paths than that one, this is the distribution of outcomes under the states you have not visited.
It delivers that advice to human engineers, who decide what to do, and who then make changes through the ordinary engineering change process with its ordinary validation.
The property that defines this end is negative and it is the important one. There is no write path from the supervisory plane to any actuator. Not a restricted one, not an authenticated one, not one behind an approval workflow. None, enforced by the hardware of the boundary rather than by the configuration of a rule.
Two consequences follow immediately. If the supervisory software crashes, hangs, corrupts its own database or produces nonsense, the plant continues operating exactly as before, because the plant was never depending on it. And if an adversary fully compromises the supervisory environment, including the cloud infrastructure it runs on, they have obtained a model of the facility, which is a real loss and I do not minimize it, but they have obtained no ability to move anything. Between them and the physics sits an absent receiver and a set of protective layers with no network interface.
That asymmetry is what the barbell buys. The complexity is not reduced; it is relocated to where being wrong is survivable.
5.5 What actually leaves the middle#
To be concrete about the evacuation, because otherwise this reads as a diagram rather than a plan.
Endpoint agents come off control hosts at the lower Purdue levels. Active scanners stop running against production control networks, and passive discovery from a span port or a network tap replaces them. Cloud-tethered collectors inside the control zone are replaced by collection on the far side of the diode. Deep packet inspection appliances that infer intent from traffic patterns, and that must sit inline to act on what they infer, come out of the inline path; the same analysis can run on a copy of the traffic on the supervisory side, where being wrong produces a false alert rather than a dropped control packet.
And the fourth item on that list is the one that is not a product at all. Compliance artifacts standing in for protective layers. A documented procedure is not an interlock. An access control policy is not a relief path. A control marked as implemented in an assessment is a statement about the assessment, and Part I said what that is worth.
6. What the safety standards actually say#
I have had the conversation in section 5 many times, and the objection I get from the security side is that hardwired interlocks are an old-fashioned security control. That misreads what they are, and the misreading matters enough to spend a section on.
A hardwired interlock is not a security control that happens to be analogue. It is a protective layer, in the sense the functional safety standards define, and its claim rests on a property those standards already require, already name, and already know how to account for. That property is independence.
6.1 What IEC 61508 and IEC 61511 are for#
IEC 61508 is the generic functional safety standard for safety-related systems built from electrical, electronic and programmable electronic technology [4]. IEC 61511 is its implementation for the process sector, covering safety instrumented systems [5]. Between them they set out how to determine how much risk reduction a protective function must deliver, how to design and verify a function that delivers it, and how to maintain the claim over the life of the plant.
The structure both standards impose is layered. A hazard is identified, the unmitigated consequence and frequency are assessed, and a required risk reduction follows. That reduction is then allocated across protective layers, each of which is credited with a share. A layer of protection analysis makes the allocation explicit: the basic process control system holds the process in range, an alarm gives an operator a chance to intervene, a safety instrumented function trips the process on a defined demand, and a mechanical relief device limits the consequence if everything above it fails.
The arithmetic of that allocation is multiplicative, and this is the hinge of the whole section. Layer credits multiply only if the layers are independent. If two layers share a failure mode, the combined probability of failure on demand is not the product of the two; it is dominated by the shared mode. IEC 61511 is explicit that protection layers must be independent of the initiating cause and of each other, and that the safety instrumented system must be separate from the basic process control system to the degree needed to keep the claim intact.
6.2 What a safety integrity level is, and is not#
Because the term is used loosely in security writing, including in the source material I am revising here, it is worth being exact.
A safety integrity level is a property of a safety function, not of a device. There is no such thing as a SIL 3 sensor in isolation. There is a SIL 3 safety instrumented function, comprising a sensor subsystem, a logic solver and a final element, with an assessed average probability of failure on demand for the whole loop. A device may be certified as suitable for use in a function of a given integrity level, which is a different and weaker statement.
The levels are bands of probability of failure on demand, averaged, for functions operating in low demand mode. SIL 1 is a probability between one in ten and one in a hundred, a risk reduction factor between ten and a hundred. SIL 2 is one in a hundred to one in a thousand. SIL 3 is one in a thousand to one in ten thousand. SIL 4 is a further factor of ten and is rare in practice, because the verification burden is severe and because the usual engineering answer at that point is to change the process rather than protect it harder.
Those numbers are about random hardware failure, and that is the second thing routinely got wrong. The probability arithmetic rests on component failure rates that are estimated from populations of identical components failing independently over time. That model is appropriate for a solenoid coil or a pressure transmitter. It is not appropriate for software.
6.3 Software failure is systematic, not random#
Software does not wear out. A software failure is a design fault that was present from the moment of release and that manifests when the inputs reach the region where the fault lives. The standards call this a systematic failure, and they are clear that systematic failures cannot be quantified the way random hardware failures can. What IEC 61508 offers instead is systematic capability: a claim about the rigour of the process by which the software was produced, verified and managed.
That is a reasonable engineering response and I do not criticize it. But note what it means for the argument of this paper. Every claim about the reliability of a software protective element is a claim about process discipline, not a measured failure rate. And every piece of software introduced into a plant carries a failure probability that nobody can put a number on, which is precisely the condition section 5.1 identified: an unreliable estimate, calling for a structural answer rather than a better estimate.
6.4 Common cause is the quantity the barbell manages#
Now put sections 6.1 and 6.3 together.
Layer independence is what makes the risk reduction arithmetic valid. Software failure is systematic and unquantified. A software component shared across two or more protective layers is therefore a common cause: not a small correction to be handled with a beta factor, but a term whose magnitude nobody can state.
Look at how often that sharing happens once you go looking. The same operating system under the control system and the safety logic solver's engineering workstation. The same vendor's firmware family in the transmitter feeding the control loop and the transmitter feeding the trip. The same network stack implementation licensed by several device manufacturers from the same supplier. The same certificate authority, the same time source, the same remote access tool used to reach both. And, most relevant here, the same endpoint agent installed on every host in the facility because the security standard asked for coverage.
This is what July 2024 demonstrated at scale, and why I said in section 3.3 that it is the argument in miniature. The organizations affected had redundancy. What they did not have was independence, because every redundant instance shared one software component and that component failed in all of them simultaneously. Redundancy without independence is not risk reduction; it is a larger number of copies of the same risk.
A protective layer implemented in mechanical or discrete analogue technology has no software, and therefore shares no software with anything. IEC 61508 recognizes this category explicitly: a protective layer that is not electrical, electronic or programmable electronic is treated as other technology, and is credited as a separate layer precisely because its failure mechanisms have nothing in common with the programmable ones. A relief valve does not care what operating system the logic solver runs.
That is the whole basis of the first end of the barbell, and it is a functional safety argument rather than a security argument. It happens to also be a security argument, which is the next point.
6.5 The security claim follows from the safety structure#
An adversary needs a path. Every attack, however sophisticated, is a sequence of reachable states ending at something that moves.
A protective layer with no network interface and no programmable content terminates every such sequence. There is no remote code execution against a spring. There is no privilege escalation on a bimetallic strip. The device cannot be reconfigured over a network because there is no configuration and no network. Whatever the adversary achieves upstream, the trip point is where it stops.
The standards have started to meet this from the other direction. The second edition of IEC 61511 requires a security risk assessment of the safety instrumented system, which is an acknowledgment that a safety function implemented in programmable technology on a connected network is exposed in a way the first edition did not contemplate. And IEC 62443, which I use constantly and which I said in Part I encodes real knowledge, is explicit that security countermeasures must not compromise essential functions [6]. Section 3 of this paper is, read charitably, a catalog of deployments that violate that requirement while claiming conformance to the standard that states it.
So the position I am arguing is not outside the standards. It is a straightforward reading of what the standards already say, applied by someone who takes the independence requirement literally.
7. Objections I take seriously#
The barbell is not free and I would not trust an argument that presented it as free. Here are the objections I meet, with what I actually say to them.
7.1 You cannot retrofit mechanical protection everywhere#
This is true and it is the strongest objection.
A greenfield facility can be designed with its protective layers in the right technology from the start, at a cost that is a small fraction of the project. A plant that has been running for thirty years cannot be re-engineered that way without an outage nobody will authorize and a capital case nobody will sign.
What I argue on brownfield sites is narrower. Identify the hazards whose consequence is unbounded, meaning the ones where the worst case is loss of the asset, release of containment, or harm to a person. That list is always shorter than people expect, usually a handful per site. For those, and only those, establish whether the final protective layer contains software, and whether that software shares anything with the control path. Where it does, the retrofit has a case that can be made on functional safety grounds alone, without any reference to security, which is usually the only way it gets funded.
Everything else stays as it is. The barbell is not a demand to rebuild the plant. It is a demand to know which layer is the last one and what it is made of.
7.2 Mechanical devices fail, and they fail silently#
Also true. A rupture disc can corrode and burst below its rating, or fatigue and fail to burst at it. A relief valve can stick after years without operation. A bimetallic contact can weld. A mechanical governor linkage can seize.
The difference is not that these devices are reliable in some absolute sense. It is that their failure mechanisms are physical, known, and countable, which is what makes proof testing meaningful. You can pull a disc and inspect it. You can lift a valve on test. You can inject a signal and confirm a trip. The standards require exactly this, with a defined proof test interval that falls out of the probability of failure on demand claimed for the layer, and a failure to perform the test invalidates the claim.
Nobody can proof test the failure modes of an endpoint agent, because nobody has the list.
There is a real failure mode here that I do want to name, because it is the one that bites: a protective layer that exists on the drawing, is credited in the layer of protection analysis, and has not been tested within its interval. That is worse than not having it, because the risk accounting assumed it worked. I have found this on sites more often than I have found a compromised controller.
7.3 A one-way boundary is not free either#
The diode is not the end of the problem, it is the beginning of an operational discipline.
Everything the supervisory side needs must be arranged to flow outward, which means the telemetry design has to be done properly up front rather than fixed later by opening a return path. Protocols that expect acknowledgment need proxying on both sides. Time synchronization has to come from a source inside the control zone rather than through the boundary. And the day will come when somebody needs to push a configuration file inward, which is the moment the entire architecture is decided: either there is a controlled physical process for that, involving a human and removable media and a change record, or somebody adds a small exception and the boundary becomes a firewall rule with extra steps.
I have seen that exception added. It is always for a good reason and it is always permanent.
7.4 Removing the agents removes the detection#
This is the objection I have most sympathy with, and the answer is not that detection does not matter.
The answer is that detection does not require an agent on the control host. A network tap or span port feeding a passive decoder on the supervisory side sees the control traffic, the protocol exchanges, the command sequences and the timing, without placing anything in the control path and without sending a single packet into it. Device state can be obtained from the control system's own data, which already exists and is already being polled. Configuration change can be detected by comparing successive read-only captures.
What this arrangement gives up is host-level visibility: process creation, file changes, memory behavior on the endpoint. That is a real loss and I state it as one. What I ask in exchange is a comparison on the correct terms. Weigh the detection you lose against the trip you no longer risk causing, and against the correlated failure you no longer own, and then ask which of those two quantities appears in your risk register. In my experience only the first one does, which is the Part II problem repeating at the level of a procurement decision.
7.5 Regulation pushes the other way#
Sometimes it does, or is read as doing so. An assessor who expects endpoint protection on every host, and who has no category for an argument about scan cycle determinism, will mark a site down for not having it.
My experience is that this is more often a problem of how the case is written than of what the standards require. IEC 62443 requires that essential functions not be compromised by countermeasures, and the functional safety standards require independence. A control implemented differently, with the engineering reason stated and the compensating measure named, is a position I have defended successfully in assessments. A blank field in a spreadsheet is not, and that is a documentation failure rather than an architectural one.
This does raise a question the engineering cannot answer on its own, which is how much of a facility's finite capital should go to each end of the barbell and on what evidence, and that question is taken up in Part V alongside the Gordon and Loeb bound on security investment [7] and the underwriting position under Lloyd's Market Bulletin Y5381 [8].
8. Physics is the only zero trust that cannot be renegotiated#
The security industry has taken the phrase zero trust and turned it into a procurement list. Identity brokers, micro-segmentation agents, policy engines, cloud access proxies. Each of them is software that must be trusted in order to do its job of not trusting anything else, and the regress is not a debating point: it is the reason the list keeps growing.
In a plant the phrase can be made to mean something exact, and the meaning is uncomfortable for the people selling the list. Zero trust means refusing to trust software to prevent physical destruction.
Software is malleable. It can be changed remotely, and if it can be changed by an authorized party then the question of destruction reduces to the question of authorization, which reduces to more software. Its failure modes are systematic, present from release, and not enumerable by anyone including the people who wrote it. Its supply chain is deep, shared, and mostly invisible from the plant floor.
Physics does not have those properties. A spring compressed to a set force releases at that force. A bimetallic strip bends at its temperature. A disc bursts at its pressure. These devices have no authorization model because they take no instructions. They have no update path because they have no content. They cannot be reached from a network because they are not on one, and that is not a configuration choice that a future change request can reverse.
Anchor the plant's survival to those, put the computation where being wrong costs nothing, and the question of what the adversary can do collapses into a much smaller question: what can they do that the physics will absorb. That is a question with an answer, and the answer does not depend on how good anyone's software turns out to have been.
9. References#
[1] N. N. Taleb, Antifragile: Things That Gain from Disorder. New York: Random House, 2012. Cited for the barbell strategy, which is developed at length in this book and is the source of the structural argument in section 5.
[2] N. N. Taleb, The Black Swan: The Impact of the Highly Improbable. New York: Random House, 2007. Cited for the earlier and less developed appearance of the barbell, and for the treatment of unreliable risk estimates that motivates it.
[3] N. N. Taleb, Fooled by Randomness: The Hidden Role of Chance in Life and in the Markets. New York: Texere, 2001. Cited only to record that the barbell does not come from this book, which is the source of the survivorship argument in Part I of this series.
[4] International Electrotechnical Commission, IEC 61508, Functional safety of electrical/electronic/programmable electronic safety-related systems. Cited for safety integrity levels as properties of functions, for probability of failure on demand, for the distinction between random hardware failure and systematic failure, and for the treatment of protective layers implemented in other technology.
[5] International Electrotechnical Commission, IEC 61511, Functional safety, Safety instrumented systems for the process industry sector. Cited for layer of protection analysis and protection layer independence, for separation between the safety instrumented system and the basic process control system, for proof test intervals, and for the security risk assessment requirement introduced in the second edition.
[6] International Electrotechnical Commission, IEC 62443, Industrial communication networks, Security for industrial automation and control systems. Cited for the requirement that security countermeasures not compromise essential functions, and, as in Part I, as a framework in general use by the author rather than as a target of this argument.
[7] L. A. Gordon and M. P. Loeb, "The economics of information security investment," ACM Transactions on Information and System Security, vol. 5, no. 4, pp. 438 to 457, 2002. Named here only to identify the allocation question handed to Part V, where it is developed.
[8] Corporation of Lloyd's, Market Bulletin Y5381: Cyber-attack exclusions, 16 August 2022. The associated model clauses LMA5564 to LMA5567 are published by the Lloyd's Market Association. Named here only to identify the underwriting question handed to Part V, where it is developed.
Note on this bibliography#
The failure mechanisms described in section 3, the account of how tooling migrates down the Purdue levels in section 2.3, the residue argument in section 4.1, and the objections and responses in section 7 are the author's own observations, drawn from assessment and engineering work in substations, battery energy storage sites, rail control rooms, water treatment plants and hyperscale data centres across the energy, rail, maritime and manufacturing sectors in North America, Australasia and Europe. They carry no citation because no published source records them, and they are offered as testimony rather than as findings.
The scan cycle periods, scheduling delays, burst pressures and trip temperatures used in section 3 and section 5.3 are illustrative figures of the order encountered in the field. They are given to make the mechanisms concrete and are not measurements of any particular installation.
The endpoint sensor failure of July 2024 referred to in sections 3.3 and 6.4 is described from public reporting of a widely observed event. It is named for its structure, which is correlated failure through a shared software component, and no vendor-specific claim is made beyond what was reported at the time.