Fooled by Best Practice, Part II: The Two Sides of the Table
J. McKenney
Paper 2 of five in Fooled by Best Practice. Part I made the argument in ordinary language. This part states the distinction it rests on, shows where the drawing and the plant come apart, and gives the turkey problem its proper form. Part III gives the method. Part IV gives the engineering response. Part V gives the decision.
Licence: CC BY 4.0. 16 September 2026.
| Field | Value |
|---|---|
| Document ID | WG-02-DT-FBP-2 |
| Slug | fooled-by-best-practice-2-two-sides |
| Working group | WG-02-DT, Digital Twin |
| Series | Fooled by Best Practice, part 2 of 5 published |
| Author | J. McKenney |
| Published | 16 September 2026 |
| Revision | |
| Status | Published |
| Preceded by | WG-02-DT-FBP-1 |
| Followed by | WG-02-DT-FBP-3, WG-02-DT-FBP-4, WG-02-DT-FBP-5 |
Executive Abstract#
Everything a security programme produces sits on one of two sides of a line, and almost all of it on the same side. On the right is the record: the passed audit, thirty-six months with no serious incident. All of it real, all generated by the one history the facility lived through, following Taleb in finance.
On the left is what an operator actually wants: not what happened but how likely something else would have. A plant running three years takes one route through an enormous space of routes, some ending with loss of control or a failed safety function. None appear in any record, because they did not occur.
This paper does three things: sharpen the line with two questions that sort any artifact in seconds; locate it inside a plant, as the distance between the designed and operated system, where controls against an unknown attack path number zero; and state the turkey as measured stability rising right up to the transition, the phase-transition language an analogy, not a model.
None of this argues against standards. I work inside IEC 62443; the standards are careful about what they claim, and the claim that goes wrong is made by the reader.
Abstract#
This paper sets out the epistemological partition the series depends on and locates it inside the operating plant. A facility over an interval realizes one history from a space carrying a probability measure. Right-side quantities are functions of the realized history, like conformance results; left-side quantities are functionals of the measure, like coverage over reachable paths. Accuracy on the right does not become relevance on the left. A two-question test classifies any artifact, and worked exercises show how buying left-side information means realizing a branch. The designed and operated system diverge continuously in one direction; drift, a distance between the two descriptions, is only ever a lower bound. The gap is the region of zero specified control coverage: the assessor stops when the document is satisfied, the adversary when he reaches something that matters. The turkey problem is a system whose measured stability increases monotonically up to the transition that destroys it; of its three mechanisms, a mean-field Ising formulation describes the one where effort flows toward the measured quantity. No parameters are measured, nothing computed is a prediction, but it supplies a testable direction: early degradation in variance and recovery time, not the mean. The distinction rewrites three operator questions without making the left side observable, which requires an instrument that samples the branches, the subject of Part III.
1. Introduction#
1.1 Where the argument stands#
In Part I, I argued that the industry I work in infers defensive competence from the absence of observed compromise, and that this is the same inference Taleb dismantled in finance [1]. I gave the arithmetic, I separated the two ways evidence goes missing, and I described four forms the error takes in the field. I am not going to repeat any of it.
What I did not do in Part I was state the distinction the whole argument rests on with enough precision to act on. I used the phrase "two sides of the table" and promised this part would define it. That is what I am doing here.
I want to be honest about why the precision matters to me and not only to the argument. I have sat in a great many closing meetings at the end of an assessment. The pattern in those rooms is consistent. Everything the client shows me is true, carefully produced, and expensive. The register is current as of the date on it. The audit finding was closed and the evidence is attached. The scan report was run against the list that was given to the scanner. None of this is theatre and I have never thought it was. What I have thought, repeatedly, is that we were all looking at a large quantity of accurate information about the wrong quantity, and that I did not have a clean way of saying so without sounding like I was calling the work worthless.
The distinction in section 2 is the clean way of saying it. It took me a long time to get it into a form short enough to use in a meeting.
1.2 What this part has to establish#
Three things, in order.
First, the partition itself, stated precisely enough that a reader can pick up any artifact in their own programme and sort it. I give a test rather than a list, because lists go out of date and because the interesting cases are the ones nobody thought to list.
Second, where the partition falls inside a plant. The abstract version of the two sides is easy to nod along with and easy to forget by Thursday. The concrete version is the distance between the system that was designed and the system that is running, and that distance is measurable, which makes it something an engineer can actually work with.
Third, the turkey, stated as a property of a curve rather than as a fable. The useful content is not that bad things happen. It is that in a specific and identifiable class of systems, every available measurement improves right up to the moment of failure. If you are in that class, your instruments are not merely insufficient. They are pointing the wrong way.
The reading level here is professional. There is a little mathematics, because two of the arguments are hard to state carefully without it, but there are no derivations and nothing that requires more than an engineering degree from some years ago.
Two boundaries I will respect. Part III owns the method for sampling the left side, so I set the problem up and stop. Part IV owns the engineering response and the case against enterprise tooling in operational technology, so where an example of mine points that way I will note it and move on.
2. The two sides of the table#
2.1 The partition#
Taleb's formulation in Fooled by Randomness is a division between the world as it is experienced, which is sequential and narrative and made of things that did happen, and the world as probability describes it, which is a space of outcomes of which the realized one is a single draw [1]. He is writing about traders. The partition is not about traders.
Here is the version I use.
Take a facility over some interval, say the last three years. Over that interval the facility realized exactly one history. Call it the path that ran. It contains everything that happened: every configuration change, every alarm, every maintenance visit, every adversary who looked at the facility and every adversary who did not, every time an operator noticed something and every time nobody did.
Now consider the set of histories the facility could have realized over that same interval, given the same plant, the same people, the same adversaries and the same world. That set is enormous and it carries a probability measure, whether or not anyone can compute it. The realized history is one draw from it.
The right side of the table holds quantities that are functions of the realized history alone. The audit result. The incident count. The mean time to detect, computed from detections that actually occurred. The number of vendor deployments that were not compromised. The patch percentage on the date of the report. The dashboard color.
The left side of the table holds quantities that are functionals of the distribution. The probability that the facility loses control of a process in the coming year. The probability that a safety function is called on and does not act. The proportion of reachable paths against which a given control actually has effect. The size of the loss at the ninety-ninth percentile rather than at the median.
The two sides are not two opinions about the same quantity. They are different quantities. A right-side number can be measured to any precision you are willing to pay for and it will still not be a left-side number. Accuracy on the right side does not become relevance on the left.
That sentence is the whole of section 2. Everything after it is application.
2.2 A test for sorting#
Two questions. Either one will do; I use both because they fail differently.
Question one. Could this artifact have been produced by someone with access only to what the plant did, and never to what it might have done? If yes, it is a right-side artifact.
Question two. Hold the realized history exactly as it was and vary the branches that did not run. Does the artifact's value change? If no, it is right side. If yes, it is left side.
Work an example. A facility passed its assessment in March with two minor findings. Question one: the assessor needed the plant's documents, its configurations and its interviews, all of which are features of what the plant did. Nothing about unrealized branches entered the assessment. Right side. Question two: imagine a world identical in every observed respect, but in which a particular contractor laptop had been carrying something in February. The assessment result in March is unchanged. Right side, confirmed.
Another. The proportion of simulated adversary campaigns in which a given segmentation boundary is reached. Question one: no. You cannot produce that number from the record, because the record contains one campaign at most and usually none. Question two: yes, obviously; it is a statement about the branches and nothing else. Left side.
A third, which is the one that catches people. A threat model. An engineer sits with the drawings and enumerates twenty attack paths, rates them, and writes them up. Is that left side? It is aimed at the left side, which is not the same thing. What it produces is a set of paths that a particular person thought of on a particular afternoon, which is a function of that person and that afternoon. It says nothing about the paths not enumerated, and the paths not enumerated are the ones that matter, because an adversary who only uses paths your engineer already thought of is not the adversary you are worried about. A threat model is a right-side artifact with left-side ambitions. It is worth doing. It is not a measurement of the distribution.
2.3 Sorting the usual artifacts#
Applying the test across the things a typical programme produces gives the following, and I include the reasoning in each row because the reasoning is the part you can transfer to artifacts I have not listed.
| Artifact | Side | Why |
|---|---|---|
| Conformance result against a framework | Right | A statement about which specified controls were present in the assessed scope on the assessed date |
| Months since last significant incident | Right | A count over the realized history and nothing else |
| Vulnerability scan results | Right | What responded to the scanner, from the list the scanner was given |
| Mean time to detect | Right | An average over detections that occurred, and silent about the events never detected |
| Patch or configuration compliance percentage | Right | A state of the estate at an instant |
| Vendor reference customers | Right | A sample conditioned on the outcome, as Part I set out |
| Training completion rate | Right | A count of completions, not a measurement of behavior under load |
| Enumerated threat model | Right, aimed left | A function of who enumerated and when |
| Tabletop or live exercise | Right, buying a little left | Deliberately realizes one branch that would not otherwise have run |
| Physical failure test, fault injection, proof test of a safety function | Right, buying a little left | Same, and considerably more informative, because the branch is realized in the plant rather than in a discussion |
| Exceedance probability over a stated period | Left | A functional of the distribution by construction |
| Proportion of reachable paths a control covers | Left | Requires the set of paths, which the record does not contain |
| Loss at a stated upper quantile | Left | Same |
The two exercise rows are the interesting ones, and they are the reason I said the test fails differently in each direction.
An exercise is an attempt to buy left-side information at retail. You take a branch that did not run naturally and you run it on purpose, and you observe what the plant and the people do. This is genuinely informative and it is the only left-side information most facilities ever acquire. A proof test of a safety instrumented function is exactly this: the demand did not arrive, so you manufacture the demand and see whether the function acts. The functional safety standards are built around that idea and they are right to be [4], [5].
The problem is arithmetic. One exercise samples one branch. The honest question to ask after any exercise is not what did we learn, it is how many branches were there and how representative was the one we picked, and the answer to the first half is a very large number. You cannot buy enough branches at that price. That limitation is precisely why the method in Part III exists, and I will not go further into it here.
2.4 Why the right side is not worthless, stated carefully#
I said this in Part I and I will say it once more in this part because it is where the argument is most often misread, and then I will not say it again in the series.
Right-side artifacts are correct. They are also narrow. An audit finds real defects, and finds them cheaply, and a facility that fixes them is better off than one that does not. A conformance programme encodes knowledge accumulated from real incidents that you would otherwise have to acquire the expensive way. None of that is in question.
There are two properties I want to give the right side explicitly, because they are its actual strengths and they get lost when people defend it on the wrong grounds.
The right side is cheap and repeatable. You can run it on a schedule, across an estate, with staff who were trained this year, and get comparable numbers. Nothing on the left side has that property.
The right side catches the low end of the distribution reliably. Facilities that have no segmentation at all, no asset register, no patch process and no logging are found by right-side instruments immediately. The instrument works. It works on the part of the problem that is visible to it, which is the part where controls are absent rather than the part where controls are present and irrelevant.
What the right side cannot do is any of the following. It cannot tell you how close you came. It cannot tell you the probability of a specific facility being compromised in a specific period. It cannot distinguish a facility whose defences held from a facility nobody capable was looking at. And it cannot tell you about a path that was not in the scope it was assessed against, which is the subject of section 3.
2.5 Why a clean record is weak evidence#
There is one piece of arithmetic that I think belongs in every discussion of this and almost never appears, and it takes a paragraph.
Evidence strength is a ratio, not a level. The question is not whether a clean thirty-six months is likely if the defences are sound. Of course it is. The question is how much more likely it is than a clean thirty-six months if the defences are not sound. Write those two as a likelihood ratio:
If both numbers are close to one, the observation carries almost no information, no matter how emphatically it is reported.
Suppose, and I want to be clear that these two numbers are illustrative and that I have not measured either, that a sound facility has a two percent chance of a serious incident in three years and an unsound one has a ten percent chance. Then the numerator is 0.98, the denominator is 0.90, and the ratio is about 1.09. Three clean years shift your prior odds by nine percent. Nine percent is not nothing. It is also not what the sentence "we have not had a significant incident" is doing when it is said in a meeting, which is closing the question.
The same point comes at you from the direction of return periods. If an event class has a return period of two hundred years at a given site, the chance of observing at least one in a three-year window is roughly
A clean three-year record is what you expect to see, with about ninety-eight point five percent probability, whether the site is well defended or not. The observation and the hypothesis are nearly independent. Nothing was learned.
This is why the observation window matters more than the observation. A right-side record becomes informative when its length is comparable to the return period of the thing you are worried about. For a nuisance event with a return period of months, three years of data is genuinely informative and you should use it. For the class of event that destroys equipment or kills somebody, the plant will be decommissioned before the record is long enough to say anything. The record cannot get there. This is not a data quality problem that better logging solves.
3. Where the map stops matching the territory#
3.1 Two different systems with the same name#
Taleb's standing warning about models is the confusion of the map with the territory [1], [2]. In a plant that confusion has a specific address, and it is not a metaphor there. It is two documents that contradict each other and one of them is a building.
Every conformance statement is made against a description of a system. This is not a criticism; it is how conformance has to work, and the standards say so plainly. IEC 62443 organizes its requirements around a system under consideration, with defined zones and the conduits between them, and a conformance claim is scoped to that definition [3]. The functional safety standards do the same through the safety requirements specification, and they go further, because their lifecycle explicitly requires that changes to the installed system be assessed against that specification rather than simply made [4], [5]. The standards know about drift. They have known about it for a long time.
So there are two systems.
The designed system is the one described: the drawings, the zone and conduit model, the asset register, the specified configurations, the reference firmware builds, the network diagram on the wall. It is a complete and consistent object, because it was constructed to be.
The operated system is the plant. It is the accumulated result of every decision taken since commissioning, most of them correct at the time, none of them coordinated with the drawing.
The assessment is performed against the first. The adversary operates against the second.
3.2 How the divergence accumulates#
The divergence is not caused by carelessness and I want to remove that explanation early, because it leads to the wrong remedy. It is caused by the plant doing its job.
A vendor needs access during a commissioning window, so an access path is opened. The window closes. The path does not, because closing it requires somebody to know it exists, to own the decision, and to be confident that nothing depends on it.
A bridge is put in between two segments because a data flow is needed on Tuesday and the properly engineered route will take six weeks. The properly engineered route is scheduled. Then it is rescheduled. The bridge is now load-bearing.
An upgrade replaces two thirds of a system, and the last third is absorbed rather than replaced, because replacing it means an outage on a process that does not have an outage window this year, or next.
Firmware is frozen because updating it requires stopping something that cannot be stopped, and the risk of the update is immediate and legible while the risk of not updating is diffuse and deferred. Every engineer reading this has made that trade and most of them made it correctly.
A card fails and is replaced like for like. The replacement carries newer silicon, a later firmware build, and a second network interface enabled by default that the original did not have. Nothing in the change record is wrong. The plant now has an interface nobody specified.
A telemetry or metering programme, run under a different contract with a different budget holder, installs cellular modems in cabinets across a site. They are not control network devices, so they are not on the control network drawing. They are in the cabinets.
A historian connection is added for a commercial reason, through a gateway commissioned by a project that closed two years ago, and the serial links behind that gateway appear on no diagram that is drawn in terms of addresses, because they do not have addresses.
I have found every one of these. Not in one plant, and not all in bad plants. The battery storage sites and the substations are where I find the modem and the absorbed legacy equipment most often. The water treatment plants are where I find the frozen firmware, for the obvious reason. The rail control rooms are where I find the access path left open, because the engineering support model demands it. The hyperscale data centres are where I find the best documentation in the industry and the divergence anyway, because the rate of change there is higher than anywhere else I work.
Now the part that makes this structural rather than anecdotal.
The plant changes continuously. The description is updated at discrete events: a project milestone, a major upgrade, an assessment that forces a walkdown. Between those events the two drift apart, and at each event the description jumps to catch up, mostly. There is no routine mechanism that pushes the operated system back toward the designed system, because the plant is rewarded for availability and nobody is rewarded for removing a bridge that is working. The forces are asymmetric, so the drift has a direction.
3.3 Drift as a measurable quantity#
If the divergence has a direction it has a magnitude, and an engineer should be able to put a number on it. Write both systems as graphs over the same node and edge types: assets, interfaces, connections, software components, trust relationships. Let be the designed system and the operated system at time . Define
for some distance over that space. A simple and perfectly usable choice is a count of edges and nodes present in one graph and absent from the other, weighted by whether the difference sits on a route that reaches something you care about, because an undocumented link between two irrelevant devices is not equivalent to an undocumented link into a safety domain.
The particular choice of distance matters much less than three properties. The quantity must be defined, so that two people measuring it get comparable answers. It must be measured on a schedule rather than at project milestones, because measuring it at milestones reproduces the very sampling that hides it. And it must be trended, because the level is less informative than the slope: a plant with a large stable drift that everyone understands is in better shape than a plant with a small drift that doubled this quarter.
Two honest limitations, and I would rather state them than have somebody find them.
The first is the one that matters. A measured drift is a lower bound. You can only measure the divergence you found. The divergence you did not find is absent from the measurement precisely because you did not find it, which is exactly the structure Part I described one level up, reappearing inside the instrument meant to address it. A drift figure that has gone down may mean the plant converged on its documentation. It may also mean this survey was less thorough than the last one. You will not be able to tell from the number.
The second is that anything can be moved to the right side of the table by scoring it. If a drift measurement becomes a percentage on a slide, reported quarterly, with a target, then within about two cycles it will be managed as a percentage on a slide. The measurement will improve. I would put no money at all on the plant improving with it. I am aware that I am proposing a metric in a paper about the failure of metrics, and the only defence I have is that this one measures a gap rather than a conformance, and that I have said out loud what will happen to it if you treat it as a score.
3.4 The gap is where control coverage is zero#
Here is the argument this section exists for.
Controls are specified against a described system. The risk assessment that justified them enumerated paths through that description. The control set was chosen to cover those paths, and the assessment of residual risk was computed over them.
A path that exists only in the gap between the described system and the operated system was not in that enumeration. It was not assessed, not rated, not accepted and not mitigated. There is no control specified against it, and there is no record of a decision to leave it uncovered, because no decision was ever taken.
This is worth separating carefully from the failure mode people expect. The common picture of a breach is that a control was present and did not hold: the firewall rule was too permissive, the detection did not fire, the patch was late. Those happen. But the failure I am describing is different in kind. Nothing failed. The control coverage in that region was zero and always had been, and the assessment reported a residual risk figure computed over a set of paths that did not include the one used.
So the drift is not a documentation defect to be tidied up when someone has a quiet fortnight. It is the region of the plant where the specified coverage is identically zero, and its size is the thing your control set has no opinion about.
There is an asymmetry in how the two parties search it, and this is the sentence I use when I need one sentence.
The assessor searches from the document toward the plant and stops when the document is satisfied. The adversary searches from the plant and stops when he reaches something that matters.
These are two different searches over two different sets. One of them is bounded by a drawing. The adversary has no drawing, does not want one, and is not looking for the same thing anyway. He enumerates what responds.
3.5 What closes the gap and what does not#
Three things that do not close it, in ascending order of how often I see them proposed.
Redrawing the diagram does not close it. It moves to where was on the day of the survey, which is useful, and then both resume their previous behavior. If the redraw happens every three years and the drift accumulates continuously, you have a sawtooth, and the interesting question is the area under it rather than the value on the day after the redraw.
Tightening the standard does not close it. The standards are already scoped to a described system and they already require change management [3], [4], [5]. The requirement is not missing. What is missing is that the described system is maintained as a project deliverable and the plant is operated as a continuous process, and no document can force those two cadences to match.
Adding a discovery tool to the network does not close it either, and this is where I have to be careful, because the case for and against instrumentation inside operational technology is Part IV's subject and I am not going to argue it here. The narrow point that belongs in this part is that a discovery tool finds what responds to it, on the segments it is placed on, using the protocols it speaks. That is a real and useful addition to and it is not the same object as . Serial links, devices that do not answer, air-gapped islands, and anything on a segment nobody put a sensor on remain outside it. You are still measuring a lower bound. You have just raised it.
What does help, in my experience, is unglamorous and organisational. Every temporary change gets an expiry date recorded at the moment it is made, and the list of temporary changes past their expiry date is a standing report that somebody senior reads. Every commissioning window closes with an explicit enumeration of what was opened. Replacement parts are checked against the defaults of the part they replaced, not against the part number. And the walkdown happens on a schedule that is not derived from the audit schedule, because a walkdown timed to the audit will find what the audit is going to look for.
None of that is exciting and none of it is a product. It is the only thing I have seen reduce the slope.
4. The turkey problem in its proper form#
4.1 The statement#
The turkey appears in The Black Swan, published in 2007, and not in Fooled by Randomness [2], [1]. Part I made that attribution and I am reinforcing it here because the misattribution is common and the two books are making related but distinct arguments. Taleb himself traces the example back to Bertrand Russell's chicken and through it to Hume's problem of induction, and says so.
The fable is short. A turkey is fed every morning by a farmer. Every feeding is a data point supporting the hypothesis that the farmer is a reliable source of food. The turkey's confidence in that hypothesis, by any reasonable statistical treatment of the evidence available to it, is highest on the afternoon before Thanksgiving.
Stated as a property rather than a story, which is how I want it for the rest of this paper:
The turkey problem is the case of a system whose measured stability increases monotonically right up to the transition that destroys it.
Every clause in that sentence is load-bearing. Measured stability, because the underlying stability was doing something else entirely. Increases, because the measurements are not merely uninformative, they are actively moving in the reassuring direction. Monotonically, because there is no dip, no wobble, nothing that a trend analysis would flag. Right up to, because the last measurement before the transition is the best one in the record.
A system with that property cannot be defended by improving the measurement. More data extends the monotone run. Better instrumentation measures the same quantity more precisely. Longer history makes the trend more convincing. Every response that the right side of the table is capable of generating makes the confidence worse and leaves the exposure where it was.
4.2 Three different mechanisms, which need separating#
The turkey shape arises in at least three ways, and they get conflated. They have different remedies, so conflating them is expensive.
Short window against a long return period. The generating process is stable, the event is genuinely rare, and the observation record is simply too short to contain one. This is section 2.5 again. The turkey is not in this category, because the farmer's behavior was not a rare draw from a stable process.
A non-stationary generating process. The distribution itself changes, and the observed record was generated by an arrangement that is ending. This is the turkey exactly. The farmer's feeding and the farmer's knife are not two draws from one distribution; they are one purpose with two phases, and the feeding phase generated all the data. In a plant, this is the threat environment, the supply chain, the connectivity and the geopolitical picture all moving underneath a record that was accumulated before they moved. Part I made this point about a practice that has held for fifteen years, and I will not restate it.
Anticorrelation through effort allocation. The measurement and the risk move in opposite directions, because effort spent improving the measured quantity is effort not spent on the unmeasured one, and because a rising measurement reduces the perceived need to look further. This mechanism is the nastiest of the three because it is generated by the organization's own competence and diligence. A team that does not care will not produce a monotone rising curve. A team that cares a great deal will.
Most plant confidence I encounter is a mixture of all three, weighted toward the second and third. The first is the only one that more data fixes, and it is the one that is least often actually operating.
4.3 The phase transition analogy#
I want a description of the third mechanism that is precise about the shape, and the cleanest one I know comes from statistical physics. I am going to state it, then state its limits, and I would rather do the second at length than have the first taken for more than it is.
Consider a large number of individuals who each hold one of two dispositions, and who influence each other. In the mean-field Ising treatment, the aggregate alignment of the population is written as an order parameter , and its evolution under a simple relaxation dynamic is
with the aggregate alignment, the strength of influence between neighbours, the number of neighbours each individual has, an external field pushing the population in one direction regardless of what its neighbours do, and an effective temperature representing everything that randomizes individual behavior.
The behavior of that equation is well understood and it has one feature I care about. With no external field, there is a critical value of the temperature, . Below it, the system has two stable solutions with away from zero, and it sits in one of them. Above it, those solutions do not exist and the only stable state has . As rises toward from below, the stable value of falls, but for much of the approach it falls slowly, and then the branch it is sitting on ceases to exist.
The mapping I am proposing is the obvious one. stands for the coherence of the organization's defensive behavior: whether an unusual reading actually gets escalated rather than noted, whether a procedure is followed when following it is inconvenient, whether the third alarm in a shift gets the same attention as the first. and stand for how strongly and how widely that behavior is reinforced between people, which is why a single conscientious engineer in an indifferent team does not hold the value up and a tight shift team does. stands for external pressure, regulatory or corporate, pushing in a direction independent of local reinforcement. stands for everything that randomizes behavior: overtime, turnover, alarm load, budget pressure, reorganization, an unfilled vacancy in its ninth month.
The analogy gives me two things the fable does not.
It gives a reason why the aggregate indicator can hold while the underlying condition deteriorates. In this class of systems, the order parameter is not a linear read-out of the control parameter. It can sit high while climbs, and then it does not.
And it gives a reason why the collapse arrives without a proportional precursor in the aggregate. Nothing in the mean value of announces the approach. The organization looks the same on paper. Its documents are the same documents, its headcount is similar, its conformance result is what it was last year, and its capacity to coordinate a response to something genuinely anomalous is not what it was.
4.4 What the analogy does not establish#
Now the limits, and I want them itemized rather than gestured at, because this is the kind of borrowing that turns into a claim if you let it.
No parameters have been measured. I have never seen , , or estimated for any facility, and I have not estimated them myself. Anybody offering you a number for your plant's organisational temperature has produced it from somewhere other than measurement.
The equation has not been fitted to anything. There is no dataset of industrial security incidents with matched organisational observations that this could be fitted against, and if such a dataset existed it would be subject to exactly the disclosure problem Part I described, which means it would be a sample of the incidents somebody was willing to report.
The model's assumptions do not hold. Individuals are not two-state objects. Influence between people is not uniform, not symmetric, and not describable by a single coupling constant. The population is not large in the sense the mean-field treatment requires, since a control room team is a dozen people, not Avogadro's number. The external field is not scalar. Every one of these is a real departure and I am not going to wave at them.
Nothing computed from it is a prediction. There is no for your plant. There is no distance-to-transition you can put on a slide. If a vendor puts one on a slide, including one who has read this paper, the number is decorative.
So what is the analogy doing? It is doing explanatory work and nothing else. It gives a precise and non-mystical account of how a system can have an aggregate indicator that stays high while its underlying condition degrades, and then lose the stable state entirely. It shows that such systems exist, that their behavior is ordinary rather than exotic, and that in them the mean of the aggregate indicator is the wrong instrument. That is a genuine contribution to how you think about your dashboard and it is not a measurement of your plant.
A reader who rejects the analogy completely loses the vocabulary and keeps the observation, which is the thing I actually have evidence for: I have not once seen a compliance score fall in the quarter before a serious incident. The observation does not need the physics. The physics tells you why you should not have expected the score to fall.
4.5 Where the warning would be, if there is one#
There is one prediction the analogy does make, and it has the merit of being checkable in principle against things a plant already records.
In systems of this class, the approach to a transition shows up in the second moment before it shows up in the first. Fluctuations grow. The time taken to recover from a small disturbance lengthens. The system becomes slower to return to its normal state after being pushed, even while its normal state looks unchanged. This is a property of the model class rather than something I am asserting about your plant.
If that carries over at all, the things worth watching are not levels. They are dispersions and recovery times.
How long does it take, now, from an anomalous reading to a closed-out disposition, and how does that compare with a year ago? Not the average, which will be dominated by the easy cases. The spread, and the tail of the spread.
How many temporary changes are open past their intended end date, and is that count growing? That is a drift measure and a dispersion measure at the same time.
How often does the same finding recur after being closed? A finding that keeps coming back is not a finding, it is a description of the plant's steady state, and it says the closure mechanism is not coupled to the thing it claims to close.
When a shift is short-handed, what stops happening? Every team has an order in which things are dropped under load. It is rarely written down and everyone on the shift knows it. What is dropped first is the best single indicator I know of where the organization's real coherence sits, and it is not on any dashboard.
I want to be plain that I do not have a validated threshold for any of these. I am not able to tell you that a dispersion of such-and-such means you are within a year of something. What I can say is that these quantities degrade before the mean does, that they are computable from records most plants already keep, and that nobody is looking at them, because they are not what the audit asks for.
5. What the distinction is worth#
The partition does not make the left side observable. I want to close by being exact about that, because a paper that sharpens a distinction can leave the impression that sharpening it was the accomplishment.
You cannot read a distribution off a plant. There is no instrument you can attach to a substation that returns a probability of loss of containment over the next twelve months. The left side of the table is not a place you can go and look, which is why the industry ended up on the right side in the first place. That was not stupidity. It was the only side with instruments on it.
What the distinction changes is what a question means, and I can give that concretely as three rewrites.
Are we compliant? becomes: what system description was this conformance statement scoped to, how far has the plant moved from that description since the assessment, and what is the slope?
Have we had an incident? becomes: how many of the branches through this period ended differently, and what would it cost to find out about even a few of them?
Is the control in place? becomes: which paths was this control specified against, which paths exist that it was not specified against, and how would we know the difference?
In each pair, the second question is the one whose answer an operator actually needs, and in each pair the second question cannot be answered by looking harder at the record. The record does not contain it. No amount of diligence applied to the right side of the table crosses over, and the cost of buying left-side information one branch at a time, through exercises and proof tests, means you will never buy enough of it that way.
Which leaves exactly one option. If the risk lives on a side of the table you cannot observe, the only honest instrument is one that samples it: something that takes the plant as it actually is, including the drift, and runs the histories that did not happen. What that requires, what it can tell you, and the considerable list of things it cannot, is Part III.
6. References#
[1] N. N. Taleb, Fooled by Randomness: The Hidden Role of Chance in Life and in the Markets. New York: Texere, 2001. Cited for the two sides of the table as an epistemological partition, for the treatment of a track record as a single sample path, and for the confusion of the map with the territory.
[2] N. N. Taleb, The Black Swan: The Impact of the Highly Improbable. New York: Random House, 2007. Cited for the turkey, which appears in this book and not in [1], and which Taleb traces to Bertrand Russell's chicken and through it to Hume's problem of induction, and for silent evidence, which Part I treats in full.
[3] International Electrotechnical Commission, IEC 62443, Industrial communication networks, Security for industrial automation and control systems. Cited for the system under consideration and the zone and conduit model, and specifically for the point that a conformance claim is scoped to a stated system description. Cited in support of the argument in section 3, not as a target of it.
[4] International Electrotechnical Commission, IEC 61511, Functional safety, Safety instrumented systems for the process industry sector. Cited for the safety requirements specification and for the lifecycle requirement that modifications to the installed system be assessed against it, which is the standards anticipating the drift problem rather than being caught by it. Also cited for proof testing as the deliberate realisation of a demand that has not occurred.
[5] International Electrotechnical Commission, IEC 61508, Functional safety of electrical, electronic and programmable electronic safety-related systems. Cited as the parent standard of [4] in the same sense.
[6] National Institute of Standards and Technology, Framework for Improving Critical Infrastructure Cybersecurity. Cited as a framework in general use whose outputs are right-side artefacts by construction, and not as a target of this argument.
[7] North American Electric Reliability Corporation, Critical Infrastructure Protection reliability standards. Cited for periodic audited conformance assessed against a defined asset scope, in the same sense as [3] and [6].
Note on this bibliography#
The mean-field Ising formulation in section 4.3 is standard statistical mechanics and is stated here without citation to a particular text, because no particular text is needed for it and because the argument does not rest on any specific treatment. Its application to an operating organisation is mine, it is an analogy, and section 4.4 states at length what it does and does not establish.
The field observations throughout section 3, the accumulation mechanisms, the sectors in which I encounter each of them, and the practices in section 3.5 that reduce the slope, are the author's own, drawn from assessment and engineering work in substations, battery energy storage sites, rail control rooms, water treatment plants and hyperscale data centres across the energy, rail, maritime and manufacturing sectors in North America, Australasia and Europe. They carry no citation because no published source records them. They are offered as testimony, and where a claim in this paper rests on observation rather than on a source, the sentence says so.
The illustrative figures in section 2.5 are illustrative. They are chosen to show the shape of the likelihood ratio argument and they are not measurements of any facility or sector.