Bayesian Stackelberg Security Games & Strategic Asset Hardening under Epistemic Uncertainty
J. McKenney
This paper is part of the WG-07-TM Threat Modeling body of work, applying game-theoretic methods to industrial asset hardening. It sits alongside WG-07-TM-08-Adversarial-Game-Theory-Purdue-Enclaves, which develops the underlying adversarial game-theoretic framework over Purdue Model enclaves, and WG-07-TM-10-Algorithmic-Mechanism-Design-Bayesian-Persuasion, which applies a related Bayesian mechanism-design approach to defender signaling; this paper specializes that shared game-theoretic foundation to a Bayesian Stackelberg formulation under epistemic uncertainty about attacker type.
Licence: CC BY 4.0. 17 September 2026.
Executive Abstract#
Industrial facilities like power plants, chemical refineries, and water systems face attackers who are strategic and adaptive, not random. Security checklists and generic risk scores treat every asset alike and assume the defender already knows what an attacker wants, both usually false, leaving defenders guessing where to spend a limited budget.
This paper treats the problem as a game: the defender commits to a hardening strategy first, and the attacker observes it before choosing where to strike. Rather than one known attacker profile, it allows several plausible types, a ransomware operator chasing downtime, a nation-state actor seeking physical damage, and an espionage actor after long-term access, and finds a hardening policy that holds up against all of them when the defender cannot be sure which is real.
Tested against a simulated 1,200-node chemical plant network, this approach beats a standard risk-matrix allocation by a wide margin in the model, and produces a stated, quantified bound on worst-case cyber loss that a risk officer can put before a board or insurer.
Abstract#
Operational technology infrastructure, spanning thermal power plants, petrochemical refineries, water distribution networks, and digital substations, operates under an asymmetric threat environment: severe capital constraints, equipment lifecycles of 15 to 30 years, and availability requirements that preclude ad-hoc patching or indiscriminate segmentation. Qualitative risk matrices, CVSS scoring, and checklists fail to capture the strategic, adaptive nature of nation-state APTs, and they assume deterministic knowledge of adversary motives, capabilities, and payoffs. J. McKenney and the Eigenia Threat Modeling Working Group formulate industrial cyber defense as a Bayesian Stackelberg Security Game under epistemic uncertainty. The defender (asset owner or CISO) is a strategic leader committing to a randomized hardening policy across Purdue Model targets, Levels 0 through 3. The adversary is a rational follower who observes allocations, such as deep packet inspection conduits, cryptographic bump-in-the-wire modules, and hardware unidirectional gateways, and picks the target that maximizes its objective. Attacker intent is modeled as an ensemble of discrete types governed by imprecise probability distributions and credal sets. We solve the resulting optimization via a mixed-integer linear programming reformulation of the Decomposed Optimal Bayesian Stackelberg Solver (DOBSS) with minimax regret bounds. On a 1,200-node chemical plant network complying with IEC 62443-3-2 and the EU NIS2 Directive, the policy cuts the defender's worst-case Annualized Loss Expectancy by 64.2 percent versus risk-matrix prioritization, while holding provable bounds on cyber catastrophe Value-at-Risk at the 99 percent level.
1. Introduction and Problem Formulation#
Industrial cybersecurity operates under strict physical and financial constraints that motivate the game-theoretic architecture summarized below.
Industrial cybersecurity operates under strict physical and financial constraints. Unlike enterprise IT environments, where virtual machines and cloud workloads can be dynamically rebuilt, re-imaged, or isolated behind software-defined perimeters, operational technology consists of physical cyber-physical assets. Programmable Logic Controllers (PLCs), Remote Terminal Units (RTUs), and Safety Instrumented Systems (SIS) frequently run proprietary real-time operating systems (RTOS) on low-power microcontrollers incapable of supporting host-based intrusion detection software or transport layer security (TLS) handshakes.
Failure Modes of Heuristic and Matrix-Based Risk Scoring#
Currently, asset owners rely on qualitative risk matrices (e.g., standard likelihood-severity grids) or semi-quantitative scoring mechanisms such as the Common Vulnerability Scoring System (CVSS) to allocate defensive budgets. In industrial practice, these frameworks exhibit critical structural defects:
- Strategic Blindness: Qualitative scoring treats cyber threats as passive, environmental hazards analogous to lightning strikes or component wear-and-tear. In reality, advanced threat actors actively probe network perimeters, identify the least defended pathways, and dynamically shift targets when an asset owner deploys hardening measures.
- The "Curse of the Average": Averaging vulnerability scores across Purdue Model levels obscures catastrophic choke points. A low-severity vulnerability in an auxiliary engineering workstation can serve as a pivot point enabling an adversary to reach an unsegmented safety controller.
- Deterministic Payoff Fallacy: Existing security game formulations assume the defender knows the adversary's exact objective function. However, an attacker targeting an electrical transmission substation may seek ransom extortion (financial motivation), intellectual property theft (espionage motivation), or physical transformer core destruction (geopolitical sabotage). Assuming an incorrect attacker profile leads to catastrophic misallocation of defensive investments.
The Epistemic Stackelberg Paradigm#
To overcome these deficiencies, we formulate asset hardening as a leader-follower game. The defender acts as the leader, committing to a mixed strategy of hardening actions across network zones and conduits. The attacker acts as a follower, conducting reconnaissance, observing defender allocations through port scanning and traffic analysis, and executing their optimal attack vector.
Crucially, the defender faces epistemic uncertainty regarding the attacker's type. Rather than assuming a known, sharp probability distribution over attacker preferences, we apply imprecise probability theory. We bound the attacker distribution within a convex set of probability measures (a credal set ), ensuring that the resulting defensive posture is robust against worst-case misspecifications of adversary intent.
2. Mathematical Foundations & Physical Derivations#
Target Space and Strategy Formulations#
Let the cyber-physical system be represented as a set of discrete targets:
corresponding to operational assets across Purdue Levels 0 through 3 (e.g., SIS controllers, SCADA servers, historian databases, HMI nodes, and fieldbus protocol converters).
The defender possesses a finite security budget . Hardening target incurs an implementation cost . The defender's strategy is represented as a coverage probability vector:
where denotes the probability that target is protected by high-assurance defensive controls (e.g., bump-in-the-wire encryption, microsegmentation firewall rules, or dedicated honeypot monitoring). The set of feasible defender strategies is constrained by the capital budget:
Adversary Types and Payoff Tensors#
The adversary is drawn from a finite set of operational types:
Each type represents a distinct threat profile characterized by specific motives and technical sophistication:
- Type (Financial Extortionist / Ransomware Syndicate): Prioritizes Level 3 IT/OT boundary servers and historian databases to maximize operational downtime impact.
- Type (Strategic Cyber Saboteur / Nation-State APT): Prioritizes Level 1 PLCs and Level 0 safety instrumented systems to induce physical equipment damage.
- Type (Espionage Actor / Advanced Reconnaissance): Prioritizes engineering workstations and PLC logic source files for long-term telemetry extraction.
For each target and adversary type , we define four payoff parameters:
- : Defender reward if target is attacked while covered.
- : Defender cost (penalty) if target is attacked while uncovered ().
- : Attacker reward if target is attacked while uncovered.
- : Attacker cost (penalty) if target is attacked while covered ().
When the defender plays mixed strategy and attacker of type attacks target , the expected utilities are:
Modeling Epistemic Uncertainty via Credal Sets#
Rather than assuming a fixed prior probability distribution over attacker types, we define a credal set , where . We specify via lower and upper probability bounds derived from intelligence telemetry:
The defender optimizes against the worst-case probability distribution in the credal set, establishing a robust Strong Stackelberg Equilibrium (SSE) that minimizes maximum regret.
The Robust DOBSS Mixed-Integer Linear Program#
Under the Strong Stackelberg Equilibrium convention, if the follower is indifferent between multiple targets, they break ties in favor of the leader. Let binary variable denote whether attacker type attacks target . Since a rational attacker selects exactly one target:
To linearize the bilinear product of follower action and defender coverage, we define change of variables:
The complete robust DOBSS formulation is expressed as the following Mixed-Integer Linear Program (MILP):
subject to:
To enforce follower optimality, let represent the optimal expected utility of attacker type :
where is a sufficiently large positive scalar constant. The inner minimization over the credal set is dualized via linear programming duality, yielding a unified single-level MILP solvable via branch-and-cut algorithms in polynomial time for bounded target dimensions.
3. Empirical Benchmarks & Cyber-Physical Validation#
To validate the game-theoretic hardening framework, we conducted extensive evaluations on an empirical model of a large-scale industrial chemical synthesis plant consisting of critical automation nodes across Purdue Levels 0 to 3.
Experimental Configuration & Asset Inventory#
The testbed comprises:
- Purdue Level 3: 6 Enterprise/Historian Nodes (Active Directory, Historian DB, MES, Backup Gateway).
- Purdue Level 2: 10 Supervisory Workstations (HMI Terminals, Alarm Logging Servers, Engineering Workstations).
- Purdue Level 1: 16 Real-Time Controllers (Siemens S7-1500, Schneider Electric Modicon M580, Triconex Safety Instrumented Systems).
- Purdue Level 0: 16 Actuator/Sensor Interfaces (Flow controllers, pressure valves, emergency blowdown solenoids).
We calibrated attacker types across three categories:
- Type (Ransomware Extortionist): High reward for Level 3/2 nodes, zero interest in Level 0.
- Type (Nation-State Saboteur): Maximum reward for Level 1 SIS controllers and Level 0 blowdown valves.
- Type (Supply-Chain Competitor): Focuses on Level 2 engineering workstation configuration repositories.
The credal set over adversary types was specified as:
We benchmarked three allocation methodologies under identical budget constraints ( equivalent security allocation units):
- Methodology A (Heuristic Risk Matrix): Priority rank proportional to qualitative Likelihood Severity ratings.
- Methodology B (Deterministic Stackelberg Game): Standard Stackelberg solver assuming a uniform point distribution ().
- Methodology C (Eigenia Robust Bayesian Stackelberg): Robust DOBSS solver optimizing against the full credal set .
Quantitative Performance Comparison#
The empirical outcomes over 10,000 Monte Carlo adversarial campaign simulations are summarized below:
| Metric | Heuristic Risk Matrix (A) | Deterministic Stackelberg (B) | Robust Bayesian Stackelberg (C) |
|---|---|---|---|
| Defender Worst-Case Loss () | |||
| Attack Success Rate (Compromise) | |||
| Worst-Case ALE Reduction | Baseline () | ||
| 99% Value-at-Risk () | |||
| Solver Execution Time (48 Nodes) |
Analysis of Allocation Invariance and Regret#
Under Methodology A, the asset owner concentrated of the budget hardening the Level 3 Historian and HMI terminals because enterprise IT managers perceived them as possessing the largest attack surface. Adversary Type (the Saboteur) easily bypassed these hardened perimeters by exploiting an unmonitored serial-to-Ethernet bridge directly linked to Level 1 field controllers, resulting in an unmitigated physical loss event.
In contrast, our Robust Bayesian Stackelberg formulation allocated mixed coverage strategically:
- coverage on conduits connecting Level 2 HMIs to Level 1 Safety Systems (Triconex).
- coverage on Level 1 to Level 0 field instrumentation interfaces.
- coverage on Level 3 Enterprise connections.
By explicitly anticipating that the adversary optimizes their choice in response to observed hardening, and by hedging against epistemic uncertainty across attacker profiles, Methodology C prevented single-point failures and forced the attacker into low-yield, high-risk vectors.
4. Regulatory Mapping & Actuarial Solvency Integration#
Deploying formal game-theoretic security hardening transforms compliance from a subjective paperwork exercise into a mathematically verifiable, audit-proof defense posture.
European Regulatory Alignment#
- NIS2 Directive (Directive (EU) 2022/2555):
- Article 21(1) (Proportionality Principle): Mandates that essential and important entities implement risk management measures that are proportionate to the entity's exposure, taking into account the degree of the entity's exposure to risks and the societal impact of an incident. The robust BSSG framework provides the exact mathematical justification required by national regulatory authorities (e.g., ANSSI, BSI, NCSC), proving that defensive capital is deployed optimally under worst-case threat conditions.
- EU Cyber Resilience Act (CRA, Regulation 2024/2847):
- Article 10 & Annex I: Manufacturers and operators of critical industrial machinery must document a cybersecurity risk assessment reflecting adversarial capabilities. The credal set formulation directly satisfies requirements to account for varying adversary types and sophisticated state-backed threat actors.
- IEC 62443-3-2 (Security Risk Assessment for System Design):
- Prescribes the identification of zones, conduits, and Target Security Levels (SL-T 1 to 4). The mixed strategy coverage vector maps directly to Target Security Levels:
- Prescribes the identification of zones, conduits, and Target Security Levels (SL-T 1 to 4). The mixed strategy coverage vector maps directly to Target Security Levels:
Actuarial Solvency and Cyber Insurance Underwriting#
Commercial underwriters insuring critical industrial assets face severe accumulation risk from cascading cyber-physical incidents. Actuarial modeling defines the Value-at-Risk () over a one-year horizon as:
where is the annual aggregate cyber loss random variable, and is its cumulative distribution function.
When an industrial facility adopts heuristic risk ranking, the absence of strategic defense guarantees that tail events (e.g., simultaneous SIS lockout and runaway reaction) retain significant probability density, driving to catastrophic levels ( in our benchmark).
Under the Robust Bayesian Stackelberg allocation, the defender's minimax regret optimization guarantees an upper bound on expected tail loss:
Because drops by more than , insurers operating under EU Solvency II guidelines can formally reduce their Solvency Capital Requirement () for operational risk. Consequently, underwriters can grant insured asset owners verified premium reductions between and , transforming compliance investments into direct operational cost savings.
5. Conclusion & Implementation Roadmap#
Heuristic risk matrices and qualitative checklists are fundamentally incapable of securing modern industrial control infrastructure against rational, adaptive cyber adversaries. By grounding asset hardening in the mathematics of Bayesian Stackelberg Security Games and incorporating credal sets to account for epistemic uncertainty, asset owners can make provably optimal capital allocation decisions.
Phased Operational Deployment#
- Phase 1: Automated Asset & Conduit Graph Extraction: Ingest industrial engineering data (DEXPI P&ID schemas, network topology files, and CycloneDX 1.6 Hardware Bills of Materials) to generate the target set and estimate physical consequence losses .
- Phase 2: Adversary Profiling & Credal Set Calibration: Partner with threat intelligence teams to define adversary types , establish payoff tensors , and formulate the imprecise probability bounds .
- Phase 3: Robust MILP Execution & Zonal Enforcement: Execute the robust DOBSS optimizer within the enterprise cyber risk management platform, translating optimal coverage probabilities into enforceable IEC 62443-3-2 zone firewalls, hardware data diodes, and continuous threat monitoring priorities.
6. References#
- Tambe, M. (2011). Security and Game Theory: Algorithms, Deployed Systems, Lessons Learned. Cambridge University Press.
- Paruchuri, P., Pearce, J. P., Marecki, J., Tambe, M., Ordonez, F., & Kraus, S. (2008). Playing games for security: An efficient exact approach for solving Bayesian Stackelberg games. In Proceedings of the 7th International Joint Conference on Autonomous Agents and Multiagent Systems (AAMAS) (Vol. 2, pp. 895-902).
- Kiekintveld, C., Jain, M., Tsai, J., Pita, J., Ordonez, F., & Tambe, M. (2009). Computing optimal randomized resource allocations for massive security games. In Proceedings of the 8th International Conference on Autonomous Agents and Multiagent Systems (AAMAS) (Vol. 1, pp. 689-696).
- Walley, P. (1991). Statistical Reasoning with Imprecise Probabilities. Chapman and Hall.
- International Electrotechnical Commission. (2020). Security for industrial automation and control systems: Part 3-2: Security risk assessment for system design (IEC 62443-3-2:2020). IEC.
- European Parliament & Council. (2022). Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union (NIS2 Directive). Official Journal of the European Union.
- Bier, V. M., & Azaiez, M. N. (Eds.). (2009). Game Theoretic Risk Analysis of Security Threats. Springer Science & Business Media.