Algorithmic Mechanism Design & Bayesian Persuasion in Adversarial Incident Disclosure
J. McKenney
This is WG-07-TM-10, a standalone treatise in the WG-07-TM Threat Modeling working group rather than an entry in a numbered series, and it names no unpublished sibling.
Licence: CC BY 4.0. 17 September 2026.
Executive Abstract#
European disclosure law now runs on hard clocks. Article 14 of the Cyber Resilience Act (Regulation (EU) 2024/2847) and Article 23 of the NIS2 Directive (Directive (EU) 2022/2555) demand a machine-readable early warning within 24 hours of awareness of an actively exploited vulnerability, then a formal notification within 72 hours. Fed to ENISA and national CSIRTs, that early warning also leaks the primitives an attacker needs to build an exploit before asset owners have deployed operational technology patches.
It models the pipeline as an asymmetric-information signaling game, with manufacturer as Sender, CSIRT as Receiver, and adversary as eavesdropping interceptor. Bayesian persuasion over vulnerability severity yields a hybrid pooling-separating rule that satisfies CSIRT verification while holding the adversary's expected exploit return below economic viability.
The rule works only under a condition checked before deployment: a remote software fault that breaks nothing physical must be worth less to the attacker than building an attack, and a fault that can wreck a turbine or a breaker worth more. Where the two are valued alike, no policy of this shape deters anything, since the balance it strikes between them is its only instrument. The paper proves that impossibility and works the condition through for three attacker classes.
Zero-knowledge proof primitives carried in CycloneDX 1.6 Vulnerability Exploitability eXchange envelopes then let the entity attest remediation progress without revealing the memory offsets or control register vulnerabilities that would arm an adversary.
Abstract#
J. McKenney casts disclosure as an asymmetric-information signaling game (manufacturer Sender, CSIRT Receiver, adversary eavesdropping interceptor) and builds an optimal policy by Kamenica-Gentzkow concavification over the state space of vulnerability severity. A unique hybrid pooling-separating rule pi-star gives enough for statutory CSIRT verification while driving the adversary's posterior expected exploit return to zero, its mixing probability in closed form. The condition is sharp: the returns on the two non-benign states must be separated by the exploit development cost, V_1 < c_dev < V_2, and the prior must carry enough non-critical mass for the pooling weight to remain a probability. The converse holds. Where the two coincide, the mitigated-signal posterior surplus is independent of the mixing probability and equals the common return less the development cost, so no policy of this form deters an adversary for whom exploitation was worth attempting. The model assumes one signal shared by both receivers; distinct private signals would require the general information design framework, not the concavification used here. Zero-knowledge proof primitives in CycloneDX 1.6 Vulnerability Exploitability eXchange (VEX) envelopes provide verifiable guarantees of remediation progress without revealing actionable memory offsets or control register vulnerabilities.
1. Introduction and The Disclosure Paradox#
The signaling architecture this paper formalizes runs from the manufacturer's private knowledge of vulnerability state through an optimal persuasion policy to calibrated signals for regulators and adversaries alike, as follows.
The regulation of cybersecurity in critical infrastructure has entered an era of statutory enforcement. For decades, vulnerability disclosure operated under informal voluntary conventions, such as Coordinated Vulnerability Disclosure (CVD) and commercial bug bounty frameworks. However, the systemic risks highlighted by supply chain compromises, such as SolarWinds, Log4j, and the targeting of industrial programmable logic controllers (PLCs), have compelled regulators to mandate binding disclosure timelines. Under EU CRA Article 14, a manufacturer of products with digital elements must notify the designated CSIRT and ENISA within 24 hours of becoming aware of any actively exploited vulnerability or severe incident, with a comprehensive notification due at 72 hours.
This regulatory regime creates a profound engineering dilemma known as the Cyber-Physical Disclosure Paradox:
- The Regulatory Enforcement Mandate: Failure to report an actively exploited vulnerability within 24 hours exposes manufacturers to catastrophic administrative fines, up to €15,000,000 or 2.5 percent of total worldwide annual turnover, whichever is higher, as well as potential commercial exclusion from the European single market under CE mark revocation.
- The Weaponization Risk: Industrial OT equipment, including protection relays, distributed control systems (DCS), and supervisory control terminals, cannot be patched instantaneously. Testing an OT firmware patch requires physical test-bench validation, safety integrity level (SIL) re-certification under IEC 61508, and scheduled plant maintenance turnarounds. If an early warning report contains granular details, such as specific buffer bounds, vulnerable function signatures, or memory offsets, the disclosure acts as a catalytic weaponization signal for adversaries who have not yet developed an exploit.
Classical information security treats disclosure as a binary policy: either total secrecy (security through obscurity) or total transparency (full disclosure). In operational environments, both extremes lead to severe failure modes: total secrecy forfeits statutory compliance and blinds collective defense networks, while full disclosure arms the adversary during the multi-week patching gap.
To resolve this paradox, vulnerability reporting must be formalized as an algorithmic mechanism design problem under asymmetric information. By applying the economic principles of Bayesian Persuasion (Kamenica & Gentzkow), the disclosing entity can design an information structure that strategically shapes the beliefs of both the regulator and the adversary. The goal is to design a signaling policy that satisfies the statutory reporting threshold for CSIRT authorities while ensuring that the adversary's posterior belief regarding exploit feasibility remains economically unviable.
2. Game-Theoretic Formulation & Player Payoffs#
We model the incident disclosure environment as a three-player asymmetric information game with an informed Sender and two strategic Receivers.
1. The State Space of Vulnerability Severity#
Let be the discrete finite state space of vulnerability states:
where:
- represents a benign or non-exploitable anomaly (e.g., local denial of service requiring physical chassis access).
- represents a remotely exploitable logic flaw without immediate kinetic blast radius.
- represents a critical kinetic hazard capable of causing catastrophic physical damage (e.g., turbine overspeed, breaker trip override, pressure vessel overpressurization).
The common prior belief distribution across the state space is denoted by , where for all , and .
2. Players and Action Sets#
- Sender ( - Disclosing Manufacturer / Asset Owner): Privately observes the true state through internal automated telemetry, static code analysis, or incident triage. The Sender does not choose a message directly; rather, the Sender commits ex-ante to an information structure (signaling mechanism) , where is a measurable signal space.
- Institutional Receiver ( - CSIRT / ENISA Authority): Evaluates the received statutory signal to determine regulatory compliance and incident response posture:
- Strategic Adversary ( - Eavesdropping Threat Actor): Intercepts the public or leaked components of signal and chooses an operational effort:
3. Player Utility Functions#
The Sender's utility function captures regulatory penalties, physical damage losses, and operational costs:
where:
- , while , where is the statutory non-compliance penalty.
- is the physical kinetic loss resulting from an adversarial exploit, satisfying , , and .
- is the cost of developing, testing, and deploying the remediation firmware.
The CSIRT Authority () seeks to enforce transparency and protect European critical infrastructure:
The Adversary () has utility reflecting the expected financial or geopolitical return of weaponization minus the operational cost of exploit engineering:
Write for the adversary's return on a weaponized remote logic flaw, for the return on a weaponized kinetic hazard, and . These are two different numbers, and the state definitions of section 2.1 are the reason. A remotely exploitable logic flaw without kinetic blast radius buys persistence, data, or a nuisance outage. A critical kinetic hazard buys turbine overspeed, breaker trip override, or pressure vessel overpressurization, which is what an extortion demand or a geopolitical effect is priced on. A rational adversary does not pay the same for the two, and section 5 gives the figures for each actor class.
Crucially, the exploit development cost is high for industrial embedded architectures. Developing an exploit against an unknown memory layout or patched firmware requires substantial reverse engineering expenditure. That cost is the whole of the deterrent. Under any posterior the adversary's expected return from weaponization is:
and the adversary aborts wherever that quantity is zero or negative. The two state returns and the development cost must stand in the order:
for the Sender to have an instrument at all. A logic flaw alone does not repay the bench work; a kinetic hazard does; and everything the Sender can do is done by choosing the weight the signal places between the two. The theorem of section 3.3 carries this ordering as a hypothesis, and the proposition that follows the theorem shows that when it fails, no signaling policy deters anything.
3. The Bayesian Persuasion Mechanism#
Under the Bayesian persuasion framework of Kamenica and Gentzkow (2011), the Sender commits to a signaling scheme before observing the true realization of .
1. Information Structure and Belief Updating#
A signaling scheme consists of a signal space and a family of probability distributions over . Upon observing signal , each player updates their beliefs from the prior to a posterior belief via Bayes' rule:
A distribution of posterior beliefs is Bayes-plausible if and only if the expectation of the posteriors equals the prior:
By the revelation principle for Bayesian persuasion, it is without loss of generality to restrict attention to straightforward signaling schemes where the signal space consists of recommended actions or calibrated risk categorizations.
The signal is single and shared. The Sender commits to one scheme and emits one realized signal per disclosure event. The CycloneDX envelope lodged with the CSIRT and the public advisory read by the adversary are two renderings of that same . The envelope adds a cryptographic proof object and the advisory adds a status string, and neither carries a belief the other does not. Both receivers therefore hold the same posterior , the belief profile is degenerate on a common posterior, and the Sender's problem stays inside the single-sender, single-signal concavification of Kamenica and Gentzkow (2011): one Bayes-plausible distribution over posteriors, with both receivers' best responses folded into the scalar objective of section 3.2. This is a property of the mechanism rather than a modeling convenience, and section 4 is what enforces it: the statutory rendering adds verifiability, not information.
A mechanism that instead sent one private signal to the CSIRT and a different private signal to the adversary would induce two distinct posteriors, and the concave closure of a scalar indirect utility would no longer be the right object. Such a design belongs to the general information design framework of Bergemann and Morris (2019), which admits an arbitrary set of receivers and arbitrary correlation across their signals, and which contains the construction used here as its single-public-signal case. That framework is cited as the tool that would be required if the shared-signal property were given up. It is not the tool used below, and nothing in sections 3.2 or 3.3 relies on it.
2. The Concavification Problem#
For any posterior belief , let:
- be the regulator's optimal decision.
- be the adversary's optimal decision.
We define the Sender's indirect utility function over posterior beliefs:
The Sender's optimal persuasion mechanism corresponds to finding the Bayes-plausible distribution of posteriors that maximizes the expected indirect utility. Geometrically, this maximum equals the concave closure (concavification) of evaluated at the prior :
3. Derivation of the Optimal Hybrid Signaling Scheme#
Because exhibits discontinuous jumps at the thresholds where the regulator shifts from to and where the adversary shifts from to , the function is non-concave.
Let be the set of posterior beliefs that satisfy the statutory standard under CRA Article 14 (i.e., where CSIRT accepts the 24-hour notification without initiating formal non-conformance audits). Let be the set of posteriors where the adversary chooses .
Theorem 1#
Structure of the Optimal Hybrid Disclosure Scheme
Write for the prior mass on state . Assume:
(H1) The prior belief satisfies , the baseline where uncalibrated disclosure either incurs regulatory sanction or invites adversarial exploitation.
(H2) The statutory verifiability threshold is met:
(H3) Severity separation. The adversary's two non-benign returns are distinct and the exploit development cost separates them:
(H4) Feasibility. The prior carries enough non-critical mass to pool against:
Then the optimal signaling policy is a two-point hybrid distribution supported on exactly two signals :
where the mixing probability pools the non-critical state with the critical state just up to the boundary of the adversary's economic indifference:
or equivalently, with the sign of the denominator made explicit by (H3):
Hypotheses (H3) and (H4) are exactly the statement that . (H3) gives and fixes the sign of the denominator; (H4) gives .
Proof of Theorem 1#
The nominal branch. Under the posterior assigns zero mass to the kinetic hazard state, since . Bayes' rule gives:
so the adversary's expected return is . Because and by (H3), that quantity is strictly negative and . Without (H3) the bound fails: if exceeded , a nominal signal with enough weight on would be worth exploiting on its own.
The mitigated branch. Under the Sender releases a machine-readable attestation confirming the active deployment of a protective mitigation interlock, such as a zero-knowledge proof of a compensatory firewall rule or a firmware sandbox, without revealing the memory offset. Bayes' rule, conditioning the state on the signal, gives the two non-zero posterior components:
and , since . Both components enter the adversary's expectation and both must be carried through the algebra; solving on one of them alone does not determine .
Write for the mass of the logic-flaw state routed to the mitigated signal, so that the unconditional probability of that signal is . The zero-surplus condition is:
Multiplying through by clears the denominators:
Collecting the terms in on one side and the terms in on the other:
By (H3) the factor is strictly negative and is strictly positive, so both sides are negative and the division is legitimate:
Substituting and solving for :
which is the expression in the statement. The subtracted term is strictly positive, giving , and (H4) is precisely the condition for that subtracted term not to exceed one, giving . The mixing probability is therefore admissible, and at the adversary's posterior surplus under is exactly zero. Under the tie-breaking convention that an indifferent adversary does not spend the development budget, the equilibrium action is under both signals.
The statutory branch. Substituting the optimum into the ratio of (H2), and using :
so (H2) is an operational test on the three adversary parameters alone: the scheme clears the statutory gate when . Because also carries cryptographically verifiable proof of incident existence and containment, the CSIRT authority verifies compliance under Article 14(2) and selects .
Optimality. Any other Bayes-plausible distribution of posteriors either places positive mass on a posterior at which the adversary exploits, incurring with positive probability, or places positive mass on a posterior failing the statutory ratio, incurring with positive probability. Both strictly reduce , and the two-point scheme above attains the concave closure at . This proves optimality.
Proposition 1: Severity separation is necessary#
If the two non-benign returns coincide, , and the paper's operating premise holds, then no signaling policy of the form of Theorem 1 deters the adversary, for any mixing probability .
Proof. By construction , so the benign state carries zero mass in the mitigated posterior and the remaining two components sum to one:
With a common value the adversary's posterior surplus collapses:
The mixing probability has vanished from the expression. The surplus is whatever the Sender chooses, so the adversary exploits on every realization of , and the Sender's only remaining option is to send that signal with probability zero, which is unavailable because and by the positive-prior axiom of section 2.1. Hypothesis (H3) is therefore not a regularity condition added for convenience. It is the boundary between a solvable problem and an unsolvable one, and it has to be checked against the adversary's own valuations before the scheme is deployed.
4. Zero-Knowledge Attestation Architecture for CycloneDX 1.6#
To execute the optimal signaling policy without leaking exploitable vulnerability details, we bind the Bayesian persuasion mechanism to an arithmetic zero-knowledge SNARK (zk-SNARK) circuit.
1. The Information Leakage Bound#
In industrial firmware vulnerability disclosures, information leakage is quantified by the mutual information between the emitted signal and the secret vulnerability parameters (such as stack buffer bounds, memory addresses, or unauthenticated register numbers):
To prevent exploit synthesis, the signaling scheme must enforce the zero-knowledge constraint:
2. Zero-Knowledge Circuit Construction#
We formulate an arithmetic circuit over the scalar field of the BN254 elliptic curve. The circuit verifies the following statement:
where:
- Public Inputs (): The public CVE identifier, the SHA-384 hash of the affected firmware binary, the statutory timestamp, and the cryptographic commitment to the mitigation state.
- Private Witness (): The exact memory address of the vulnerability, the vulnerability type, the exploit preconditions, and the proprietary remediation patch diff.
The resulting Groth16 proof is precisely in length. This proof is embedded directly into the declarations block of a standard CycloneDX 1.6 Vulnerability Exploitability eXchange (VEX) JSON document, as specified in the schema below:
{
"bomFormat": "CycloneDX",
"specVersion": "1.6",
"serialNumber": "urn:uuid:8b3e8e24-9b51-4f11-9a72-6a4a0c8b3211",
"version": 1,
"components": [
{
"type": "firmware",
"bom-ref": "eigenia-firmware-substation-relay-g4-2.4.1",
"group": "eigenia",
"name": "substation-relay-g4",
"version": "2.4.1"
}
],
"vulnerabilities": [
{
"id": "CVE-2026-4921",
"source": { "name": "NVD", "url": "https://nvd.nist.gov/vuln/detail/CVE-2026-4921" },
"analysis": {
"state": "in_triage",
"justification": "protected_by_mitigating_control",
"response": ["workaround_available"],
"detail": "Automated Bayesian Persuasion Mitigation Envelope v1.2"
},
"affects": [
{
"ref": "eigenia-firmware-substation-relay-g4-2.4.1",
"versions": [
{ "version": "2.4.1", "status": "affected" }
]
}
],
"properties": [
{
"name": "eigenia:cra:article14:proof_scheme",
"value": "Groth16-BN254"
},
{
"name": "eigenia:cra:article14:zk_proof",
"value": "0x1a8f...39bc"
},
{
"name": "eigenia:cra:article14:statutory_timestamp",
"value": "2026-09-14T08:12:00Z"
}
]
}
]
}5. Empirical Simulation & Benchmarking#
The algorithmic disclosure mechanism was evaluated across a simulated corpus of modeled on historical industrial control system advisories (ICS-CERT and CISA KEV catalogs between 2021 and 2026).
1. Evaluation Setup#
Each cycle simulated:
- True state distribution: (low/moderate), (remote logic flaw), (critical kinetic overpressure or electrical trip hazard).
- Adversarial profiling: three actor classes, each carrying its own exploit development cost and its own pair of state returns, tabulated below.
- Baseline comparison:
- Full Immediate Disclosure: Immediate publication of complete vulnerability details at .
- Uncalibrated Minimal Reporting: Vague text notifications without cryptographic proof at .
- Bayesian Persuasion Mechanism (): Proposed hybrid signaling with zero-knowledge verification.
Each actor class is priced on the two non-benign states separately, following the state definitions of section 2.1. The remote logic flaw is worth what persistence, data theft, or a nuisance outage on a substation relay is worth to that class of actor. The critical kinetic hazard is worth what a credible threat of turbine overspeed or breaker trip override is worth, which for every class is a different and larger number. The development cost is the bench cost of reaching a working exploit against proprietary embedded firmware: acquiring the hardware on the secondary market, standing up a test rig, and reverse engineering the target.
| Adversary class | (logic flaw, ) | (kinetic hazard, ) | |
|---|---|---|---|
| Script kiddie or commodity operator | |||
| Commercial cybercriminal or extortion crew | |||
| Advanced persistent threat |
Before the deterrence claim can be made for any class, hypotheses (H3) and (H4) of Theorem 1 have to hold for that class, and the resulting has to lie in . With the simulated prior , , , the check runs as follows. Write and , so that .
| Adversary class | (H3) | (H4) | ||||
|---|---|---|---|---|---|---|
| Script kiddie or commodity operator | holds | yes | ||||
| Commercial cybercriminal or extortion crew | holds | yes | ||||
| Advanced persistent threat | holds | yes |
Substituting each back into the mitigated posterior confirms that the surplus is driven to zero rather than merely reduced, which is the quantity the theorem claims and the only quantity that supports the abort conclusion.
| Adversary class | Posterior surplus | Nominal surplus | |||
|---|---|---|---|---|---|
| Script kiddie or commodity operator | |||||
| Commercial cybercriminal or extortion crew | |||||
| Advanced persistent threat |
The surplus is zero under and strictly negative under for all three classes, so under the tie-breaking convention of Theorem 1 the adversary aborts under either realization of the signal, for every class, at the tabulated valuations. The statutory side of the check is , which evaluates to , and for the three classes in order, so a CSIRT tolerance of clears all three.
Had the two non-benign states been priced alike within a class, , Proposition 1 would apply instead: the posterior surplus under would equal for every mixing probability, the deterrence claim would be unavailable at any , and the row would have to be struck rather than recomputed.
2. Empirical Results#
| Metric / Disclosure Policy | Full Immediate Disclosure | Uncalibrated Minimal Reporting | Bayesian Persuasion () |
|---|---|---|---|
| Statutory CRA Art. 14 Compliance | () | (Rejected by CSIRT) | () |
| CSIRT Formal Audit Inquiries | (High friction) | (Automated ZK Clearance) | |
| Mean Time to First Exploit Attempt | No Exploits Attempted () | ||
| Opportunistic Exploit Attempts | (99.5% reduction) | ||
| Mean Operator Financial Loss | € (Kinetic damage) | € (Regulatory fines) | € (Triage/patch testing only) |
| ZKP Proof Generation Overhead | N/A | N/A | (Groth16 on CPU) |
| CSIRT Verification Runtime | N/A | (Manual) | (Automated REST API) |
The experimental data confirms the theoretical predictions:
- Full immediate disclosure guarantees compliance with regulatory deadlines but unleashes a wave of opportunistic exploits (), resulting in extensive physical equipment damage before asset owners can schedule maintenance outages.
- Uncalibrated minimal reporting shields operators from exploit weaponization but fails statutory CSIRT verification in of cases, triggering extensive regulatory scrutiny and threat of administrative penalties under CRA Article 14.
- The Bayesian persuasion mechanism with zero-knowledge attestation achieves statutory compliance within the mandatory 24-hour window, eliminates manual CSIRT audit friction through cryptographic verification, and reduces adversarial exploit attempts by .
6. Regulatory Synthesis & Statutory Timelines#
Deploying algorithmic mechanism design provides critical infrastructure operators with a legally defensible blueprint under European Union statutory timelines:
- Phase 1: 0 to 24 Hours (The Early Warning): The operator detects the anomaly, computes the optimal signal , generates the Groth16 zero-knowledge proof , and transmits the CycloneDX 1.6 VEX envelope to ENISA and the designated national CSIRT. This satisfies the strict 24-hour statutory requirement under CRA Article 14(2) without publishing actionable exploit primitives.
- Phase 2: 24 to 72 Hours (The Incident Notification): The operator provides updated mitigation telemetry, verifying that physical interlocks, network micro-segmentation, or firewall virtual patching rules are active across production assets.
- Phase 3: Final Comprehensive Report (Within 14 Days of Patch): Once the verified vendor firmware patch has undergone full hardware-in-the-loop (HIL) safety testing and field deployment, the complete technical documentation is finalized and archived.
7. Conclusion#
Vulnerability disclosure in cyber-physical systems cannot remain an uncoordinated, binary choice between total secrecy and reckless publication. By framing statutory reporting as a Bayesian persuasion mechanism, this treatise proves that information asymmetry can be strategically engineered to satisfy rigorous regulatory oversight while systematically starving adversaries of actionable weaponization intelligence. The proof is conditional, and the condition is where the engineering judgment sits: the mechanism has an instrument only where the exploit development cost separates what the adversary will pay for a remote logic flaw from what it will pay for a kinetic hazard. Where an adversary prices those two alike, Proposition 1 rules the construction out entirely, and the operator's answer has to come from somewhere other than disclosure design. Integrating zero-knowledge proof primitives directly into CycloneDX 1.6 VEX envelopes provides the mathematical bridge between legal accountability and physical infrastructure protection.
8. References#
- Kamenica, E., & Gentzkow, M. (2011). Bayesian Persuasion. American Economic Review, 101(6), 2590-2615.
- McKenney, J. (2026). Asymmetric Information and Signaling Equilibria in Critical Infrastructure Vulnerability Disclosures. Eigenia Research Technical Reports, WG-07-TM-10.
- European Parliament and Council. (2024). Regulation (EU) 2024/2847 on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act). Official Journal of the European Union.
- European Parliament and Council. (2022). Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union (NIS2 Directive). Official Journal of the European Union.
- Groth, J. (2016). On the size of pairing-based non-interactive arguments. Advances in Cryptology, EUROCRYPT 2016, 305-326.
- OWASP Foundation. (2024). CycloneDX v1.6 Specification: Vulnerability Exploitability eXchange (VEX).
- Bergemann, D., & Morris, S. (2019). Information design: A unified perspective. Journal of Economic Literature, 57(1), 44-95.