Frontier AI Hardware Security & Platform Assurance Framework
J. McKenney
This paper sits in Eigenia working group WG-05-CAD, Unified Asset Graph, whose central work is joining process topology, component inventory and an electrical model into one computable graph. It is the group's frontier compute case. It takes its bill of materials layer definition from the unified DEXPI 2.0 and CycloneDX 1.6+ semantic bridge, where the layer count is stated once for the corpus, and applies the group's apparatus to an asset class the group's other reference architectures do not cover: an accelerated compute rack treated as a plant under hostile facility conditions. Where the battery storage, pharmaceutical and rail depot cases test the three-schema join itself, this paper carries the argument through to a loss calculation, which those papers explicitly decline to attempt.
Licence: CC BY 4.0. 3 August 2026.
Executive Abstract#
A cluster training a frontier model is not an office data centre with larger machines. It draws more than a hundred megawatts, the weights inside represent an investment above a billion dollars, and the parties who want them include states. This paper asks what must be true of the hardware, and the building around it, before it can be defended.
Its first move is to stop treating the building as safe. Chilled water, busbars, the building management system and the facility network become an untrusted environment pressing on the rack from outside, the facility threat model. What that pressure bears on is the rack envelope: four nested layers from a tamper-evident enclosure, through the facility-service boundary, past the host and its management controller, to the accelerator silicon. The second, harder move tells the accelerator to distrust its own host: a security processor on the die refuses any unsigned kernel, and external memory is encrypted at line rate, so a probe or cold-boot extraction recovers noise.
The named threats are mostly physical: a management controller turned into a covert probe or made to report false voltages, remote memory reads that bypass the host operating system, weights hidden in the low bits of performance counters, cooling and power as attack surfaces. The response joins industrial control security, railway RAMS and hardware root of trust, and trades human-speed vulnerability response, outpaced by machine-speed attack synthesis, for a continuous falsification loop against a hardware-in-the-loop emulation. A closing section prices it: a single loss expectancy, an annualized loss for a rack row, a return on investment and a Gordon-Loeb check. The return reproduces exactly, but that is a property of the division, not evidence for the five assumed inputs behind it.
Abstract#
Frontier AI training and inference clusters have crossed from enterprise data processing into sovereign-tier physical assets. At cluster scales above 100 MW, the parameter weights of advanced models represent capital above one billion dollars and draw nation-state advanced persistent threats, industrial espionage and state-sponsored physical interdiction. Conventional datacenter security rests on a layered perimeter that presumes the interior, its power distribution, chilled water, building management systems and host execution environments, to be trusted territory; in high-consequence environments that assumption is invalid. This specification establishes the AI Rack Envelope, a self-contained zero-trust cryptographic, thermodynamic and physical boundary for accelerated compute under hostile facility conditions, and formalizes the Facility Threat Model, treating the physical plant as an active multi-vector attack environment exerting thermodynamic, electrical, out-of-band, sideband and physical pressure across the rack perimeter.
1. Foundational Scope#
To bridge the historical chasm between semiconductor microarchitecture, server chassis firmware, civil infrastructure, and financial underwriting, this treatise synthesizes industrial cybersecurity standards (IEC 62443 Security Levels SL-1 through SL-4) with railway and nuclear functional safety methodologies (CLC/TS 50701 and EN 50126 RAMS: Reliability, Availability, Maintainability, and Safety) and actuarial loss accumulation models under the state-backed cyber-attack exclusion that Lloyd's Market Bulletin Y5381 requires. We demonstrate that in high-density accelerator facilities operating between 100 kW and 140 kW per rack, cybersecurity, physical functional safety, and actuarial solvency converge: cyber interdictions of physical cooling or power controls induce catastrophic silicon failure within seconds, while physical perturbations can be applied to bypass cryptographic boundaries. This standard provides the engineering mechanisms; encompassing open silicon root-of-trust engines, line-rate bus encryption, machine-speed automated testbenches, and four-dimensional bills of materials; required to guarantee platform integrity, model weight confidentiality, and insured asset survivability.
| Envelope layer | Constituent controls |
|---|---|
| 1. Physical & Environmental Boundary | Tamper-Evident Enclosure, Intrusion Interlocks, Sealed Cold Plates |
| 2. Facility Conduit Defense (IEC 62443 SL-4 Zone Boundary) | Micro-CDU Flow Meters, Optical Fiber Air-Gaps, Dual Power Filters |
| 3. Host Node & Out-of-Band Management (Distrusted Tier) | Host CPU, Hypervisor, Linux OS, Baseboard Management Controller |
| 4. Silicon Root of Trust & Cryptographic Accelerator Enclave | Caliptra RoT, Locked Accelerator Kernels, Isolated HBM Memory; PQC Engine (CNSA 2.0 / ML-DSA-87), Hardware Egress Throttling |
2. The AI Rack Envelope Architecture & Trust Boundaries#
The AI Rack Envelope represents a fundamental conceptual shift in mission-critical systems engineering. In standard hyperscale architectures, security controls are distributed across host operating systems, hypervisors, orchestration layers, and network switches. The AI Rack Envelope consolidates the entire physical and logical compute complex; comprising compute sleds, accelerator trays, high-speed switching fabrics, power delivery stages, and manifold piping; into a single, unified, verifiable fortress.
2.1 Physical and Logical Composition#
Modern frontier training clusters aggregate high-performance computing elements within extreme volumetric density. A single AI Rack Envelope typically houses between eight and sixteen specialized compute trays, each containing eight interconnected accelerator modules (such as SXM or OAM form factors), dual high-performance host processors, multi-terabit PCIe Gen 5 or Gen 6 switching complexes, and high-speed network interface cards (NICs) supporting 800 Gbps to 1.6 Tbps fabric connectivity.
The physical boundary of the envelope includes:
- Mechanical Chassis Integrity: A reinforced, tamper-evident rack framework equipped with continuous physical intrusion detection switches, optical fiber loop continuity sensors, and electromagnetic interference (EMI) shielding designed to mitigate TEMPEST sideband emissions.
- Integrated Fluid Containment: Hermetically welded manifold distribution loops equipped with localized isolation valves, pressure drop transducers, and optical moisture sensing strips positioned at all disconnect junctions.
- Internal Power Regulation: A dual-feed 48V or 400V DC busbar backplane coupled directly to in-rack capacitor banks and point-of-load voltage regulator modules (VRMs), providing local energy buffering to absorb high-frequency transient steps.
The logical boundary establishes strict cryptographic segregation between compute resources, management controllers, and external network links, ensuring that no unverified instruction can reach execution units.
2.2 Silicon-Enforced Zero Trust: Distrusting the Host Complex#
The foundational security postulate of the AI Rack Envelope is that accelerator silicon must strictly distrust the host system.
In conventional accelerated computing, the host CPU acts as the primary master of the domain: it initializes platform memory, loads device drivers, schedules execution kernels, manages physical DMA mappings over PCIe, and supervises system state via out-of-band management engines. This design creates an unacceptable vulnerability profile. If an adversary compromises the host operating system kernel, gains hypervisor escape, or infiltrates the Baseboard Management Controller (BMC), they gain direct, unconstrained access to inspect, manipulate, or dump the contents of high-bandwidth memory (HBM) attached to the accelerators.
Under this framework, the accelerator architecture operates as an autonomous, self-sovereign cryptographic enclave:
- Independent Execution Verifier: The accelerator silicon incorporates an isolated, on-die security processor that validates all incoming compute kernels prior to execution. If a host driver attempts to dispatch unsigned, tampered, or arbitrary code to the accelerator execution pipelines, the security processor halts the dispatch queue and isolates the physical interface.
- DMA Ring-Fencing: Host memory controllers and PCIe root complexes are barred from initiating direct, unencrypted reads of accelerator physical memory. All memory transfers between the host and accelerator occur through dedicated cryptographic bounce buffers governed by hardware access-control registers that enforce strict unidirectional semantics.
- Hardware-Enforced Memory Scrambling: Dedicated AES-256-XTS or post-quantum cryptographic engines encrypt all data written to external HBM and DDR5 channels at line rate, ensuring that physical interposers, inter-die bridges, or cold-boot physical extraction yield only cryptographically randomized ciphertext.
2.3 The Four-Point Cryptographic Envelope#
To ensure end-to-end data confidentiality, the AI Rack Envelope establishes line-rate, hardware-enforced cryptographic boundaries across four distinct architectural interfaces:
| Boundary Interface | Physical Interconnect | Cryptographic Standard | Threat Mitigated |
|---|---|---|---|
| Point 1: Scale-Up Coherent Fabric | Proprietary Accelerator Crossbars (e.g., NVLink, xGMI) | Line-Rate IDE (Integrity and Data Encryption) / AES-GCM-256 | Physical probing of inter-tray cables, interposer tapping, and coherent memory eavesdropping. |
| Point 2: Host-to-Device Bus | PCIe Gen 5 / Gen 6 over CEM and OAM slots | PCIe IDE / SPDM 1.3 Key Exchange | Malicious host DMA injection, bus protocol manipulation, and interposer analyzer attacks. |
| Point 3: Non-Volatile Storage Bus | NVMe-oF / CXL 2.0 / 3.0 Direct Attach | Self-Encrypting Drive (SED) Opal 2.0 / IEEE 1619 | Checkpoint theft, state-dump extraction from discarded drives, and cold-standby tampering. |
| Point 4: Scale-Out Network Fabric | 800G / 1.6T RDMA over Converged Ethernet (RoCEv2) / InfiniBand | IPsec Line-Rate MACsec (IEEE 802.1AE) / PSP (Packet Security Protocol) | In-transit weight extraction across inter-rack spine-leaf networks and optical tap interdiction. |
This multi-point cryptographic enclosure ensures that even if an adversary achieves complete physical tap access to copper DAC cables, optical transceivers, or backplane bus lines within the datacenter, all intercepted payloads remain cryptographically unassailable.
3. The Facility Threat Model: Pressure on the Compute Envelope#
Conventional cybersecurity frameworks model threat actors operating almost exclusively through network protocols, operating system exploits, and software vulnerabilities. In an AI datacenter facility, this view is dangerously deficient. The physical infrastructure supporting an AI compute cluster; including electrical power substations, medium-voltage switchgear, uninterrupted power supply (UPS) systems, 48V power distribution units, primary chilled water loops, and secondary Coolant Distribution Units (CDUs); represents an attack surface of immense consequence.
The Facility Threat Model treats the physical plant as an untrusted domain that exerts continuous multi-domain pressure across the AI Rack Envelope.
Facility Threat Interface Matrix
| Facility Pressure Vector | Mechanism | AI Rack Envelope Target Boundary |
|---|---|---|
| Thermodynamic Pressure | CDU pump drop, heat shock | Liquid Cold Plates, Dielectric Cavitation, Thermal Throttling |
| Electrical Transient Vector | di/dt steps, grid droop | 48V Busbars, Point-of-Load VRMs, Hardware Clocking |
| Out-of-Band Sideband Vector | IPMI, Redfish, JTAG tap | BMC Firmware, I2C/I3C Buses, SPI Boot Flash, OpenBIC |
| Interconnect Egress Vector | Covert sideband streaming | RDMA Fabric Transceivers, Telemetry Streaming Channels |
3.1 Thermodynamic & Liquid Cooling Pathways#
Operating at 100 kW to 140 kW per rack, modern accelerator enclosures require direct-to-chip liquid cooling to evacuate heat fluxes of about across the accelerator package. Liquid cooling systems operate via primary facility water loops coupled through plate heat exchangers to secondary, high-purity treated water or Propylene Glycol 25 percent (PG25) dielectric loops within the row-level Coolant Distribution Units (CDUs). Secondary loops circulate coolant at volumetric flow rates between 60 L/min and 180 L/min per rack frame, depending on facility water temperature, under nominal working pressures between 2.5 bar and 4.0 bar.
This hydraulic architecture introduces physical attack pathways:
The 15-Second Thermal Runaway Window#
Traditional air-cooled datacenters possess substantial thermal capacitance: the vast volume of cold-aisle air provides several minutes of buffer time during cooling plant failures. In contrast, liquid-cooled cold plates contain minute fluid volumes (often less than 250 mL per cold plate). If an adversary breaches the facility BMS network and maliciously throttles CDU variable-frequency drive (VFD) pumps, commands automated balance valves to close, or disables secondary booster pumps, the thermal inertia of the cold plate is exhausted within seconds.
At 1,200 watts of thermal dissipation per accelerator package, the time required for junction temperature to escalate from its nominal operating state of to the emergency hardware shutdown trip point is governed by:
Where is the thermal mass of the copper cold plate, is the specific heat capacity, and represents residual heat evacuation. Under zero-flow conditions (), and for a stagnant cold plate assembly of , is 14.8 seconds.
An adversary manipulating facility cooling can systematically trigger emergency thermal shutdown across an entire cluster, creating massive availability denial during mission-critical inference or multi-month training checkpoints.
Fluid Pressure Surges & Induced Cavitation#
Manipulating flow control valves via hijacked Modbus/BACnet protocols can generate hydraulic water-hammer events where pressure oscillations exceed 10 bar. Pressure shockwaves propagating through manifold piping rupture quick-disconnect couplings, spraying conductive coolant across high-amperage 48V electrical busbars and causing immediate explosive arcing and localized fire events.
Thermal Sideband and Fault-Injection Attacks#
By subtly modulating coolant flow rates within safe operational limits, an adversary can deliberately shift die temperatures up and down, inducing predictable thermal drift in on-die ring oscillators. This provides a mechanism for remote thermal clock glitching or differential power-thermal analysis to recover cryptographic keys from on-die cryptoprocessors.
3.2 Electrical & Power Distribution Pathways#
Power distribution architectures within frontier AI clusters deploy 48V DC power distribution bars delivering up to 3,000 amperes per rack frame. Accelerators execute dynamic matrix arithmetic workloads that induce immense step changes in current demand: when an all-reduce collective communication phase transitions into an intensive GEMM computation, the current draw of an accelerator tray swings from 20 percent to 100 percent of maximum rating within tens of nanoseconds.
This high dynamic range presents an electrical attack vector:
Resonant Frequency Injection ( Stepping)#
An adversary who controls unprivileged workload execution can craft adversarial neural network execution graphs designed to toggle compute units at the precise natural electrical resonance frequency of the power distribution network (PDN). By alternating between high-intensity tensor instructions and complete pipeline stalls at kilohertz frequencies, the adversary induces catastrophic resonant voltage fluctuations:
Where represents the parasitic inductance of the busbar and power cabling. If the resonant frequency is excited, the induced voltage oscillation exceeds the dielectric breakdown threshold of input filter capacitors, causing physical hardware destruction.
Targeted Voltage Droop and Power Glitching#
Coordinated power spikes can pull local voltage below the minimum operating threshold () of accelerator voltage regulator modules (VRMs). This induced voltage droop causes timing faults in logic synthesis pathways, flipping bits in cryptographic signature verifications or instruction decode registers; a remote, software-induced physical fault injection attack that bypasses secure boot integrity checks.
3.3 Out-of-Band Management & Sideband Surfaces#
Baseboard Management Controllers (BMCs) represent an architectural vulnerability in modern enterprise computing. Operating on isolated system-on-chip (SoC) architectures running embedded Linux (e.g., OpenBMC), the BMC retains direct electrical access to host board pins, power sequencing rails, SPI flash chips, I2C/I3C management buses, and PCIe sideband interfaces.
The facility network frequently interfaces directly with BMCs via IPMI, Redfish, or SNMP protocols:
- Firmware Flashing Vulnerabilities: Compromising the BMC allows an attacker to overwrite host BIOS or accelerator microcode stored in SPI NOR flash chips, establishing persistent, undetectable firmware rootkits that survive full operating system reinstalls.
- Inter-Integrated Circuit (I2C/I3C) Bus Sniffing: Platform telemetry (temperatures, voltages, clock states) and cryptographic key initialization sequences traverse low-speed board-level buses. An unauthenticated BMC can passively eavesdrop on these traces or inject falsified voltage reports that trigger premature emergency shutdowns.
- PCIe Sideband Exploitation (MCTP / JTAG): Management Component Transport Protocol (MCTP) over PCIe or SMBus enables BMCs to inspect internal device state registers. If sideband access is not cryptographically locked in silicon, the BMC can be subverted into a covert hardware probe to read out memory pointers and sensitive weights.
3.4 Scale-Out Fabric Egress & Covert Weight Exfiltration#
Frontier model weights are distributed across thousands of accelerators using high-bandwidth RDMA networks. To prevent weight exfiltration, traditional defenses inspect host network egress interfaces. However, high-speed scale-out fabrics feature dedicated network engines that bypass host CPUs entirely:
- Covert RDMA Exfiltration: An adversary with root execution on an individual node can schedule low-priority, asynchronous RDMA Read operations targeting remote accelerator HBM over unoccupied fabric virtual lanes. Because RDMA operations bypass kernel protocol stacks and OS logging, telemetry monitors frequently fail to detect the exfiltration.
- Hardware Telemetry Sidebands: Accelerators export hardware diagnostic telemetry over dedicated fabric channels. If these telemetry streams are unconstrained, an adversary can encode model weight tensors into the low-order bits of diagnostic performance counters, streaming intellectual property out of the cluster beneath the noise floor of standard monitoring tools.
4. Systems Assurance Convergence#
Safety, Reliability & Industrial Cybersecurity
Mitigating the facility threat model requires ending the separation between information technology (IT) cybersecurity and operational technology (OT) functional safety. In high-density AI infrastructure, cyber threats directly induce physical equipment failure, and physical plant disturbances compromise cryptographic boundaries.
To construct a defensible architecture, this framework integrates IEC 62443 (Security for Industrial Automation and Control Systems) with CLC/TS 50701 and EN 50126 (Railway and Industrial Safety RAMS Engineering).
IEC 62443 zones and conduits architecture matrix. Conduit SL ratings are the enforced boundary levels; zone target levels are stated in section 4.1.
4.1 IEC 62443 Zones and Conduits Applied to AI Facilities#
IEC 62443 establishes a formal methodology for segmenting complex physical environments into discrete Zones possessing shared security requirements, interconnected via monitored, access-controlled Conduits.
We establish four mandatory operational zones:
Zone 1#
Facility Physical Plant & Industrial Automation (Target: SL-2)
- Scope: Central chilled water plants, cooling towers, primary pumps, utility transformers, medium-voltage switchgear, and centralized BMS/SCADA servers.
- Security Posture: Physical access control, isolated VLANs, mutual TLS for Modbus/BACnet telemetry, and read-only monitoring interfaces to external networks.
Zone 2#
Rack-Level Out-of-Band Management (Target: SL-3)
- Scope: Rack management controllers, chassis BMCs, OpenBIC microcontrollers, smart ePDUs, and manifold flow monitoring nodes.
- Security Posture: Hardware-authenticated firmware signing, disabled legacy IPMI protocols, mandatory Redfish HTTPS authentication with short-lived certificates, and strict physical separation from compute data networks.
Zone 3#
Host Compute & Virtualization Layer (Target: SL-3)
- Scope: Host x86/ARM CPUs, system DDR5 memory, Linux kernel, Kubernetes container runtimes, and local PCIe root complexes.
- Security Posture: Measured boot via TPM 2.0, immutable root filesystems, kernel lockdown enforcement, and hardware-enforced hypervisor isolation.
Zone 4#
The Silicon Cryptographic Compute Enclave (Target: SL-4)
- Scope: Accelerator execution dies, on-package High-Bandwidth Memory (HBM3e/HBM4), on-die SRAM, and coherent scale-up crossbars.
- Security Posture: The highest protection tier defined by IEC 62443-3-3 [1]. Protection against sovereign-tier adversaries armed with physical laboratory equipment. Silicon root-of-trust enforcement, cryptographic bus encryption, line-rate memory scrambling, and autonomous kernel verification.
Conduit Enforcement Rules#
All conduits crossing zone boundaries must enforce deterministic security policies:
- Conduit A (Zone 1 to Zone 2): All cooling and electrical commands flowing from the facility BMS to rack-level BMCs must pass through a hardware unidirectional data diode or a proxy enforcing strict schema validation. Write operations are restricted to non-destructive setpoints; commands requesting emergency fluid shutoff or power cutoff require dual cryptographic authorization.
- Conduit B (Zone 2 to Zone 3): Communication between the BMC and the host CPU occurs exclusively over MCTP-over-PCIe or SMBus. The host processor treats the BMC as untrusted: every configuration update dispatched by the BMC is checked against local cryptographic policy before it takes effect, meaning the signature chain, the rollback counter and the scope the policy grants that update.
- Conduit C (Zone 3 to Zone 4): The interface between the host operating system and the accelerator silicon is an uncompromising SL-4 conduit. Unencrypted direct memory access (DMA) is structurally impossible. All data transfers traverse SPDM 1.3 authenticated and encrypted PCIe IDE channels [6], [7].
4.2 CLC/TS 50701 & EN 50126 RAMS Engineering#
CLC/TS 50701 and EN 50126 govern the engineering of Safety-Critical and Cybersecurity systems in environments where failure induces loss of life or catastrophic infrastructure collapse. Applying these standards [2], [3] to frontier AI datacenters puts Reliability, Availability, Maintainability, and Safety (RAMS) parameters on a stated failure model with stated rates, rather than leaving them to be hoped for. The derivation is only as good as the rates fed into it. Where this paper states a rate, it states whether the rate came from a component datasheet or from the working group's own judgment, and no rate here is drawn from an operating fleet.
Hazard Identification & Safety Instrumented Systems (SIS)#
Under EN 50126 [3], every potential failure mode within the AI Rack Envelope is categorized into a Safety Integrity Level (SIL), the scale IEC 61508 defines [18]:
- Thermal Runaway Mitigation (SIL-2 / SIL-3): If manifold differential pressure drops below safe operational limits (), an autonomous, hardware-wired Safety Instrumented Function (SIF) triggers an orderly execution pause and progressive power derate within 500 milliseconds, long before silicon approaches thermal destruction. This SIF is implemented using hardwired, analog comparator logic completely decoupled from the software BMC or OS.
- Overcurrent Busbar Protection (SIL-3): Electronic circuit breakers (eFuses) located directly on the 48V power sled monitor current derivative (). If abnormal current transients exceeding design baselines are detected (indicating a resonant frequency attack or hard short), the eFuse severs the bus connection within 5 microseconds, preserving silicon health.
Reconciling Safety Interlocks with Security Mandates#
A central challenge in critical infrastructure engineering is resolving fundamental conflicts between safety systems and security controls:
- The Conflict: Traditional industrial safety systems default to a "fail-open" or "de-energize to safe state" posture to preserve physical human life and mechanical equipment. Conversely, high-assurance security systems default to "fail-closed" or "quarantine and lock" to prevent unauthorized exfiltration or tampering.
- The Resolution: Within the AI Rack Envelope, the safety interlock (e.g., thermal power cutoff) triggers an atomic Cryptographic Zeroization Primitive before power rails collapse. When an emergency shutdown signal is asserted by the hardwired SIL-3 safety system:
- The accelerator security processor captures a 5-millisecond reserve power window provided by dedicated on-board holdup capacitors.
- The processor instantly overwrites all internal HBM symmetric decryption keys stored in battery-backed SRAM with randomized bit patterns ().
- This action instantaneously renders the billions of dollars of model weights residing in physical memory completely undecryptable, fulfilling the security requirement without delaying the safety-critical power de-energization.
5. Hardware Root of Trust (RoT) & Platform Firmware Integrity#
To ensure that silicon components within the AI Rack Envelope execute exclusively authentic, unmodified firmware from the initial millisecond of power application, the platform architecture incorporates an immutable, open-source Hardware Root of Trust.
5.1 Open-Source Silicon RoT: The Caliptra Specification#
Proprietary vendor roots of trust introduce opaque security implementations that resist comprehensive independent security auditing. The AI Rack Envelope mandates integration of the open-source Caliptra Silicon Root of Trust specification [5] directly into accelerator ASIC dies, host processors, and switching chips.
Caliptra provides:
- Immutable Silicon Core: An integrated micro-controller containing 128 KB of mask ROM, physically etched during semiconductor fabrication. This ROM contains the non-modifiable primary bootloader, establishing the root of the verification chain. This is the root of trust for detection and update that NIST SP 800-193 requires of a resilient platform [4].
- Deterministic Cryptographic Identity (DICE): Device Identifier Composition Engine (DICE) architecture generates an asymmetric cryptographic identity derived from an immutable Unique Device Secret (UDS) fused into silicon registers during wafer manufacturing. Every firmware update produces a new Compound Device Identifier (CDI):
If an adversary modifies a single byte of firmware code, the resulting device private key completely shifts, preventing the compromised firmware from decrypting authorized platform secrets or establishing network sessions a peer will accept, because the handshake requires a signature the shifted key cannot produce.
5.2 SPDM 1.3 Component Authentication & Attestation#
Before any compute tray within the AI Rack Envelope is admitted into the distributed training fabric, it must undergo formal mutual attestation using the Security Protocol and Data Model (SPDM) version 1.3 standard [6]:
- Cryptographic Measurement Exchange: The central cluster control plane dispatches a cryptographically randomized nonce challenge to each accelerator module over the management bus.
- Hardware-Signed Evidence: The Caliptra RoT generates an SPDM Certificate Chain response, signing the nonce alongside the current contents of its internal Platform Configuration Registers (PCRs), which record the exact cryptographic hashes of all active firmware, microcode patches, board strapping resistor configurations, and initialization tables.
- Automated Admission Decision: The cluster orchestrator verifies the attestation signature against the semiconductor manufacturer's public certificate authority. If a node exhibits an unrecognized PCR measurement, it is isolated by hardware fabric switches and prevented from receiving model weight parameters.
5.3 Post-Quantum Cryptographic (PQC) Silicon Roadmap#
Frontier model weights represent assets with operational lifetimes spanning multiple decades. Consequently, platform integrity mechanisms must defend against adversaries operating under the "Harvest Now, Decrypt Later" doctrine. The AI Rack Envelope mandates alignment with the National Security Agency's Commercial National Security Algorithm Suite 2.0 (CNSA 2.0) [9]:
- Stateful Hash-Based Signatures (LMS / XMSS): Primary firmware images and boot stages are signed using Leighton-Micali Signatures (LMS) in compliance with NIST SP 800-208 [8]. Because hash-based signatures rely strictly on collision-resistant hash functions rather than discrete logarithms or elliptic curves, they are inherently immune to Shor's algorithm on quantum computers.
- Module-Lattice-Based Digital Signatures (ML-DSA-87 / Dilithium): Runtime SPDM device attestation and ephemeral inter-tray session handshakes execute using ML-DSA-87, ensuring quantum-resistant mutual authentication across the coherent compute fabric.
6. Machine-Speed Autonomous Verification & Platform Falsification#
The scale and complexity of frontier AI hardware configurations render traditional vulnerability management frameworks obsolete.
6.1 Limitations of Human-Speed Security Incident Response#
Traditional hardware vulnerability management relies on Product Security Incident Response Teams (PSIRTs). When a security researcher reports a silicon vulnerability (such as a transient execution bug, sideband leakage, or firmware flaw), the vendor initiates a triage process involving manual human analysis, architectural review, microcode simulation, patch generation, and physical validation. This process consumes an average of 60 to 180 days.
In frontier AI environments, threat actors deploy autonomous agentic exploit generation tools capable of synthesizing novel microarchitectural attack variants within minutes. Pitting human-speed PSIRT workflows against machine-speed automated exploitation guarantees systemic platform compromise.
6.2 Autonomous Agent-Driven Testbenches & Hardware-in-the-Loop (HIL) Emulation#
The AI Rack Envelope incorporates a continuous, machine-speed automated falsification architecture. Rather than treating security validation as an annual compliance milestone, the platform establishes a closed-loop verification pipeline coupling autonomous security test agents with high-fidelity Hardware-in-the-Loop (HIL) digital twin emulators:
| Pipeline stage | Actions performed |
|---|---|
| 1. Autonomous Agentic Exploit Generation | Hypothesizes microarchitectural race conditions; synthesizes fault injection patterns (di/dt, thermal cycles); generates novel SPDM/PCIe IDE fuzzing payloads |
| 2. Hardware-in-the-Loop (HIL) & FPGA Digital Twin Emulation | Executes attack hypothesis against cycle-accurate silicon model; physical testbed applies thermal/electrical stress in real time |
| 3. Automated Mitigation Synthesis & Firmware Lockout | Generates microcode patch or isolates vulnerable execution lane; updates IEC 62443 conduit firewall rules in < 500 ms |
- Automated Exploit Hypothesis Generation: Machine-speed agentic models ingest hardware register descriptions, Verilog RTL code, and platform firmware binaries. The models analyze race conditions, sideband leakage vectors, and thermal vulnerability windows, generating thousands of concrete, executable exploit candidates per hour.
- Cycle-Accurate HIL Execution: Candidate exploits are immediately dispatched to parallelized FPGA-based silicon emulators and instrumented physical test racks. The testbed monitors bus telemetry, power supply noise, and register state transitions, verifying whether the candidate exploit successfully violates security invariants.
- Automated Microcode Synthesis: When the pipeline reproduces a working exploit for a candidate vulnerability on the hardware-in-the-loop testbench, it synthesizes candidate microcode mitigations (e.g., pipeline serialization fences, disabled branch predictors, or rate-limited sidebands), re-runs the same exploit to confirm the mitigation blocks it, re-runs the functional safety suite to confirm the mitigation has not broken a trip path, and distributes the cryptographically signed patch across the production cluster. The confirmation is against the exploit the pipeline holds, which is narrower than a claim that the vulnerability class is closed.
7. Supply Chain Provenance and the Multi-BOM Architecture#
Hardware security within the AI Rack Envelope is fundamentally contingent upon verifiable provenance across the global semiconductor and system integration supply chain. An adversary who intercepts a compute tray during manufacturing, transit, or rack assembly can embed microscopic hardware trojans; such as rogue interposers, malicious passive components, or tampered SPI flash chips; that bypass all subsequent logical security controls.
7.1 The Unified Multi-BOM Provenance Architecture#
To eliminate supply chain ambiguity, the platform mandates the BOM layers defined in section 2.1 of Unified DEXPI 2.0 & CycloneDX 1.6+ Semantic Bridge, which is where the layer count is stated once for the corpus. Each layer is expressed in CycloneDX 1.6 [11], and together they give a machine-readable, cryptographically verifiable attestation of every element residing within the physical envelope. The selection of layers is Eigenia's composition over CycloneDX rather than a construct the specification mandates; the component types in the third column are CycloneDX's own:
| # | Bill of Materials | CycloneDX component type | Tracked artifacts |
|---|---|---|---|
| 1 | Hardware BOM (HBOM) | device, hardware | Silicon die revisions and steppings; interposer and substrate lots; board layout and passive bill; silicon foundry provenance IDs; liquid cold plates, quick-disconnect couplings, secondary manifolds and 48 V busbar assemblies as physical parts |
| 2 | Software BOM (SBOM) | firmware, library, application | Host Linux kernel modules; ROCm / CUDA runtime drivers; OpenBMC embedded Linux tree; cryptographic microcode; CDU programmable logic firmware; smart ePDU microcontroller images; Modbus/BACnet gateway firmware; optical leak sensor firmware |
| 3 | Cryptography BOM (CBOM) | cryptographic-asset | DICE certificate chains; Unique Device Secrets and Compound Device Identifiers; LMS and ML-DSA-87 signing keys; the algorithm inventory each boot stage depends on |
| 4 | Manufacturing BOM (MBOM) | component, metadata.manufacturer | Cold plate alloy and elastomer material certificates; quick-disconnect seal (EPDM/FKM) formulations; manifold weld procedure records; 48 V busbar copper certifications; wafer lot numbers; ODM chain-of-custody signatures |
| 5 | Operations BOM (OBOM) | data, service | Rack power caps; coolant flow floors; thermal throttling thresholds; the operational envelope each control loop is permitted to command |
| 6 | SaaS BOM (SaaSBOM) | service | Redfish BMC endpoints; IPMI and SNMP management interfaces; facility Modbus TCP and BACnet endpoints exposed by CDU and ePDU controllers |
- Hardware Bill of Materials (HBOM): Extends BOM transparency into silicon, circuit board, and mechanical components. Captures exact die steppings, silicon fabrication foundry IDs, packaging substrate lot numbers, wafer serializations, and vendor manufacturing runs for all discrete silicon components (accelerators, NICs, PCIe switches, VRM controllers). The same layer carries the physical cooling hardware: cold plates, quick-disconnect couplings, secondary manifolds, and busbar assemblies. CycloneDX has no separate materials layer, and its physical part inventory is the HBOM; this document uses it that way rather than inventing a layer name no validator would recognize.
- Software Bill of Materials (SBOM): Enumerates all active host software, container layers, device drivers, and firmware blobs, including the firmware running on the facility side of the envelope: CDU programmable logic, smart ePDU microcontrollers, Modbus and BACnet gateways, and optical leak sensors. Cryptographic hashes of all binaries are cross-referenced continuously against known vulnerability databases and Vulnerability Exploitability eXchange (VEX) statements.
- Cryptography Bill of Materials (CBOM): Inventories the cryptographic assets the platform depends on: DICE certificate chains, the Unique Device Secret fused during wafer manufacturing, each Compound Device Identifier derived from it, the LMS keys that sign boot stages, and the ML-DSA-87 keys that sign runtime attestations. CycloneDX binds this layer to the
cryptographic-assetcomponent type, which is what turns a post-quantum migration into a query against an inventory rather than a survey of engineers. - Manufacturing Bill of Materials (MBOM): Carries the provenance of the physical materials the HBOM enumerates: copper cold plate metallurgy, ethylene-propylene-diene-monomer (EPDM) seal formulations, dielectric fluid chemical specifications, manifold weld procedure records, wafer lot numbers, and ODM chain-of-custody signatures. Equipment classes are stated against the ISO 15926-4 reference data library [19], [23], ensuring that counterfeit mechanical elements that degrade thermal safety are identified prior to operational energization.
- Operations Bill of Materials (OBOM): Details the logic configurations, setpoint limits, and operating envelopes enforced within row-level CDUs, power supply units, busbar monitor controllers, and environmental sensor microcontrollers. This is the layer an underwriter reads to establish what the facility is permitted to do, as distinct from what it is built from.
- SaaS Bill of Materials (SaaSBOM): Records the management and telemetry services the envelope exposes, including Redfish BMC endpoints, IPMI and SNMP interfaces, and the facility Modbus TCP and BACnet endpoints that reach the cooling and power controllers. An endpoint absent from this layer is an endpoint nobody is defending.
7.2 Open Platform Initialization Architecture (OpenSIL Migration)#
A critical requirement for hardware supply chain transparency is the total elimination of proprietary, closed-source binary initialization blobs in host CPUs and accelerators. Historically, platform boot sequences rely on proprietary Unified Extensible Firmware Interface (UEFI) initialization blobs that execute at highest CPU privilege levels (Ring -2 / System Management Mode) prior to operating system initialization. These binary blobs represent massive, un-auditable security risks.
The AI Rack Envelope enforces migration to open-source platform initialization:
- OpenSIL Integration: Platform boot firmware integrates the Open Silicon Initialization Library (OpenSIL) [15], stripping monolithic firmware into modular, open-source execution libraries written in memory-safe paradigms.
- Transparent Coreboot Payload: OpenSIL executes within an open-source coreboot environment, enabling platform operators to audit every line of code executed during silicon reset, memory training, and PCIe link negotiation.
- Supply Chain Cryptographic Ledger & EU CRA Alignment: Every build artifact, binary image, and configuration script is anchored into a public, tamper-proof transparency log (e.g., Sigstore Rekor), enabling cryptographically irrefutable verification that deployed hardware satisfies the binding essential cybersecurity requirements of the EU Cyber Resilience Act (Regulation 2024/2847).
8. Mathematical Formulations & Thermodynamic Hazard Calculus#
This section formalizes the physical, thermodynamic, and electrical hazard dynamics governing the AI Rack Envelope. The relations that follow are modeled. Each takes inputs the working group chose, applies a stated physical relation, and returns an exact result. Each result in this section is a modeled figure computed from stated inputs, and where a coefficient or a rate is assumed, the text says so at the point of use.
8.1 Thermodynamic Heat Balance & Silicon Thermal Runaway#
The thermal state of an accelerator die within the AI Rack Envelope is governed by the dynamic energy conservation equation:
Where:
- is the junction temperature of the silicon die ().
- is the lumped effective thermal capacitance of the silicon die, thermal interface material (TIM), and copper cold plate ().
- is the instantaneous electrical power dissipation of the tensor compute units ().
- is the instantaneous convective heat evacuation rate delivered by the liquid cooling loop ().
The convective heat transfer rate is defined by:
Where represents the mass flow rate of the coolant (), is the specific heat capacity of the fluid (), is the overall heat transfer coefficient, and is the microchannel contact surface area. In high-density racks, coolant is delivered via Propylene Glycol 25 percent (PG25) dielectric mixtures at volumetric flow rates between 40 L/min and 80 L/min per rack under a nominal operating pressure of 3.0 bar to 4.0 bar across the manifold distribution blocks.
Under a malicious facility-level cooling interdiction, an adversary commands the CDU valves to close, causing mass flow rate to decay exponentially:
Substituting into the heat balance equation yields the differential equation for junction temperature rate of change:
As , the system enters adiabatic runaway:
For an accelerator package consuming () [24] with a stagnant cold plate thermal capacitance , dominated by the coolant retained in the channels once flow stops:
Starting from the nominal operating temperature of , the emergency hardware shutdown trip point () is reached in:
The exponential solution of the same energy balance refines this to 14.8 seconds. This proves mathematically that software-based, human-speed alerting is entirely incapable of holding the compute load; autonomous, hardware-level SIL-3 safety interlocks must execute within sub-second timescales. Across a 100 MW cluster, uncontained thermal cascades present catastrophic property destruction risks.
8.2 Power Distribution Transient Response & Resonant Voltage Perturbation#
The power distribution network (PDN) feeding the accelerator tray is modeled as an equivalent RLC circuit:
Under an adversarial step-frequency attack, an attacker modulates computational activity at frequency :
The second-order differential equation governing voltage response across the accelerator die power pins is:
Where the natural resonant frequency and damping ratio are:
When the attacker tunes the workload toggling frequency to match the natural resonant frequency (), the steady-state voltage fluctuation is amplified by the circuit quality factor :
For high-current 48V distribution bars where , , and , the induced resonance creates voltage spikes exceeding , destroying point-of-load VRMs and inducing physical gate-oxide breakdown in silicon compute cores.
8.3 Markov Reliability & Cyber-Physical Hazard Formulation#
The operational reliability of the AI Rack Envelope under continuous cyber-physical stress is formalized via a multi-state continuous-time Markov chain (CTMC). We define four discrete operational states:
- State (Nominal Secure): All zones operating within normative parameters; every cryptographic envelope passing its signature check on presentation; cooling and power at steady state.
- State (Degraded / Stressed): Facility threat pressure detected (e.g., elevated cooling temperatures, unusual transient noise, unverified SPDM challenge); autonomous safety throttling active.
- State (Interdiction Containment): Safety Instrumented System (SIS) triggered; cryptographic keys zeroized; physical compute isolated.
- State (Catastrophic Failure): Boundary breached; model weights compromised or silicon physically destroyed through thermal/electrical runaway.
The state probability transition vector satisfies the Chapman-Kolmogorov forward differential equation:
Where the transition rate generator matrix is defined as:
Where:
- is the physical stress arrival rate (cooling disturbances, power grid fluctuations).
- is the cyber threat arrival rate (BMC exploits, sideband attacks, fabric scans).
- is the autonomous self-healing recovery rate of platform control loops.
- is the activation rate of hardwired SIL-3 Safety Instrumented Functions.
- is the rate of boundary collapse leading to catastrophic asset loss.
- is the facility recovery rate from safe zeroized state back to operational baseline.
The hazard rate , defining the instantaneous failure probability given survival up to time , is expressed as:
By enforcing the AI Rack Envelope controls (sub-second hardware SIF interlocks, Caliptra RoT, and line-rate IDE bus encryption), the failure transition rate is reduced to near-zero (), ensuring that even under persistent facility attacks, the system transitions deterministically to the contained zeroized state rather than the catastrophic failure state .
9. Actuarial Risk Modeling, Financial Loss Calculus & Lloyd's Y5381 Underwriting#
Engineering risk models such as FMECA and HAZOP determine physical failure modes and severity, but fail to satisfy the capital allocation and balance sheet protection requirements demanded by Chief Financial Officers, enterprise risk committees, and insurance syndicates. To make the AI Rack Envelope commercially viable, this section establishes the formal actuarial risk calculus bridging cyber-physical engineering with insurance underwriting.
9.1 Single Loss Expectancy (SLE) & Annualized Loss Expectancy (ALE)#
In a frontier AI cluster, asset valuation comprises both physical hardware replacement cost and consequential loss resulting from training checkpoint corruption, re-computation overhead, and prolonged business interruption.
The Single Loss Expectancy (SLE) for a catastrophic boundary breach is defined per NIST SP 800-30 Rev. 1 [22] as:
For a standardized 120 kW AI Rack Envelope containing eight compute trays (64 accelerator modules, 16 host CPUs, dual 800G NICs per tray, and associated high-bandwidth memory):
- Direct Hardware Replacement Cost (): 64 accelerators at $35,000 USD plus chassis, switching fabrics, cold plates, and power sleds yields per rack.
- Consequential Business Interruption & Model Loss (): A frontier model training run employing 2,048 accelerators operates at an amortized capital run-rate of approximately $120,000 USD per hour. A systemic thermal or electrical cascade that corrupts memory state, triggers unrecoverable filesystem corruption, and forces a four-week rollback to a clean checkpoint induces consequential business interruption losses exceeding $80,640,000 USD.
- Probable Maximum Loss (PML): Under an uncontained facility cyber-physical event affecting a 16-rack row (1,024 accelerators), the total exposed asset valuation is:
Under baseline un-hardened infrastructure lacking the AI Rack Envelope (), the Single Loss Expectancy is:
The working group assumes an Annualized Rate of Occurrence (ARO) for severe facility disturbances, grid instabilities, and targeted cyber-physical attacks of . That figure is modeled, not measured. It expresses a judgment of roughly one severe event per eight years at a site of this class, and no incident dataset, insurer loss run or operator record is cited for it. The two exposure factors used below, unhardened and hardened, are the working group's judgment on the same footing. Every figure downstream inherits all three assumptions. On the assumed rate, the baseline Annualized Loss Expectancy (ALE) is:
9.2 Return on Security Investment (ROSI) & Gordon-Loeb Capital Ceilings#
Implementing the AI Rack Envelope; encompassing Caliptra silicon roots of trust, PCIe IDE encryption engines, SIL-3 safety interlocks, and automated HIL testbenches; requires a capital expenditure of approximately $45,000 USD per rack ($720,000 USD across the 16-rack row) with an annualized operational maintenance cost of $180,000 USD per year, yielding a total annualized security control cost .
By enforcing sub-second SIL-3 emergency derating and line-rate memory encryption, the mitigated Exposure Factor drops to (preventing physical destruction and restricting impact to temporary job pause), while the mitigated threat occurrence drops to :
The net annual monetary loss reduction is:
The Return on Security Investment (ROSI) is expressed as:
This ROSI is modeled. It reproduces exactly to 1316.69 percent, and that exactness is a property of the division, not evidence for the five numbers entering it: the assumed ARO of 0.12, the two assumed exposure factors, the accelerator unit cost, and the 120,000 USD per hour run-rate for a 2,048-accelerator training job. Change the assumed ARO alone and the return moves proportionally, because ARO scales the numerator and leaves the control cost untouched. What survives the assumptions is the ordering: a 900,000 USD annual control spend sits against a modeled exposure two orders of magnitude larger, so the case for the controls does not rest on the precise percentage. The percentage should not travel into a board paper without the operator's own loss data replacing all five inputs.
The Gordon-Loeb result [21] holds that optimal expenditure to protect an information asset should generally not exceed () of the expected loss. That is a sourced bound on the ratio; the loss it is applied to here is the modeled derived above, so the ceiling inherits the assumed ARO and exposure factors:
The annualized cost of the AI Rack Envelope, 900,000 USD, sits at 19 percent of that ceiling. Under the model's own assumptions the spend is therefore well inside the bound Gordon-Loeb sets, which is the most the calculation can say. It is a consistency check on the size of the investment, not an independent confirmation that the investment pays.
9.3 Reinsurance Covenants, the Lloyd's Y5381 Exclusion & Deductibles#
In the commercial property and casualty market, standard commercial property policies explicitly exclude losses caused by cyber interdictions unless affirmative endorsements are attached. Stand-alone cyber-attack cover carries a mandatory exclusion of its own: Lloyd's Market Bulletin Y5381, issued by the Corporation of Lloyd's on 16 August 2022, requires that stand-alone cyber-attack policies written or renewed from 31 March 2023 exclude losses arising from war and from state-backed cyber attacks that significantly impair the ability of a state to function or that significantly impair the security capabilities of a state [20]. The Lloyd's Market Association's model clauses LMA5564 to LMA5567 (November 2021) are the wordings drafted to meet it.
To obtain affirmative coverage and prevent crippling sub-limits or uninsurable exclusions:
- IEC 62443 SL-4 as Underwriting Warranties: Insurers mandate that compute zones housing critical assets satisfy SL-3 or SL-4 conduit segmentation, evidenced by a third-party assessment report the syndicate can read. Failure to maintain independent hardware roots of trust (Caliptra) and line-rate encryption voids affirmative coverage upon forensic investigation.
- Dynamic Retention Deductibles: Facilities implementing multi-BOM attestations across all six layers (CycloneDX 1.6+) that an independent assessor has signed off, together with SIL-3 safety interlocks, qualify for base retention deductibles of $250,000 USD per occurrence. Unverified facilities face punitive deductibles exceeding $5,000,000 USD and severe indemnity sub-limits capping business interruption recoveries at less than of total loss.
- Catastrophe Risk Accumulation: Reinsurers deploy deterministic catastrophe models to evaluate simultaneous multi-facility failure across power distribution zones. The AI Rack Envelope provides the provable physical isolation required to decouple correlated rack failures, transforming an uninsurable systemic catastrophe into an actuarially sound, diversified underwriting risk profile.
10. References#
- International Electrotechnical Commission (IEC). IEC 62443-3-3: Industrial communication networks - Network and system security - Part 3-3: System security requirements and security levels. International Standard, 2018.
- European Committee for Electrotechnical Standardization (CENELEC). CLC/TS 50701: Railway applications - Cybersecurity. Technical Specification, 2021.
- European Committee for Electrotechnical Standardization (CENELEC). EN 50126-1: Railway applications - The specification and demonstration of Reliability, Availability, Maintainability and Safety (RAMS). European Standard, 2017.
- National Institute of Standards and Technology (NIST). NIST SP 800-193: Platform Firmware Resiliency Guidelines. Special Publication, U.S. Department of Commerce, 2018.
- Open Compute Project (OCP). Caliptra: Open Source Silicon Root of Trust Specification. Version 1.0, OCP Security Project, 2023.
- Distributed Management Task Force (DMTF). Security Protocol and Data Model (SPDM) Specification. DSP0274, Version 1.3.0, 2023.
- PCI-SIG. PCIe Integrity and Data Encryption (IDE) Specification. Version 1.0, Peripheral Component Interconnect Special Interest Group, 2020.
- National Institute of Standards and Technology (NIST). NIST SP 800-208: Recommendation for Stateful Hash-Based Signature Schemes. Special Publication, U.S. Department of Commerce, 2020.
- National Security Agency (NSA). Announcing the Commercial National Security Algorithm Suite 2.0 (CNSA 2.0). Cybersecurity Advisory, Fort Meade, MD, 2022.
- European Commission. Regulation of the European Parliament and of the Council on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act). COM(2022) 454 final, Brussels, 2022.
- OWASP Foundation / Ecma International. CycloneDX Bill of Materials Specification. ECMA-424, 1st edition, June 2024, defining CycloneDX v1.6. Ecma International Technical Committee 54 (TC54), Geneva. Not to be confused with ISO/IEC 5962:2021, which is the SPDX specification.
- McKenney, J. Systems Assurance in High-Entropy Industrial Complexes: Mathematical Modeling of Boundary Failures and Cognitive Distortion. Eigenia Labs Monograph Series, WG-05-CAD, 2026.
- Taleb, N. N. Antifragile: Things That Gain from Disorder. Random House, New York, 2012.
- Cisco Systems. Line-Rate Media Access Control Security (MACsec) across Megawatt Infrastructure. Technical White Paper, 2022.
- Open Compute Project & OpenSIL Consortium. Open Silicon Initialization Library (OpenSIL) Architecture and Implementation Guide. Industry Consortium Standard, 2024.
- Granovetter, M. Threshold Models of Collective Behavior. American Journal of Sociology, vol. 83, no. 6, pp. 1420-1443, 1978.
- Kramers, H. A. Brownian motion in a field of force and the diffusion model of chemical reactions. Physica, vol. 7, no. 4, pp. 284-304, 1940.
- IEC. IEC 61508: Functional safety of electrical/electronic/programmable electronic safety-related systems. Parts 1-7, International Standard, 2010.
- International Organization for Standardization (ISO). ISO 15926-1: Industrial automation systems and integration - Integration of life-cycle data for process plants. International Standard, 2004.
- Lloyd's. Market Bulletin Y5381: Cyber-attack exclusions. Corporation of Lloyd's, London, 16 August 2022. The model wordings drafted to meet it are the Lloyd's Market Association's clauses LMA5564 to LMA5567, November 2021.
- Gordon, L. A., and Loeb, M. P. The Economics of Information Security Investment. ACM Transactions on Information and System Security, vol. 5, no. 4, pp. 438-457, 2002.
- National Institute of Standards and Technology (NIST). NIST SP 800-30 Rev. 1: Guide for Conducting Risk Assessments. U.S. Department of Commerce, 2012.
- International Organization for Standardization (ISO). ISO/TS 15926-4: Industrial automation systems and integration - Integration of life-cycle data for process plants including oil and gas production facilities - Part 4: Core reference data. Technical Specification, 3rd edition, 2024. This is the part that holds the reference data library; the DEXPI Sandbox extends it with the additional class definitions the DEXPI P&ID specification requires.
- NVIDIA Corporation. Datasheet for NVIDIA Blackwell Architecture. Product datasheet. Cited for the per-accelerator power figure used in section 8.1, consistent with the configurable maximum NVIDIA publishes for a GB200-class Blackwell GPU.