Skip to content
Vulnerabilities

AI Data Center Security Checklist: GPUs, Firmware and Supply Chains

The hardware under AI workloads is now a major security concern. GPUs, baseboard management controllers, firmware, drivers and the components behind them can all give attackers a path to persistent, low-level access. Recent guidance and active exploitation have raised the stakes. NIST published an H...

· Jul 29, 2026 · 6 min read · 👁 0 views
AI Data Center Security Checklist: GPUs, Firmware and Supply Chains

The hardware under AI workloads is now a major security concern. GPUs, baseboard management controllers, firmware, drivers and the components behind them can all give attackers a path to persistent, low-level access.

Recent guidance and active exploitation have raised the stakes. NIST published an HPC security overlay, the G7 released AI SBOM guidance and CVE-2024-54085 was added to CISA’s Known Exploited Vulnerabilities catalog.

This checklist turns that fast-moving picture into practical controls for AI data center security at the firmware and hardware layers.

Securing that layer is its own discipline. Eclypsium works below the operating system, giving teams continuous visibility into the firmware, hardware and low-level components of servers, network devices and endpoints, the layer where endpoint detection and legacy vulnerability tools have little reach.

The controls below map to that same firmware and hardware layer.

What changed in 2025 and 2026

Several developments moved firmware and supply chain risk higher on the priority list for platform teams.

NIST SP 800-234, published May 4, 2026, tailors 60 SP 800-53 controls to HPC environments that support AI workloads.

NIST states that securing HPC is essential to protecting AI models and sensitive data.

G7 AI SBOM guidance defines seven clusters covering metadata, models, KPIs, infrastructure, security properties, system-level properties and dataset properties.

CVE-2024-54085, disclosed by Eclypsium research in March 2025, allows remote authentication bypass through the Redfish Host Interface in AMI MegaRAC SPx.

CISA added it to the KEV catalog on June 25, 2025, reported as the first BMC vulnerability to be listed.

Protected PCIe and Secure AI features for NVIDIA platforms introduced specific firmware and platform prerequisites.

Compliance expectations now extend to AI supply chains and partners, not only operators. That shift shapes every section below.

GPU platform hardening and attestation

Start by proving hardware and firmware are in a known state before any workload runs.

Require remote attestation before workloads. The NVIDIA Attestation Suite includes NRAS and a RIM service to cryptographically verify GPU hardware and firmware integrity across fleets.

Enforce firmware baselines. For Protected PCIe mode, NVIDIA advises HGX Hopper firmware 1.7.0 to avoid a known issue in 1.6.0 that can cause the GPU to fall off the bus during boot.

For Secure AI, once you are on firmware 1.8.0 or later, rollback to 1.7.1 or earlier is not supported.

Confirm Protected PCIe prerequisites. It requires H100 or H200 GPUs on HGX 8-GPU systems, a CPU trusted execution environment such as AMD SEV-SNP or Intel TDX, a CUDA 12.8 or later driver and firmware 1.7.0 or later.

Bind key and secret release to verified attestation, and log attestation and firmware change events in your SIEM for review.

Track GPU driver and vGPU advisories on a schedule, since one outdated driver can weaken an otherwise clean baseline.

Keep inventory evidence tied to firmware integrity checks and attestation logs.

This matters most in shared fleets, where components move between tenants during intake, re-assignment and after maintenance windows.

Eclypsium inventories NVIDIA GPUs, verifies firmware integrity and flags vulnerable GPU drivers across those fleets.

That inventory doubles as audit evidence when a tenant asks what state the hardware was in before their workload ran.

BMC and server-board controls

Baseboard management controllers are frequent targets because they persist below the operating system.

Researchers have reported BMC firmware vulnerabilities and exploit paths that can support persistent access in data centers.

Isolate BMC interfaces on management networks and block internet exposure. Patch and verify BMC firmware, and confirm that any No Auth settings are disabled.

Treat Redfish access as high-risk and review vendor advisories for MegaRAC and OpenBMC components.

NVIDIA’s DGX B200 firmware guide notes two U-Boot vulnerabilities in OpenBMC, CVE-2024-57258 and CVE-2024-57256. It provides patched firmware for both.

Monitor BMC activity as security telemetry and enforce UEFI and Secure Boot protections.

Supply chain integrity and bills of materials

You cannot secure components you cannot account for. Bills of materials give teams a structured way to track those components.

Require SBOMs and MBOMs for models, datasets, firmware, drivers and hardware components.

Canada’s Cyber Centre advises generating and maintaining SBOMs and, where possible, MBOMs that capture provenance and versioning for AI components, and updating them whenever systems change.

Verify GPU authenticity and firmware integrity on receipt and between tenancies. Contract for vulnerability notification SLAs and secure update channels.

Map supplier responsibilities to NIST and CISA guidance and capture provenance for audits.

Shared compute sanitization and decommissioning

Shared clusters swap tenants and workloads often, so firmware state should be re-checked at each handoff.

Between training runs or tenant swaps, re-verify GPU and BMC firmware state. Clear device state and validate against known-good RIMs before re-assignment.

For end-of-life systems, apply cryptographic erasure and certified destruction, then revoke orphaned credentials and remove dangling DNS or out-of-band access.

Retirement deserves the same rigor as intake. A structured approach to data center decommissioning reduces the chance that old credentials or forgotten management paths remain after hardware leaves the floor.

Governance tie-ins

Map firmware and hardware controls to NIST SP 800-234 overlays. Use AI SBOM minimum elements to set procurement checklists for models, datasets and infrastructure.

Document exceptions and mitigations so auditors and responders can trace decisions.

How to start in 30, 60 and 90 days

30 days: Inventory GPUs and BMCs, collect firmware hashes and enable attestation pilots.

60 days: Enforce management network isolation, patch BMCs and adopt intake checks for GPUs.

90 days: Require AI SBOMs and MBOMs from suppliers and enforce RIM-based acceptance for new GPU trays.

Frequently asked questions

What companies secure AI data center servers and GPUs?

Securing AI servers and GPUs is mostly about the firmware and hardware layer, below the operating system, where attackers seek persistent access.

The work covers GPU and BMC firmware integrity, remote attestation, supply chain provenance and sanitization between tenants.

Eclypsium focuses on that layer, providing continuous firmware and hardware monitoring across servers, network devices and endpoints.

Why is standard endpoint security not enough for AI infrastructure?

Endpoint detection and legacy vulnerability tools operate at or above the operating system, so they miss firmware, BMC and component-level risks.

AI infrastructure needs visibility into the low-level layers where GPUs, controllers and drivers actually run.

How is Protected PCIe different from standard confidential computing?

Protected PCIe adds platform prerequisites beyond a baseline confidential computing setup.

It requires H100 or H200 GPUs on HGX 8-GPU systems, a CPU trusted execution environment such as AMD SEV-SNP or Intel TDX, a CUDA 12.8 or later driver and NVIDIA firmware 1.7.0 or later.

How do we handle BMC firmware on shared clusters?

Re-verify BMC firmware state between tenants, keep management interfaces isolated and track advisories for MegaRAC and OpenBMC components.

Use patched firmware for known issues such as the OpenBMC U-Boot vulnerabilities documented for the DGX B200.

What should procurement request from AI infrastructure suppliers?

Request current SBOMs and MBOMs, firmware provenance, secure update details, vulnerability notification terms and evidence that GPU and BMC components can be verified at intake.

Source: CybersecurityNews.com

Follow ShomoySoft for more: Follow on Facebook

💬 Comments (0)

Login to join the discussion.

No comments yet. Be the first!

Related Articles

Recommended for you