AI factories are scaling faster than the people who run them
Industry forecasts project more than 100 gigawatts of new AI compute capacity added globally through 2030. The experienced operations staff to commission and run those facilities is not scaling on that curve.

According to the iMasons State of the Digital Infrastructure Industry 2026 report, the digital infrastructure industry has reached 373 gigawatts of total power capacity today, with 280 gigawatts more in the development pipeline. In the next three years alone, more capacity will come online than was built in the past thirty years. Industry analysts forecast total capacity could triple, quintuple, or more before the end of the decade. The experienced operations staff to commission and run those facilities is not scaling on that curve.
This is the workforce problem AI infrastructure is not solving.
New technologies, still building the operational playbook
Two technology transitions are hitting data centers in parallel: direct liquid cooling and high-voltage DC power distribution. Both are being deployed to support the compute density that modern GPU clusters require. Both are new enough that the safety standards, training programs, and operational practices that govern them are still being extended for this deployment context.
Liquid cooling means coolant distribution units, rear-door heat exchangers, direct-to-chip cold plates, and manifold networks running through the raised floor and overhead. Each of these has failure modes, isolation sequences, and emergency procedures that differ from anything in a traditional air-cooled facility. High-voltage DC distribution, which NVIDIA and the broader industry are adopting for next-generation AI factory deployments, is a different safety regime than the AC systems most data center staff trained on: different protective equipment, different isolation requirements, different consequences when something goes wrong.
The workforce that knows how to operate these systems safely is a small group. It is not growing as fast as the number of facilities that need it. According to the iMasons State of the Digital Infrastructure Industry 2026 report, the U.S. alone sees more than 80,000 electrician job openings annually; based on forecasts for U.S. data center growth, those openings could increase by a factor of ten or more. A single gigawatt-scale project under construction today requires more than 7,000 workers on site each day.
Where the knowledge lives today
A modern AI factory requires hundreds of operating procedures: standard operating procedures for routine tasks, method of procedure documents for one-time changes, and emergency operating procedures for fault conditions. Based on Visum’s work across AI factory facilities, the count runs above 800 per facility, and it grows as liquid cooling and high-voltage distribution come in.
Each procedure requires a senior engineer or operator to pull together information from multiple sources: equipment manuals from vendors, engineering drawings from the design team, safety standards from NFPA and ASHRAE, and site-specific configuration from commissioning records. None of these sources share a common format. A vendor manual documents the equipment in general terms; the specific configuration at this facility is in a commissioning report. A drawing shows the isolation point; the procedure for using it is in a different document, and the safety interlock that governs it is in a third.
That knowledge lives in the senior engineer’s head. They know which documents to apply, how to read them together, and which gaps to fill from experience. Writing one correct procedure takes days. A procedure that reaches operations without its source context cannot be verified or updated by anyone who was not in the room when it was written.
What existing tools do not address
The data center industry has invested heavily in two categories of tooling. Digital twins model geometry and physics: useful for design and capacity planning, but silent on what the operator should do when a pump fails or a leak alarm fires. Sensor and telemetry platforms optimize cooling and predict equipment failures: useful for runtime efficiency, but also silent on the human procedure side. A facility can be modeled in full detail and monitored continuously, and the technician on the floor is working from a document that carries no reference back to the drawings, manuals, and standards used to write it.
Neither category addresses the workforce problem directly.
The deeper problem: knowledge that cannot travel
The staffing shortage and the data problem reinforce each other. Operational knowledge is distributed across vendor manuals, engineering drawings, commissioning reports, and site-specific configuration records, none of which share a format or a common reference vocabulary. That knowledge cannot be transferred systematically across facilities, verified by someone who was not in the room when a decision was made, or updated reliably when equipment is replaced or a safety standard changes.
Uptime Institute research on human-factor outages has documented procedure-following failures as a recurring contributor. One documented pattern: operators diverge from the written procedure when the procedure no longer matches the specific site. Manuals update. Equipment gets replaced. Commissioning-era configuration drifts. The procedure written at handover becomes a liability over time, and the senior engineer who knows what changed is not always available.
This is a traceability and transfer problem, not just a staffing problem. The information exists. It is in manuals, drawings, and standards. The missing piece is a structured layer that keeps those sources connected to each other and to the procedures they back, so that the knowledge can be verified, updated, and used by people who did not author it.
What Visum AI is working on
Visum AI is collaborating across the data center industry, standards bodies, and research institutions to build that layer. The goal is a structured operational graph for AI factory facilities: every procedure traceable to the source documents that back it, every step bound to the specific equipment it applies to, every interlock and isolation sequence machine-readable rather than buried in narrative text.
The next post will cover what that layer looks like: a machine-readable schema that connects procedures, equipment, interlocks, isolation points, and source documents into a graph a query engine can traverse and the operations team can verify and act on.