Dark Gigawatts
Undisclosed operational risk in the AI power buildout.
The conversation is on X today!
“Apparently no one knows how to operate the onsite power for volatile data center loads. It is hard to believe any of them will reach reliability targets.” That worry is an open secret among energy and power engineers. We’ve heard repeatedly there are no “Dark GPUs” in comparison to the “Dark Fiber” concerns during the dot-com crash, so this time is different (at least for now). However, Dark Gigawatts may be an even more esoteric and pernicious concern considering the lack of public disclosure and the cost of getting it wrong. It’s clear that every powered GPU is currently in very high demand. Whether those GPUs keep delivering contracted compute when the power behind them falters, “Firm Compute”, is a far murkier picture.
SpaceX has attached $1.25 billion a month of Anthropic capacity fees to roughly 325,000 GPUs across Colossus and Colossus 2.[1] At the disclosed full rate, the fees annualize to $15 billion, an average of about $41 million per calendar day. A power event can reduce delivered compute without producing a headline campus outage; whether it reduces billed fees depends on contract terms that remain undisclosed. Major operators are building credible mitigations, yet public investors and stakeholders rarely see the combined power, controls, workload-recovery and contract case that shows who absorbs the shortfall.
Announced onsite generation alongside major U.S. AI campuses now totals more than ten gigawatts across the twelve projects we track, a mix of contracted equipment, permits and announced plans that lets developers reach power years before a conventional grid connection might arrive. How much of it is running today cannot be verified from public disclosure: commissioning records are not published, and the fastest-moving campuses are also running rented temporary turbines that are not publicly disclosed. Only one of the onsite powered datacenter projects that Occam Edge tracks publishes enough data to check whether its generation fleet holds meaningful reserve capacity, the spare generation margin that keeps the computers running when a turbine fails. Our review of public sources located none that publishes a complete plan for keeping the computers powered through an equipment failure.
Investors are valuing these projects on the permanent plants described in permits and announcements. At the SpaceX Colossus site where agency records allow verification, the onsite power is a temporary rental fleet with different equipment, different emission controls and a reserve that cannot be determined from the public record.[2] Longhorn, the Stargate campus in Abilene, is the second site where there are accessible agency records for review. The missing evidence to determine Firm Compute matters because power shortfalls can impact customer delivery, project cash flow, and the stability of the surrounding grid.
FIRM COMPUTE | Contracted compute that remains deliverable after the largest turbine fails, measured at one minute, one hour, one day and five days.
Turbine failure can be calculated
Every multi-unit power plant is built around routine equipment unavailability. On a 41-turbine island, an illustrative 3 to 5% forced-outage sensitivity implies one or two turbines down at any given time before planned maintenance. The frequency of any specific trip depends on the model and its history. The fleet-level math is well understood, and recent events show single-point failures reaching well-run islanded systems:
Kauai’s island utility blacked out island-wide in January 2024 after losing generation at a single station[3], twenty-six months after a single plant trip removed 60% of its supply[4].
A single turbine or instrumentation fault at Chevron’s Wheatstone LNG plant in 2023 cut a quarter of the facility’s output for roughly three days.[5]
The math gets harder before any failure occurs, because nameplate capacity overstates what a fleet can deliver. Reciprocating generator sets carry duty-class ratings under ISO 8528-1: standby units are built for limited backup hours, prime units for variable load, continuous units for full output around the clock. A standby rating assumes the set is backing up a normal utility source during interruptions, not serving as the primary source.[6] Gas turbines carry OEM site ratings with their own duty limits, start counts and maintenance clocks. Under either regime, a rented unit pressed into base load around the clock is working against duty limits the nameplate does not show. The restatements happen on the record: at Longhorn, Crusoe registered twelve turbines in August 2024, then in January 2025 cut the fleet to ten, raised the Titan 350 site rating from 35 to 38 megawatts, lowered the LM2500 from 35 to 34.1 megawatts, and raised the LM2500 fleet’s expected duty from 5,880 to 8,760 hours a year. Every unit is now expected to run every hour of the year.[31]
High temperatures cut output further: turbines are rated at 59 degrees Fahrenheit, and vendor sensitivity data show output falling 5% to 10% for every 18 degree rise, putting an aeroderivative fleet roughly a quarter below nameplate on a 113 degree afternoon. Model-specific site curves govern the actual number.[7] The hottest days are also the days cooling load peaks. The true reserve is the spare capacity left after adjusting ratings for duty class and summer heat. None of the projects we track disclose that number.
The starting test for Firm Compute borrows the utility N-1 criterion: can the campus lose its largest running unit at peak load and keep the computers on. Some critical loads warrant stricter criteria; the public record does not let even this basic test be checked. Engineers from Chevron and Schweitzer Engineering, writing up decades of islanded industrial practice, found that how reserve is spread across machines matters as much as how much reserve exists[8]. A cushion concentrated in one or two units can prove useless in the event itself. SpaceX’s Southaven permit for Colossus 2 documents how hard these fleets are expected to work: up to three starts per turbine per day with no annual hour caps, a load-following duty cycle[2].
How a campus absorbs a turbine failure
A campus survives a turbine trip only if the right equipment was sized and commissioned years in advance. Four systems respond to the failure, ordered here from fastest to slowest. In practice their responses overlap, and each must hold the load until slower systems catch up.
Server-level storage responds first, within milliseconds. NVIDIA’s newest server hardware includes capacitors that absorb the sharpest swings of a training load before they leave the cabinet. NVIDIA’s measured results from one GB300 rack test show up to a 30% cut in the peak demand the rack places on the site.[9]
Uninterruptible power supplies respond in under a second. Their capability depends directly on the batteries behind them. Vertiv’s published lab data shows an undersized battery bank smooths an AI load only briefly before it drains, while doubling the battery cabinets sustained the modeled smoothing profile without depleting the batteries.[10]
Plant-scale batteries respond within seconds. Megapack-class units at the plant or the interconnection absorb whatever the buildings cannot.
Reserve generation responds last, over minutes. Spare turbines pick up the load for the remainder of the event.
The same Vertiv study reports the constraint that matters most for a turbine island. Onsite generators tolerate far smaller load swings than a utility grid. In the configuration Vertiv modeled, load changes reaching the generators had to be held near 1% per second, because a faster swing causes the generators’ own protection to read it as a fault and disconnect the load. The allowable ramp at any real site is set by its specific fleet and protection settings.[10]
A supervisory controller coordinates these four systems. The industry calls it a microgrid controller or energy management system. It dispatches the fast machines and the storage against transients, sequences generation, and coordinates fault handling across the fleet. Caterpillar, Tesla, and INNIO Jenbacher each sell these coordinators. There is also a standard framework for proving one works: IEEE 2030.7 specifies what a microgrid controller must do, and IEEE 2030.8 defines the disturbance tests. The framework demonstrates performance only when a project publishes its test results.[11]
When Kauai Island lost a plant carrying 60% of its load, batteries at four of its plants responded within 50 milliseconds and arrested the event, with under 4% of customer load shed[4]. This demonstrates that properly commissioned coordination works.
The coordination process is more challenging for onsite “islanded” power. There is no larger network lending stability, so frequency moves fast and far. Fault current is scarcer, because inverter-based sources feed as little as 1.1 times their rated current into a short circuit. Protection built for utility fault levels can become insensitive to a real fault, as documented by Sandia National Laboratories a decade ago.[12] Power-electronics-dense load distorts a generator’s voltage waveform far more than it distorts a utility’s as it heats the machines from the inside. Resolving the distortion requires engineering rather than oversizing. None of this is exotic, and all of it is verifiable:
1. Largest-unit-loss calculation,
2. Battery specification with C-rate and duration,
3. Controller acceptance test against defined disturbances.
Our review of public sources through July 2026 located none of those documents for the projects we track.
The computer load is part of the reliability system
The largest operators are using batteries, rack-level controls and workload scheduling to make the computer load respond when power supply changes. Public evidence shows five approaches.
Plant-scale battery smoothing. Colossus combines grid power, onsite generation and Tesla Megapacks. Tesla reports that its grid-forming Megapack controls can cut high-frequency load variability by more than 70% in selected tests, and identifies Colossus as a Megapack deployment[13]. The tests demonstrate load shaping. The campus response to the loss of its largest running turbine remains undisclosed.
Rack-level storage and power controls. NVIDIA reports up to a 30% reduction in measured peak input from rack-level storage and power controls in a specific rack test[9].
Workload scheduling and demand response. Google protects latency-sensitive services by throttling lower-priority batch work and uses advance notice for demand response[14]. These systems cover rack-level variability and scheduled reductions. Evidence for a no-notice campus event remains limited.
Failure testing and recovery drills. Meta has de-energized large production regions to test batteries, failover and recovery[15].
Geographically distributed training. Google DeepMind trained a model across four U.S. regions using an asynchronous research architecture[16]. These demonstrations move reliability into software and network design, while consuming spare capacity and recovery time.
A common software layer runs through the workload side of these approaches. The cluster scheduler, the software that assigns work to GPUs, coordinates workload placement and curtailment alongside the independent rack, UPS, battery, protection and plant control systems.
xAI, now part of SpaceX, stated publicly in November 2023 that it built its Grok training stack on Kubernetes, the industry-standard container-orchestration platform. Whether the current Colossus stack retains that architecture is undisclosed[17].
Google fulfills its utility demand-response agreements, which now cover a gigawatt of contracted demand-response capacity, by having its schedulers limit or shift machine-learning work.[18] The scheduler is also where flexibility gets priced.
Microsoft’s published research shows its LLM training clusters ran within about 3% of their provisioned power limit. In separate Microsoft tests on earlier-generation GPU servers, frequency capping cut peak server power 22% at about a 10% performance cost.[19]
Google reports that a 1% improvement in delivered training time on a 20,000-GPU job is worth more than a million dollars.[20]
Scheduler-driven flexibility is measurable, and utilities are already buying it. The implication for Firm Compute is direct: work the scheduler sheds or shifts to protect the site reduces delivered compute the same way a power failure does, so Firm Compute must account for both.
These systems lower the probability of a visible outage. Their cost appears elsewhere: reduced training output, spare GPUs, delayed runs, network capacity or customer remedies. Ordinary cluster reliability already absorbs recovery capacity.
Meta reported 419 unexpected interruptions in 54 days while training Llama 3 on 16,384 GPUs, roughly one every three hours, nearly all from routine hardware and software faults rather than site power, with automated recovery holding effective training time above 90% [21].
A campus power event would draw on the same recovery capacity. Internal training generally offers more flexibility than contracted third-party compute and latency-sensitive inference. Investors need a clear allocation of which workloads absorb the event and which obligations remain exposed.
Firm Compute is Undisclosed
Installed generation and installed GPUs are inputs. The commercial output is compute that remains deliverable through a defined power event, Firm Compute. It should be measured by duration because each protective system expires at a different time: batteries cover seconds or minutes, spare generation covers longer events, and workload recovery governs what happens after curtailment. A single uptime percentage can hide the point at which a technical derate becomes a customer shortfall. Published Model FLOPS Utilization figures, roughly 40 to 55% in the best documented large training runs, measure computational efficiency during useful work. Availability, goodput and commercially delivered GPU-hours are separate quantities, and none of the four is disclosed for the projects we track.[22]
That shortfall lands differently across the capital stack. The customer loses GPU-hours or waits for recovery. The lender and insurer see revenue interruption and a test of the contingency case. Public shareholders see installed capacity produce less billable output. The utility sees a large load disappear, island or return abruptly. The first minute, first hour and first day may land on different stakeholders.
The disclosure gap across twelve projects
Occam Edge built a failure-absorption tracker across twelve onsite-power projects representing about 10.6 gigawatts of announced capacity. The census is scoped by architecture, not by size: it covers campuses where onsite or behind-the-meter generation is the primary serving supply, not standby backup.[1] Campuses served by dedicated utility-built plants sit outside it, including Meta’s Hyperion in Louisiana, where Entergy is building roughly 7.5 gigawatts of generation on the utility side of the meter. Crusoe’s Microsoft campus at Abilene is the clearest public benchmark: 900 megawatts of generation against two 336-megawatt critical buildings, a 1.34 nameplate-generation-to-critical-IT ratio, before facility overhead, heat derate, unit availability or largest-unit size.[23] Even there, the largest-unit-loss calculation, battery duration, controls acceptance test and workload-recovery plan are not public. The other eleven disclose less. Their resilience remains unverified from public evidence.
[1] The test is the role of onsite generation, not its presence. Nearly every large campus carries onsite backup generation, typically rated for ISO 8528-1 standby duty: limited-hour coverage of interruptions to a normal utility source, with the utility carrying reserve, fault current, inertia and operating discipline. The twelve census projects run onsite generation as the primary serving source, in prime or continuous duty, so the campus carries those responsibilities itself. Hybrid campuses that lean on both the grid and load-bearing onsite generation are included where the onsite share is material to reliability. Grid-disturbance risk, the NERC and ERCOT material later in this briefing, applies to large campuses regardless of architecture.
What the record shows at Colossus 2
One project in the census shows why the disclosure gap matters in practice. SpaceX’s Colossus complex in the Memphis area is one of the largest onsite power buildouts we track, and its permit docket is the deepest public record available on any of the twelve projects. In March 2026, Mississippi regulators approved permits covering 41 gas turbines for the power island SpaceX is building in Southaven to feed Colossus 2: three turbine models from two vendors and 1.24 gigawatts behind the meter. The permit describes the authorized plant. The full docket, obtained from the same regulators, shows a materially different plant operating today.[2]
The records show Colossus 2 running today on sixty trailer-mounted rental turbines totaling 1,445 megawatts, more capacity than the permanent plant the permit authorizes. The rental fleet spans three manufacturers and six model variants, from 13-megawatt Solar units to four vintages of GE TM2500 and Mitsubishi FT8s. For 46 of the 60 units, the filed inventory lists demineralized-water injection as the emissions control and none of the catalytic controls required of the permanent fleet.[2] Eighteen units arrived in August 2025, before the permit application was filed. Thirty-three arrived after the permit issued. Construction notices show 24 of the 41 permanent turbines started between late April and early July, so the permanent plant is rising around a live gigawatt load served by rented machines.
The disclosed capacity fees attach to Colossus and Colossus 2 together; the public record does not show which billed GPUs the Southaven units serve. What the record does show is an interim fleet whose reserve against actual load cannot be determined, whose six-variant mix may increase integration and maintenance burden, and which runs on a countdown: Mississippi’s portable exemption applies only while a unit has been on site less than twelve months. The first units reach that limit on August 1, 2026, and eighteen units totaling 374 megawatts reach it during August. Removing, replacing or re-permitting a quarter of the operating fleet by defined dates is required for compliance. Nothing in the record shows either fleet, interim or permanent, carrying a representative full-load contingency test. Colossus 2 is the one project where the record allows this comparison, and the comparison shows announced configurations are an unreliable guide to operating reality.
Reliable islanded power has a long track record
Industry has run islanded onsite power for critical loads for decades, and the credible operators leave visible track records.
EDL’s microgrid at the Agnew gold mine reports 99.99% reliability while drawing half its energy from renewables, built five technologies deep with a thermal backbone underneath.[24]
At the Fekola mine, the battery was sized for one specific job: letting three of six generators shut down safely instead of idling.[25]
Gruyere’s owners commissioned their island in stages, thermal first, renewables layered on afterward.[26]
Even rental turbines have a long-duty record when they are managed for it. TM2500 units serve Puerto Rico’s grid today under FEMA emergency-generation authorizations running to 2027, a program that began after Hurricane Maria in 2017, with the emissions and permitting tail explicitly planned and funded.[27]
The track record of success is consistent: deep redundancy, staged commissioning, tested controls, published performance. Power market institutions now demand the same evidence.
NERC documented 1,500 megawatts of data-center load tripping itself off during a routine transmission fault, escalated to its highest alert tier this spring after the pattern kept repeating, and has opened a separate project to draft mandatory standards. The alert itself carries no penalty-backed obligation.[28]
ERCOT now advises grid planners, where ride-through is unproven, to model data-center load as tripping when voltage sags below 0.75 per unit for at least 20 milliseconds.[29]
The burden of proof has already moved. What successful operators publish voluntarily and regulators now require is the same evidence our analysis finds missing at eleven of twelve projects.
The most likely cost is a reliability tax: lower training output, checkpoint recovery, power caps, spare GPUs, additional turbine wear and customer make-up. These losses can accumulate without taking the entire campus dark. The public record contains no headline campus outage. The operating history is short, mitigation may already be absorbing events quietly, and operators rarely publish derates or lost compute output. Those conditions make the absence of reported failures weak evidence. They also will not last. The permit requires operating reports for the permanent fleet that count every startup and shutdown and time how long each unit runs before its controls come up, and the portable fleet’s twelve-month clocks surface in the same docket. This phase is auditable on a schedule.
The risk reaches contracts and the grid
SpaceX’s Anthropic agreement turns the engineering risk into a contract-allocation issue. The prospectus discloses pricing and termination rights.[1] Treatment of curtailed compute, make-up capacity and service shortfalls remains undisclosed. Those provisions decide whether a turbine trip stays inside operating expense or reaches revenue and customer retention.
Grid authorities need the same operating evidence for a different reason. A site that drops or restores hundreds of megawatts without warning can create a disturbance beyond the campus fence. NERC requests data on ramp rates, workload composition, UPS settings, onsite generation and disturbance behavior.[28] Dominion asks how long batteries and backup generation can carry the load, whether the site can island, how much server load can move elsewhere and how many seconds the transfer takes.[30] That checklist is also a practical template for investor diligence because it exposes the duration, recovery and concentration risks that component counts miss.
Two limits to our analysis deserve equal prominence:
Counterparties may hold nonpublic evidence: owners, customers, lenders, insurers and equipment vendors can possess the one-line diagrams, acceptance tests, models and operating telemetry that answer these questions privately, and security or competitive concerns can justify keeping some of that material unpublished.
Most of the assessed twelve projects are also pre-operational, so missing operating records are partly a lifecycle effect.
Neither limit closes the gap for outside investors: evidence that cannot be examined cannot support a valuation.
What underwriting should require
Eight records connect the engineering design to the commercial obligation.
The power case
Loss-of-largest-unit calculation, with a unit-availability sensitivity.
Battery power, duration, and ride-through specification.
Controls acceptance test at representative full load.
The compute case
Workload curtailment by service class and response time.
Checkpoint, restart, spare-compute, and geographic-transfer plan.
Contracted availability, make-up rights, and termination triggers.
The bridge case
Interim fleet composition, swap schedule, and the reserve margin held through unit rotations.
The operating-mode reports already filed with regulators, which count starts, durations, and control warm-up for every unit.
Together, these records show how a shortfall travels from turbine to contract. The power case establishes the immediate megawatt deficit and the bridge available to cover it. The compute case establishes which workloads can move or pause and which customer obligations remain firm. That evidence supports contingency sizing, covenant design, business-interruption pricing and commissioning gates before financial close. The bridge case is the easiest to obtain: the cycling record already exists as a regulatory filing.
Over a multi-year operating life, unit trips are expected design-basis events. Whether a trip reaches contracted compute depends on design and contract choices that are undisclosed in the public record. Occam Edge is tracking Firm Compute project by project. Dark Gigawatts are avoidable with the right system design decisions. This time could be different if stakeholders demand that those decisions come to light.






Google, Amazon, Microsoft, and Meta have been operating data centers for decades and mostly have serious programs for generation. It is hard for me to imagine newer entrants to understand the rigor that you are calling out (as missing for the best-in-class operators).
With that said, Batch Zero and its Provisional Controllable Load Resource program will force most entrants to deal with grid load allowances that change every five minutes, for years until the actual grid gets built out. Folks are going to have to learn fast how to manage it.
Wow, amazing work. Your FOAK engine approach is a powerful tool on durability.