Inference accounts for 80-90% of total AI compute consumption. Unlike training workloads, inference runs 24/7 at low latency — requiring firm, sustained power rather than burst capacity. 

This is the number that matters. The trillion dollar bet isn’t on training. It’s on inference.

Training is periodic — weeks or months, then done. Inference is forever. Forever is too long.

The Technologist had seen the SpaceX IPO number.

Unimpressed, he returned to the problem at hand.