Back to Blog

The Compute Squeeze: Why Compute Won't Get Cheaper

<p>In March 2026, we tried to run a 24-hour AI inference workload for a new feature. Our cloud provider throttled us after 11 hours. The reason wasn&#39;t chips. It was power.</p> <hr> <p>That same week, <a href="https://www.nscale.com/press-releases/nscale-acquires-american-intelligence-power-corporation" target="_blank">Nscale</a> announced it had acquired the Monarch Compute Campus in West Virginia. America&#39;s first state-certified AI microgrid, with up to 8 gigawatts of potential onsite power. Eight nuclear plants&#39; worth of electricity, dedicated to AI.</p> <p>Jensen Huang has been calling these facilities &quot;AI factories&quot; for over a year. The Monarch acquisition is the first time one of those factories owns its own grid.</p> <p>The deal was about securing power supply before the squeeze hits, not about cheaper AI for the rest of us.</p> <hr> <h2>Microsoft&#39;s 1.35GW Bet on Vera Rubin</h2> <p>Nscale <a href="https://www.prnewswire.com/news-releases/nscale-acquires-american-intelligence--power-corporation-creating-a-full-stack-ai-hyperscaler-integrated-from-energy-to-compute-302715207.html" target="_blank">signed a letter of intent with Microsoft</a> for 1.35 gigawatts of NVIDIA Vera Rubin NVL72 GPUs. One of the largest GPU commitments in history.</p> <h3>Vera Rubin, Briefly</h3> <p>NVIDIA&#39;s next-generation AI chip architecture, named after the astronomer who discovered dark matter. The NVL72 configuration packs 72 GPUs into a single system linked by NVLink and hits 1.4 exaflops of AI performance. One rack can train models in days that took months on the previous generation.</p> <h3>Microsoft Isn&#39;t Buying GPUs for Fun</h3> <p>They&#39;re securing capacity because they see what&#39;s coming: demand for AI compute that current infrastructure cannot satisfy. The <a href="https://www.nscale.com/press-releases/nscale-acquires-american-intelligence-power-corporation" target="_blank">deployment starts in late 2027</a>. The companies sitting on compute capacity in 2028 will set the terms for everyone renting it.</p> <hr> <h2>The Memory Numbers Worth Tracking</h2> <p>While Nscale was announcing their acquisition, <a href="https://www.tomshardware.com/pc-components/dram/micron-enters-high-volume-production-of-hbm4-for-nvidia-vera-rubin" target="_blank">Micron entered high-volume production of HBM4</a>. The RAM that powers these AI supercomputers.</p> <table> <thead> <tr> <th>Metric</th> <th>HBM3E (Previous)</th> <th>HBM4 (New)</th> <th>Improvement</th> </tr> </thead> <tbody><tr> <td>Bandwidth</td> <td>~1.2 TB/s</td> <td>2.8+ TB/s</td> <td><strong>2.3x faster</strong></td> </tr> <tr> <td>Power Efficiency</td> <td>Baseline</td> <td>20% better</td> <td>Less heat, less cost</td> </tr> <tr> <td>Stack Capacity</td> <td>24GB</td> <td>36GB</td> <td>50% more memory</td> </tr> </tbody></table> <p>Memory bandwidth is the hidden bottleneck in AI training. You can have the world&#39;s fastest GPU, but if you can&#39;t feed it data fast enough, it sits idle.</p> <p>HBM4 changes that. Training runs that took weeks now take days. Models can be larger. Inference becomes cheaper at scale, for the companies that have access.</p> <p>For SaaS founders the question flips from price to availability. Can you get the compute at any price at all?</p> <hr> <h2>Energy Is the Real Constraint</h2> <p><a href="https://www.nscale.com/press-releases/nscale-acquires-american-intelligence-power-corporation" target="_blank">Nscale didn&#39;t just buy a data center</a>. They bought a power plant.</p> <p>The Monarch Compute Campus is certified as an AI microgrid. It can generate its own power, up to 8GW. Data center power is now the bottleneck on AI expansion, and owning your own electrons is the only way to guarantee supply.</p> <h3>The Power Crunch Is Already Here</h3> <p>Northern Virginia, the world&#39;s largest data center market, is running out of electricity. Major AI training facilities are being delayed because there isn&#39;t enough power available. Some regions have multi-year waitlists for new capacity.</p> <p>Power plants take 5-10 years to build. AI demand is growing faster than the grid can expand. The companies that secured energy supply now (Nscale, Microsoft, Meta) will have compute to sell. Everyone else will be bidding for scraps.</p> <h3>More Infrastructure Doesn&#39;t Mean Cheaper Compute</h3> <p>Demand is inelastic. AI training runs consume whatever power is available, so when 50+ gigawatts of new data center demand hits a grid that can&#39;t expand fast enough, prices spike instead of drop.</p> <p>The cheap power doesn&#39;t stay cheap either. Nscale&#39;s West Virginia site has affordable power today, but as data centers cluster near any energy source they bid up local prices. Cheap gets expensive fast.</p> <p>And there&#39;s the opportunity cost: that 8GW could power roughly 6 million homes. As data centers eat more of the grid, residential and commercial rates rise to clear the market, or governments step in with restrictions. Cheap, abundant compute is not arriving on schedule.</p> <p>Nscale locked down a resource that will only get scarcer and more expensive, and tied it to Microsoft. That&#39;s the real play.</p> <hr> <h2>What We&#39;re Doing About It</h2> <p>We learned this the hard way. When our inference workload got throttled in March, we had to move the feature behind a queue and cap usage per customer. The cost didn&#39;t break us. The unavailability did. Our budget was fine. The capacity simply wasn&#39;t there.</p> <p>That experience changed how we plan. Three things we&#39;re doing now:</p> <ol> <li><p><strong>Cost-modeling for scarcity, not abundance.</strong> We stopped assuming inference costs will drop 10x. Our 2026 projections use $0.15 per query as a floor, not $0.10. If energy prices spike, we have headroom. When they drop, we pass savings to customers.</p> </li> <li><p><strong>Building for graceful degradation.</strong> When GPU capacity runs short, our AI features fall back to cached or pre-computed responses instead of failing. The March throttle taught us that availability matters more than latency.</p> </li> <li><p><strong>Auditing our dependencies.</strong> We mapped every AI feature to its provider and region. Two of them route through the same Northern Virginia cluster that&#39;s already capacity-constrained. We&#39;re moving one to a different region before it becomes a problem.</p> </li> </ol> <p>The mistake we made was treating compute like electricity. We assumed it would always be there when we needed it. Compute is closer to water in a drought. You don&#39;t notice the constraint until it&#39;s your tap that runs dry.</p> <hr> <h2>The Tiered Market Ahead</h2> <p>Cheap AI for everyone may not arrive. What&#39;s coming instead is a tiered market: hyperscalers with secured energy supply offer reliable (if not cheap) compute, and everyone else gets rationing and rising costs.</p> <p>If you run a SaaS, don&#39;t assume AI compute gets cheaper. A business model that needs 10x cost drops to work is a bet on abundance that may not pay off. The constraint that decides who wins is energy supply.</p> <p>We&#39;re still figuring out what this means for our product roadmap. One open question we haven&#39;t resolved: should we pre-pay for reserved GPU capacity 18 months out, or stay flexible and risk being shut out during peak demand? I don&#39;t have a clean answer. Neither does anyone I&#39;ve asked.</p> <hr> <p><em>If this was useful, we publish one post a week on building software when the constraints are real. No &quot;10 AI tools that will transform your workflow.&quot; Just what we&#39;re learning and the mistakes we&#39;re making. <a href="https://thinkcodeship.com" target="_blank">Subscribe here</a>.</em></p>