Jim Chanos has a point about CoreWeave.

The famed short-seller recently argued that neocloud moats are “both dug and filled” by Nvidia. His reasoning is straightforward: Nvidia supplies the GPUs that neoclouds rent to customers, influences their allocation and pricing, and possesses the ability to weaken or bypass the intermediaries whenever it chooses. Chanos went so far as to call neoclouds “financial conduits, not technology companies.”

That criticism identifies a genuine vulnerability. But it may understate the larger problem.

Nvidia does not need to turn against CoreWeave or any other neocloud for their original business model to come under pressure. The success of AI itself will do that.

Neoclouds were created by the scarcity of AI compute. When Nvidia GPUs were nearly impossible to obtain, possessing them at scale was a meaningful differentiator. Customers paid a premium for immediate access, and companies capable of financing, installing and operating large GPU clusters looked as though they had developed durable technology moats.

But scarcity is not a moat if the entire industry is spending trillions of dollars to eliminate it.

That is the neocloud version of the indispensability trap. The infrastructure becomes essential. Capital floods in to expand it. Production increases, alternatives emerge, efficiency improves and access becomes more widely distributed. The infrastructure may become vastly larger and more important even as the scarcity premium enjoyed by its original providers begins to erode.

Jevons Paradox Meets the Indispensability Trap

This is not an argument that demand for AI compute will decline. I believe the opposite is much more likely.

As AI becomes cheaper and more efficient, we will use far more of it. Tasks that made no economic sense at $20 or $30 per million tokens become obvious candidates for automation at a fraction of that cost. Agents will perform multiple steps where a person previously issued a single prompt. AI will analyze more documents, generate more software, process more video and operate continuously inside business workflows.

That is Jevons paradox in action. Greater efficiency makes a resource cheaper to use, which encourages enough new uses that total consumption increases rather than decreases.

We received a particularly timely example after OpenAI reduced the price of GPT-5.6 Luna by 80% and Terra by 20%. TD Cowen analyzed usage data from OpenRouter and found that Luna’s effective usage price fell approximately tenfold while consumption increased roughly fourteenfold. Estimated Luna revenue still rose 34%. Terra’s effective price fell approximately threefold, usage increased fivefold and estimated revenue rose 45%. The observation covered only about two weeks, so it is far too early to treat those numbers as a settled long-term elasticity measurement. But the initial pattern is striking: Prices fell, consumption soared and aggregate revenue increased.

Jevons paradox explains why AI compute demand will keep increasing. The indispensability trap explains why the companies supplying that compute may nevertheless capture a declining share of the value it creates.

The latest Futurum data puts numbers behind both sides of that argument.

Futurum’s 1H 2026 Data Center Semiconductors Market Sizing & Five-Year Forecast projects the overall data-center semiconductor market growing from $241 billion in 2025 to approximately $1.2 trillion in 2030 under its base case. GPU revenue alone is forecast to increase from $157.6 billion to $604.5 billion.

That is not a demand slowdown. It is an explosion.

Custom silicon, however, is forecast to grow even faster. Futurum expects XPU revenue—which includes purpose-built accelerators such as Google TPU, AWS Trainium, Meta MTIA and Microsoft Maia—to increase from $37.4 billion in 2025 to $237.2 billion in 2030.

Within combined GPU and XPU accelerator revenue, merchant GPUs are forecast to decline from approximately 80.8% of the market to 71.8%. XPU share rises from 19.2% to 28.2%. That represents approximately nine percentage points of forecast share migration from merchant GPUs to custom accelerators, even as the total accelerator market becomes several times larger.

This is the trap expressed quantitatively. GPUs may sell in vastly greater numbers while becoming less singularly indispensable.

The current market remains exceptionally concentrated. Futurum’s quarterly data shows Nvidia holding approximately 95% to 96% of merchant data-center GPU revenue throughout 2024 and 2025, ending the fourth quarter of 2025 at approximately 95.5%. That supports Chanos’s immediate point. Neoclouds remain downstream from a supplier that controls almost the entire merchant market for their most important asset.

But the five-year forecast shows the ecosystem surrounding that dominant position becoming more diverse. In Futurum’s 1H 2026 Data Center Semiconductor Decision Maker Survey, 824 respondents anticipated increasing XPU spending by an average of 39.96% in 2026, compared with 35.33% for GPUs and 30.91% for CPUs. XPU spending starts from a much smaller base, but it is growing faster.

The more AI succeeds, the more capital will pour into alternatives to its most expensive bottlenecks.

More Compute, in More Places

The same pattern is visible in where enterprises expect to run AI.

In Futurum’s 1H 2026 AI Platforms Decision Maker Survey, 736 respondents were asked to select their primary generative-AI deployment environments. Because multiple selections were permitted, the results describe a hybrid architecture rather than a winner-take-all market.

Managed cloud was selected by 63.86% of respondents. SaaS-embedded AI followed at 42.66%, private VPC at 40.08%, on-premises infrastructure at 30.71%, neoclouds at 18.61%, and edge or device deployment at 17.66%.

This is a single-period snapshot, so it would be wrong to claim that it proves a historical “shift” away from neoclouds. It does show that neoclouds are one part of a distributed AI infrastructure market. They are not the inevitable destination for every workload.

A separate Futurum question asking respondents for their single primary AI workload location found public cloud leading at 40.66%, but enterprise-owned data centers were close behind at 35.56%. Colocation represented 13.47%, bare-metal or HPC environments 5.58%, and edge deployments 4.73%.

None of this suggests that phones, PCs or enterprise systems will replace centralized AI factories. Frontier training, the largest reasoning systems and many high-volume inference workloads will continue requiring enormous centralized clusters. But inference will be placed wherever privacy, latency, economics, data gravity and available hardware make the most sense.

Smaller models are already part of that equation. Futurum found small language models such as Phi and Nano in production at 24.73% of surveyed organizations, with DeepSeek at 19.84%. Those numbers do not tell us how many GPU-hours have been displaced, but they establish that enterprises are using alternatives to the assumption that every AI task requires the largest available model running in a centralized cloud.

The hardware is advancing just as rapidly. Epoch AI estimates that AI-chip performance per dollar improved approximately 37% annually across its longer historical series. A newer analysis estimates that the average dollar spent on AI chips has purchased approximately 49% more performance each year since 2023. Nvidia says its Rubin platform can reduce inference token costs by as much as tenfold and use four times fewer GPUs to train certain mixture-of-experts models than Blackwell.

Nominal GPU-hour prices do not have to fall in a straight line for the trap to operate. Prices can remain high or temporarily increase when demand outruns supply. What matters is the cost of accomplishing a comparable unit of useful work. A newer accelerator can cost more per hour while still producing much cheaper tokens, faster training and greater performance per watt.

We will consume more compute. Each unit of compute will accomplish more. Both can be true.

CoreWeave Is Protected—For Now

CoreWeave is the natural case study because its latest results illustrate both the strength of present demand and the financial consequences of supplying it.

CoreWeave reported second-quarter 2026 revenue of $2.575 billion, more than double the prior-year period. Its adjusted EBITDA reached $1.51 billion, representing a 59% margin. But adjusted operating income was only $128 million, a 5% margin, and adjusted net loss reached $567 million. Its GAAP net loss attributable to common stockholders was $626 million. Net interest expense alone was $640 million.

The company had $103.7 billion in unsatisfied remaining performance obligations as of Jun 30, 2026. Forty-one percent is expected to become revenue within 24 months, another 39% during months 25 through 48, and the remaining 20% during months 49 through 78. CoreWeave also said that figure excluded more than $25 billion in net new customer commitments added during early Q3.

Those contracts matter. Anyone expecting CoreWeave’s facilities to empty next quarter is arguing against the available evidence. The company has secured substantial near-term utilization and revenue visibility.

But it has also accumulated approximately $35.6 billion in debt principal and $46.7 billion in net property and equipment. Depreciation and amortization reached $1.4 billion during the second quarter. CoreWeave depreciates technology equipment over six years, while its weighted-average operating-lease term has reached approximately 12 years.

Customer concentration remains substantial. CoreWeave’s three largest disclosed customers accounted for approximately 72% of quarterly revenue. Its filing also says that all GPUs currently deployed in its infrastructure are Nvidia GPUs because customers contractually specified Nvidia hardware.

This does not establish that CoreWeave’s loans broadly outlast its customer contracts. Public disclosures do not permit contract-by-contract matching of every debt facility with its associated customer agreement, and we should not pretend otherwise.

The more important question is what happens when those contracts renew.

What will five- or six-year-old technology equipment earn when customers can obtain far more performance per dollar from newer Nvidia systems, AMD accelerators, custom silicon or infrastructure they control themselves? Can older GPUs be redeployed profitably? Can CoreWeave finance continuous fleet refreshes while servicing its debt and carrying longer-lived facilities, leases and power commitments?

Margin compression alone does not destroy an infrastructure company. Margin compression combined with leverage, rapid technological depreciation and continuous capital requirements can.

Moving Up the Stack Before the Premium Disappears

CoreWeave is not standing still. Its investments in Mission Control, SUNK and SUNK Anywhere show that it understands its long-term value cannot rest exclusively on access to Nvidia GPUs.

Mission Control adds observability, security, automated remediation and operational expertise. CoreWeave says the platform can deliver as much as 96% training goodput and 20% higher model utilization. SUNK Anywhere is especially revealing because it extends CoreWeave’s orchestration system to customer infrastructure outside the CoreWeave cloud. The software is beginning to separate from ownership of the underlying compute.

The same pattern is visible across the category. Nebius has introduced a model that allows infrastructure partners to deploy its AI-cloud software stack in their own data centers. Nebius explicitly describes it as a high-margin revenue stream requiring minimal incremental capital.

Nscale is seeking to acquire Anyscale, extending its offering from power and data centers through the Ray-based software platform used to run production AI. IREN is adding software, orchestration and support capabilities while emphasizing the value of its secured power portfolio. Lambda now sells managed Kubernetes, Slurm, orchestration, hybrid-cloud integration and engineering services along with GPU capacity.

These are rational attempts to move up the stack. Whether they create durable moats remains unanswered. Much of the software is still adjacent to Nvidia and CUDA. Futurum’s chipsets survey found that 53.88% of respondents were at least somewhat likely to switch their primary accelerator vendor within 24 months. The leading factor that could persuade them to switch was the software ecosystem, selected by 50%.

That suggests the deepest lock-in may still belong to the chip and software-platform provider, not the company renting access to its hardware.

Scarcity May Move Rather Than Disappear

The strongest counterargument is that even if GPUs become more plentiful, power, grid connections, HBM, advanced packaging, cooling and high-performance networking will remain scarce.

That is almost certainly true for some time.

The International Energy Agency reported that data-center electricity demand increased 17% in 2025. It expects total data-center electricity consumption to double by 2030 and AI-focused data-center electricity use to triple. Grid constraints could delay approximately 20% of planned global data-center capacity through the end of the decade.

Futurum’s survey provides additional perspective. Accelerator availability was identified as the leading current limit on cluster expansion by 32.65% of respondents. Memory and storage followed at 23.91%, with grid interconnection and utility capacity at 16.63%. In a separate question, budget and capital expenditures were the leading general scaling constraint at 23.30%, followed by networking lead times at 16.63% and power and cooling availability at 14.93%.

Power is a genuine constraint, but it is not the only one, nor does the current survey establish it as the single dominant bottleneck.

If power, land, cooling and grid access become the neoclouds’ most durable advantages, they may preserve considerable value. They may also look increasingly like specialized utilities, data-center operators, REITs or infrastructure-finance companies rather than software companies.

Those can be enormous, essential and highly profitable businesses. The indispensability trap does not require bankruptcy. It means that becoming essential can constrain extraordinary returns as competition, substitution and public dependence reshape the market around the infrastructure.

Neoclouds may sell more compute than ever while earning less of a premium for each unit they provide. Total revenue can increase even as unit economics, margins and strategic leverage deteriorate.

Today’s scarcity gives neoclouds a window. It does not guarantee them a permanent moat.

The winners will use the capital, customer relationships and operating expertise generated during this scarcity period to build software, orchestration, managed services, distribution and workflow integration. The losers will confuse temporary possession of a scarce asset with lasting control of the market.

We will consume vastly more AI compute. We simply will not continue paying scarcity prices for undifferentiated access to it.

The Indispensability Trap is Alan Shimel’s forthcoming book, due in September 2026. It examines how the companies that build society’s essential infrastructure often enable greater fortunes above them while seeing their own returns constrained as that infrastructure becomes indispensable.