HomeArtificial IntelligenceAI GovernanceREPORT - The Price of Intelligence: AI Compute Scarcity, Hardware Inflation and...

REPORT – The Price of Intelligence: AI Compute Scarcity, Hardware Inflation and the Emerging Global Access Divide

Introduction — From Universal Digital Service to Rationed Computational Capability

Master Abstract

Artificial intelligence is frequently presented as an infinitely replicable software technology, yet every economically useful output depends on a finite and unusually concentrated physical system: leading-edge semiconductor fabrication, advanced packaging, high-bandwidth memory, accelerators, networking equipment, data-centre capacity, electricity, cooling infrastructure and specialised engineering labour. The resulting market cannot be evaluated through consumer subscription prices alone. Its true structure begins upstream, where technical and financial barriers restrict the number of firms capable of manufacturing the most advanced components, and continues downstream, where a small group of hyperscalers and model developers determine which users receive access, at what performance level, with which latency, context window, privacy safeguards and contractual limitations. Evidence available in August 2026 confirms an extraordinary expansion of demand and investment, but it does not yet prove that scarcity has been artificially engineered or that unlawful pricing conduct has occurred. NVIDIA reported fiscal-year 2026 revenue of $215.9 billion, an increase of 65%, with a fiscal-year GAAP gross margin of 71.1%; in the quarter ended 26 April 2026, Data Center revenue reached $75.2 billion, 92% above the corresponding prior-year period. These results demonstrate exceptional value capture in a supply-constrained technological segment, but margins and revenue growth must subsequently be decomposed into innovation rents, intellectual-property returns, scarcity rents, product-mix effects and possible market power. NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026 – NVIDIA – February 2026 NVIDIA Announces Financial Results for First Quarter Fiscal 2027 – NVIDIA – May 2026

The upstream constraint is not reducible to GPU design. AI accelerators require leading-edge logic, large quantities of HBM, complex interposers and advanced packaging technologies capable of connecting logic dies with multiple memory stacks at extreme bandwidth. TSMC reported that the annual capacity of facilities managed by the company and its subsidiaries exceeded 17 million 12-inch-equivalent wafers in 2025, while its official 2025 first-quarter investor transcript stated that the company was working to double CoWoS capacity during that year in response to customer demand. TSMC has separately announced that a 9.5-reticle-size CoWoS implementation planned for volume production in 2027 is intended to integrate twelve or more HBM stacks in one package. These disclosures show how frontier AI performance increasingly depends on co-optimisation across fabrication, memory and packaging, rather than on an isolated processor. Memory producers are consequently moving investment toward higher-value AI products: Micron projected in December 2025 that the HBM total addressable market could rise from approximately $35 billion in 2025 to around $100 billion in 2028, while Samsung announced plans to invest more than KRW 110 trillion in facilities and research and development during 2026, explicitly including HBM, foundry production and advanced packaging. These figures establish strong demand and rapid capacity formation, but they do not by themselves demonstrate permanent scarcity. The central analytical problem for this report will therefore be whether additions to fabrication, packaging and memory capacity can outrun the combined growth of training, inference, multimodal services, agentic workloads and national sovereign-compute programmes. TSMC 2025 Annual Report – Taiwan Semiconductor Manufacturing Company – April 2026 TSMC First Quarter 2025 Earnings Conference Transcript – Taiwan Semiconductor Manufacturing Company – April 2025 TSMC Unveils Next-Generation A14 Process at North America Technology Symposium – Taiwan Semiconductor Manufacturing Company – April 2025 Micron Fiscal First Quarter 2026 Financial Results – Micron Technology – December 2025 Samsung Electronics Shareholder Value Enhancement Plan – Samsung Electronics – March 2026

The downstream investment cycle indicates that today’s apparently inexpensive access to powerful models cannot automatically be treated as a stable long-term equilibrium. Alphabet reported $91.4 billion in 2025 capital expenditure, compared with $52.5 billion in 2024, and stated that it expected a significant further increase in technical-infrastructure investment during 2026; capital expenditure reached $35.7 billion in the first quarter of 2026 alone, compared with $17.2 billion in the corresponding 2025 period. This investment does not prove that consumer or API prices must rise, because higher utilisation, algorithmic efficiency, model compression, custom accelerators and competition can reduce unit costs. It does, however, establish that AI services are supported by a capital base whose depreciation, financing, electricity, network and replacement costs must eventually be recovered through advertising, subscriptions, API billing, enterprise contracts or cross-subsidisation. Energy is becoming a parallel constraint. The International Energy Agency projects global data-centre electricity consumption of approximately 945 TWh in 2030, more than twice the current level represented in its base case and just under 3% of projected global electricity consumption. The IEA further estimates average data-centre electricity-demand growth of about 15% annually between 2024 and 2030, while AI-optimised facilities account for the principal incremental driver. The report must consequently model not merely a price per token, but a multidimensional tariff in which model capability, input and output length, reasoning intensity, latency, reserved capacity, geographic location, energy availability and privacy requirements become separate price determinants. The Bitcoin analogy will be tested rather than accepted: compute can be scarce and dynamically priced, but unlike Bitcoin it is heterogeneous, depreciating, geographically constrained, continuously producible and consumed in delivering a service. Alphabet Annual Report on Form 10-K for the Fiscal Year Ended December 31, 2025 – Alphabet and U.S. Securities and Exchange Commission – February 2026 Alphabet Quarterly Report on Form 10-Q for the Quarter Ended March 31, 2026 – Alphabet and U.S. Securities and Exchange Commission – April 2026 Energy and AI: Energy Demand from AI – International Energy Agency – April 2025

For households, students, researchers and smaller organisations, the decisive divide is not between possessing and lacking nominal AI access, but between different levels of quality-adjusted computational capability. A used workstation may run a quantised model locally, yet remain inadequate for long contexts, multimodal processing, high concurrency, rapid generation or specialised scientific workloads. The RTX 3090, introduced at a starting price of $1,499, contains 24 GB of GDDR6X memory; this remains useful for defined local workloads but imposes a hard memory boundary before software overhead, context cache and concurrent inference are considered. A 2025 Mac Studio with M4 Max can be configured with 128 GB of unified memory, enabling a much larger theoretical model-memory envelope, although unified-memory capacity must not be equated with NVIDIA accelerator performance, software compatibility or effective application throughput. GeForce RTX 3090 Family Specifications – NVIDIA – September 2020 Mac Studio 2025 Technical Specifications – Apple – March 2025 The international distributional problem begins even earlier: the International Telecommunication Union estimated that 94% of people in high-income countries used the internet in 2025, compared with only 23% in low-income economies, while 2.2 billion people remained offline. Advanced AI therefore arrives on top of an unresolved connectivity, affordability, electricity, device and skills divide. Measuring Digital Development: Facts and Figures 2025 – International Telecommunication Union – November 2025 Europe’s response illustrates the strategic scale of the issue: the European Commission launched a July 2026 call intended to establish as many as seven AI Gigafactories and unlock more than €30 billion, while its architecture defines such facilities as infrastructures bringing together more than 100,000 advanced AI processors. EU Launches AI Gigafactories Call to Boost Europe’s Computing Capacity and Unlock More Than €30 Billion – European Commission – July 2026 AI Factories – European Commission – July 2026 The five-year investigation will determine whether this investment broadens access or merely relocates concentrated capacity, and whether AI becomes a productivity equaliser, a metered industrial utility or a cumulative mechanism through which capital-rich users acquire an enduring cognitive and economic advantage.

Ai Compute Access Dashboard
Integrated evidence model · 2026–2031

AI compute, affordability and access

Navigate the ten analytical layers, stress-test five scenarios and change policy intensity without converting estimates into observed facts.

Current classification
HIGH
Upward risk toward severe
Enterprise AI expenditure
300
Index, 2026 = 100
Quality-adjusted API price
32
Index, 2026 = 100
Global AI Access Gini
0.215
Central 2031 estimate
Inference demand
1,100
Index, 2026 = 100

Annual trajectory

Central, low and high paths

Selected scenarioLow referenceHigh reference
Select a point for its exact value.

Critical focal points and falsification thresholds

The classification uses the highest persistent risk dimension corroborated by at least one additional dimension.

Access Gini
0.215
Affordability
8.0%
Capability gap
28 pp
Lead time
11 mo

Evidence integrity ledger

FindingStatusConfidenceInterpretive constraint
Units: indexes use 2026 = 100; Gini is a 0–1 coefficient; affordability is annual adequate-access cost as a share of exposed-group disposable resources. Forecast values are scenario estimates, not observed facts.

The Price of Intelligence: How Compute Is Becoming the New Infrastructure of Economic Power

Artificial intelligence is no longer principally a contest between algorithms. It is a contest over access to accelerators, high-bandwidth memory, advanced packaging, electricity, data centres and the capital required to assemble them into usable systems. The decisive inequality is therefore shifting from internet connectivity to computational capability: who can train, adapt and operate advanced models; who must rent that capability; and who remains confined to weaker systems. This transition has consequences for industrial productivity, scientific sovereignty, education, defence and the distribution of economic power. Europe has begun constructing a public response, but the strategic question remains unresolved: whether AI compute will develop as broadly accessible infrastructure or as a scarce, vertically controlled service whose price and availability are determined by a small group of technology providers.

The New Economic Input

Compute cannot be classified simply as another industrial input. It is simultaneously a capital good, a metered service, a component of critical infrastructure and an instrument of geopolitical leverage. Its economic value depends not only on processor throughput but also on accelerator memory, bandwidth, interconnects, software compatibility, electricity, cooling, network latency and utilisation. A nominally powerful accelerator without adequate memory or software support may be economically inferior to a slower system capable of completing the required workload reliably.

This distinction is essential when assessing claims of hardware inflation. Sticker prices alone cannot establish whether AI capability has become more expensive. A valid comparison must determine the cost of completing the same workload under equivalent requirements for accuracy, context length, latency, concurrency, privacy and reliability. Semiconductor improvements can reduce the cost of an arithmetic operation while the total cost of deploying a useful model rises because the workload requires more memory, more energy, specialised personnel or a larger cluster.

The result is an apparent paradox: aggregate computing power can become more abundant while strategically useful compute remains scarce. Scarcity becomes especially acute when organisations require frontier-level performance, sovereign data handling, guaranteed capacity or local operation. For a student, an independent researcher or an SME, the relevant barrier is not whether a small model can technically run on a computer. It is whether that system can produce competitive work within acceptable time, quality and reliability constraints.

Adoption Without Equal Access

European enterprises are already entering this transition from profoundly unequal starting positions. Eurostat reported that 19.95% of EU enterprises employing at least ten people used one or more AI technologies in 2025, while the corresponding share among large enterprises reached 55.03%. Use of Artificial Intelligence in Enterprises – Eurostat – 2025official statistical publication.

The disparity matters because AI adoption is cumulative. Organisations with data, engineers, cloud agreements and integration budgets can test more applications, train employees, construct proprietary workflows and learn from failure. Smaller organisations must absorb licences, cybersecurity, compliance, data preparation, integration and human-supervision costs before productivity gains are certain. AI can consequently become a de facto mandatory input: firms may have to purchase it because competitors use it, even when their own return on investment remains difficult to measure.

This creates a structural asymmetry. Large companies can negotiate capacity reservations, diversify providers and finance local infrastructure. SMEs frequently buy retail access under standard contracts and carry greater switching costs relative to revenue. If model capability, latency and access priority are increasingly differentiated by price, the market will not merely divide users into subscribers and non-subscribers. It will divide them according to the quality, speed, context and reliability of the intelligence they can afford.

Europe’s Infrastructure Response

The European Commission’s AI Continent Action Plan, presented on 09/04/2025, placed physical infrastructure at the centre of European policy. The Commission identified EUR 200 billion to be mobilised for AI investment, including EUR 20 billion intended to finance as many as five AI gigafactories. It also identified a network of 19 AI Factories supporting start-ups, industry and research. AI Continent – European Commission – 09/04/2025official programme page.

That network has since been complemented by 13 AI Factory Antennas, while the European High Performance Computing Joint Undertaking states that the AI Factories offer free, customised support to SMEs and start-ups. AI Factories – European High Performance Computing Joint Undertaking – 2026official infrastructure portal. On 30/07/2026, EuroHPC launched its AI Gigafactories call, defining the planned facilities as integrated environments combining AI-optimised supercomputers, advanced data centres, high-capacity storage, high-speed networks, secure cloud access and specialist support. The EuroHPC Joint Undertaking Launches the AI Gigafactories Call – EuroHPC Joint Undertaking – 30/07/2026official call announcement.

The architecture is strategically correct because it recognises that purchasing accelerators is insufficient. Europe needs an operating ecosystem able to allocate capacity across borders, support industrial users, protect sensitive data and convert computational resources into deployable applications. The danger is fragmentation. National installations that cannot exchange workloads, use compatible interfaces or provide predictable access could reproduce at public expense the same lock-in that Europe seeks to reduce.

The Energy Constraint

AI infrastructure is also energy infrastructure. On 20/12/2024, the United States Department of Energy reported that American data centres consumed 176 terawatt-hours of electricity in 2023, equivalent to approximately 4.4% of national electricity consumption. The Department’s cited range for 2028 was 325–580 terawatt-hours, or approximately 6.7–12% of US electricity consumption. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – United States Department of Energy – 20/12/2024official publication.

These figures are not a forecast for Europe, but they expose the physical scale of the challenge. Data-centre policy cannot be separated from generation capacity, transmission networks, substations, transformers, cooling water, permitting and local-system reliability. A country may possess capital and technical expertise yet remain unable to deploy new AI capacity within commercially useful timescales because grid connections are unavailable.

Public policy must therefore distinguish between socially valuable infrastructure and private congestion costs. Data centres should contribute to the network investments required by their load, while public authorities should assess flexibility, emissions, location and the economic value of workloads. Subsidising compute without financing the electrical system that supports it would move the bottleneck rather than remove it.

Memory as Geopolitics

High-bandwidth memory demonstrates how a specialised component can become both an industrial bottleneck and a security instrument. On 02/12/2024, the United States Department of Commerce’s Bureau of Industry and Security announced controls covering 24 types of semiconductor-manufacturing equipment, three categories of software tools, HBM and 140 additions to the Entity List. BIS explicitly described HBM as critical to AI training and inference at scale. Commerce Strengthens Export Controls to Restrict China’s Capability to Produce Advanced Semiconductors – Bureau of Industry and Security – 02/12/2024official announcement.

The decision illustrates why compute cannot be treated as a globally fungible commodity. Its availability depends on origin rules, export licences, technical thresholds, fabrication geography and political alignment. Unlike Bitcoin, compute is heterogeneous, location-dependent and depreciating. Unused capacity is perishable; old hardware loses relative value; and two nominally equivalent systems can produce different economic results because of memory, software or network constraints.

The European Chips Act, which entered into force on 21/09/2023, was designed to reinforce the Union’s semiconductor ecosystem, improve supply-chain resilience and reduce external dependencies. Its stated objectives include research leadership, design, manufacturing, advanced packaging, production capacity and skills. European Chips Act – European Commission – 21/09/2023official policy framework. The strategic lesson is that leading-edge fabrication alone is not enough. Packaging, memory, substrates, testing, equipment, software and energy must be treated as an interdependent industrial system.

The Cost of Dependence

Cloud AI solves one access problem while creating another. It allows organisations to use advanced systems without owning expensive infrastructure, but it transfers control over pricing, capacity, model continuity and data-processing conditions to the service provider. Providers may differentiate prices according to input, output, context, latency, reasoning intensity, tool use, reserved capacity or service priority. This resembles utility pricing more than the sale of conventional software.

The danger is not metering itself. Metering can improve efficiency and transparency. The danger arises when customers cannot compare the full cost of equivalent workloads or move their data, embeddings, evaluation records and applications without material loss. The European Data Act, Regulation (EU) 2023/2854, has applied since 12/09/2025 and establishes requirements intended to facilitate switching between cloud and edge services and promote interoperability. Data Act Explained – European Commission – 15/12/2025official explanation.

The next regulatory step should be standardised effective-price disclosure. Providers should separate inference, storage, retrieval, networking, tool use, reserved capacity and migration charges. Public procurement should require exportable operational data, documented interfaces, model-version notice periods and tested exit procedures. Competition authorities, meanwhile, should examine self-preferencing, data access and cloud dependencies without presuming that vertical integration is inherently unlawful. The Commission’s DMA review, published on 28/04/2026, identified interoperability, potential self-preferencing, access to data and cloud dependencies among the AI-related issues raised during consultation. Commission Staff Working Document SWD(2026) 123 final – European Commission – 28/04/2026official document.

The Global Capability Divide

The AI divide begins before access to an accelerator. According to the International Telecommunication Union’s figures released on 17/11/2025, approximately six billion people were online in 2025, while 2.2 billion remained offline. The ITU estimated 5G coverage at 84% of the population in high-income countries, compared with 4% in low-income countries. Facts and Figures 2025 – International Telecommunication Union – 17/11/2025official release.

The internet-use divide is equally stark. ITU data published on 15/10/2025 recorded internet use at 94% in high-income countries but 23% in low-income economies; the corresponding figure for Africa was 36%. Facts and Figures 2025: Internet Use – International Telecommunication Union – 15/10/2025official statistical analysis.

AI can lower barriers by translating information, supporting education and distributing sophisticated analytical tools. Yet those benefits presuppose connectivity, electricity, compatible languages, payment capacity and digital skills. Where these foundations are absent, the arrival of more powerful models can widen the productive distance between countries rather than close it. The relevant development objective is therefore not the installation of prestige clusters. It is reliable, maintainable and affordable access linked to local skills, regional infrastructure and public-service demand.

A Capability Floor

The most effective response is not universal ownership of frontier hardware. It is a guaranteed capability floor combined with competitive provision. Students and researchers should receive workload-denominated compute entitlements usable across approved public, cloud and open-weight systems. Universities should pool procurement and publish utilisation, queue times and completed-workload costs. SMEs should gain access through cooperative purchasing and the AI Factory network. Privacy-sensitive institutions should be able to choose local, sovereign or hybrid deployment without accepting an automatic penalty in capability.

Hardware life also matters. Directive (EU) 2024/1799 requires Member States to apply the European right-to-repair framework from 31/07/2026. The Commission states that manufacturers of covered products must offer repair within a reasonable time and at a reasonable price, cannot use hardware or software techniques to obstruct repair, and must provide access to spare parts. Transposition, Entry into Force or Application of EU Legislation: Key Milestones – European Commission – 01/07/2026official implementation calendar. Extending useful hardware life and creating trustworthy refurbished markets can lower access costs without suppressing semiconductor innovation.

The Strategic Choice

Europe’s AI strategy will succeed only if it measures access rather than announcing capacity. The decisive indicators are not the number of installed accelerators or the nominal value of investment. They are the share of capacity reaching SMEs, universities and public-interest users; waiting times; cost per completed workload; migration success; energy availability; local-language performance; and the proportion of students and researchers able to use systems adequate for their tasks.

The emerging market will not price intelligence in the manner of a uniform commodity. It will price differentiated access to capability. Model quality, memory, context, latency, privacy and guaranteed capacity will determine what users can accomplish and what they must pay. If that architecture remains concentrated and difficult to contest, AI will magnify existing advantages in capital, education and geography. If Europe combines infrastructure, portability, competition, repair, open models and targeted access, compute can instead become the foundation of a broader productive system. The central political decision is therefore not whether to regulate or subsidise AI. It is whether advanced computational capability will remain purchasable privilege or become accessible economic infrastructure.

AI Compute Access Observatory · 2026–2031
Scarcity Transmission Engine
Interactive, source-anchored scenario model connecting compute demand, supply expansion and electricity pressure. Outputs are scenario indices, not observed market-price forecasts.
MODEL ACTIVE
Five-Year Scenario Controls
Demand index Capacity index Access-cost pressure
Dynamic Risk Matrix
Scarcity58
Demand growth relative to supply expansion.
Affordability49
Higher value indicates greater access pressure.
Energy exposure43
Sensitivity to electricity-cost escalation.
Access divide55
Composite inequality transmission indicator.
Local AI Hardware Envelope
Installed accelerator / unified memory 24 GB
Official architecture reference GDDR6X
Theoretical four-bit weight envelope ≈40B parameters
Illustrative ceiling after a 20% weight-loading allowance; context cache, runtime overhead and usable throughput reduce practical capacity.
Method: theoretical parameter envelope = installed memory × 8 ÷ 4 bits ÷ 1.20. It is not a performance benchmark and does not imply that a model of the stated size will run at acceptable speed. The RTX 3090 and Apple unified-memory architecture are not performance-equivalent. Scenario indices use 2026 = 100 and respond only to the user-selected assumptions; they must not be interpreted as audited forecasts or market quotations.

Chapter 1 — The Political Economy of AI Compute

1.1 Purpose, Scope and Analytical Position

AI compute is neither a single product nor a homogeneous quantity. It is a layered production capability created by combining semiconductor-design intellectual property, fabrication capacity, advanced packaging, high-bandwidth memory, servers, interconnects, storage, data centres, electricity, cooling, software frameworks, trained models and specialised labour. The economic object purchased by the final user is therefore not simply a GPU or a token. It is access to a time-bounded, workload-specific configuration of capital, energy, memory bandwidth, software compatibility and model capability. This distinction is essential because different forms of apparent scarcity can arise at different layers. A shortage of accelerator boards may reflect insufficient leading-edge fabrication, constrained advanced packaging, unavailable HBM, export restrictions, hyperscaler procurement, distributor behaviour or a temporary demand shock. A high API price may instead reflect training-cost recovery, inference expense, reserved capacity, latency guarantees, model scarcity, commercial segmentation or pricing power. The report will not infer unlawful market manipulation merely from increasing prices, exceptional margins or a concentrated supplier structure. It will determine whether prices exceed competitive economic benchmarks after controlling for product quality, capacity cycles, depreciation, input costs, demand growth, risk and innovation. The European Commission has independently identified data, AI accelerators, computing infrastructure, cloud capacity and technical expertise as possible barriers to entry in generative-AI markets, while cautioning that a bottleneck becomes a competition concern only in the relevant economic and legal context. Competition Policy Brief: Competition in Generative AI and Virtual Worlds – European Commission – September 2024

The principal proposition of this report is therefore conditional: if the supply of quality-adjusted AI compute expands more slowly than demand, if effective competition remains restricted at several complementary layers, and if access to superior systems produces cumulative educational or economic gains, then compute scarcity can become a mechanism of social and industrial stratification. The contrary outcome remains possible. Fabrication capacity, alternative accelerators, open-weight models, quantisation, distillation, inference optimisation, edge deployment and public compute infrastructure could lower the cost of a fixed level of capability even while the nominal price of premium hardware increases. Chapter 1 establishes the framework needed to discriminate between these trajectories. It defines six economically distinct interpretations of compute, maps the value chain, locates potential sources of scarcity and bargaining power, formalises the concepts used throughout the report and specifies eight falsifiable hypotheses. No hypothesis will be treated as confirmed in this chapter. In particular, claims concerning used-market prices, Apple or Samsung product inflation, the real cost of local AI and the sustainability of current token pricing require longitudinal datasets that will be constructed in Chapters 2, 4 and 5. The evidentiary status at the beginning of the investigation is consequently “open”: upstream capacity pressure, extraordinary investment and unequal adoption are already observable, while persistent real hardware inflation, systematic second-hand premiums, above-normal rents and future token-price escalation remain matters for formal testing.


1.2 The Six Economic Identities of AI Compute

AI compute can assume six economic identities simultaneously, but the relative importance of each identity changes according to the workload, purchaser and time horizon. For a software developer purchasing several hours of inference, compute resembles a variable industrial input. For a hyperscaler committing billions of dollars to accelerators, data centres and power procurement, it is a durable but rapidly depreciating capital good. When delivered continuously through standardised APIs with metering, capacity tiers and service-level commitments, it acquires utility-like characteristics. When governments rely on it for defence, intelligence, healthcare, energy-system optimisation or public administration, it becomes strategic infrastructure. When access is conditioned by export controls, domestic fabrication, allied supply chains or national cloud capacity, it becomes a geopolitical resource. Finally, when future capacity is reserved through long-term contracts, priority tiers, capacity options or transferable claims, it can acquire commodity-like and partially financialised attributes. None of these classifications is complete in isolation. Unlike electricity, compute is heterogeneous: one accelerator-hour cannot be substituted perfectly for another because memory, precision, architecture, software, networking, geography and permitted workloads differ. Unlike oil, compute is not stored and later burned; unused capacity is a perishable service flow generated by depreciating equipment. Unlike Bitcoin, it is produced through industrial investment and has no fixed algorithmic supply. Nevertheless, compute can still be rationed, reserved, traded contractually and priced dynamically, making the commodity analogy economically relevant within narrow limits.

Economic identityDefining characteristicTypical purchaserPrincipal price unitMain scarcity mechanismMost appropriate regulatory lens
Normal industrial inputConsumed in producing another good or serviceEnterprise, developer, laboratoryAccelerator-hour, token, task or API callShort-run workload demandInput-cost pass-through and productivity
Scarce capital goodRequires large upfront investment and depreciatesHyperscaler, state, large enterpriseServer, cluster, installed MW or five-year TCOFabrication, financing and installation lead timeInvestment, depreciation and capacity formation
Utility-like serviceContinuously delivered and meteredHousehold, SME, institutionSubscription, token or reserved capacityCongestion and service availabilityTransparency, reliability and nondiscrimination
Strategic infrastructureSupports essential economic or state functionsGovernment, critical infrastructure, hospitalSovereign capacity and service availabilityResilience, security and domestic controlCritical-infrastructure governance
Geopolitical resourceAccess depends on jurisdiction and international alignmentState, defence firm, sanctioned or controlled entityLicensable performance and permitted end useExport controls and territorial concentrationTrade, security and industrial policy
Potentially financialised commodityFuture capacity is reserved, tiered or contractedHyperscaler, model provider, financial investorForward capacity, priority or availability commitmentExpectations of future scarcityMarket design, disclosure and systemic leverage

The classification produces an immediate methodological consequence: there can be no universal “price of AI compute.” The price vector must instead be expressed as P = f(H, M, B, L, E, G, S, Q, R, T), where H denotes hardware architecture, M available memory, B memory and network bandwidth, L latency, E energy burden, G geography, S software ecosystem, Q model or task quality, R reliability and T contract duration. Two services charging the same amount per million tokens may have radically different effective prices if one produces more useful output, requires fewer retries, supports longer context, preserves privacy or supplies guaranteed capacity. Similarly, two computers with identical purchase prices may have different economic costs because one consumes more electricity, has lower utilisation, lacks software compatibility or becomes obsolete sooner. The report will therefore use both nominal prices and quality-adjusted prices, and it will measure access as a bundle rather than a binary condition.


1.3 Compute as an Industrial Input and Capital Good

As an industrial input, AI compute enters a production function alongside labour, data, conventional capital and organisational knowledge. Its marginal value depends not simply on how much compute is consumed, but on whether the enterprise possesses complementary assets capable of converting model output into reliable decisions or revenue. A firm without clean data, process redesign, evaluation systems, cybersecurity, legal governance and trained personnel may purchase large volumes of tokens without achieving measurable productivity gains. Conversely, a specialised small model integrated into a well-designed process may produce a greater return than a more expensive frontier model used without organisational adaptation. A general production specification for enterprise i in period t is Yi,t = Ai,t × F(Ki,t, Li,t, Ci,t, Di,t, Oi,t), where Y is output, K is conventional capital, L is labour, C is quality-adjusted compute, D is usable data, O is organisational capability and A captures other productivity determinants. Compute may complement highly skilled labour, substitute for routine tasks or generate no material return if organisational capability is absent. This complementarity helps explain why a nominally equal subscription does not create equal economic access: large companies can spread integration, governance and evaluation costs across more users and transactions. Eurostat’s 2025 enterprise survey supplies an early observable indicator of this scale effect. AI technologies were used by 55.03% of large EU enterprises, compared with 30.36% of medium-sized enterprises and 17.00% of small enterprises. The survey does not establish that compute prices caused the difference, but it supports testing affordability, expertise and scale as contributing mechanisms. Use of Artificial Intelligence in Enterprises – Eurostat – December 2025

As a capital good, compute is characterised by high initial expenditure, rapid technological obsolescence, uncertain residual value and strong utilisation effects. A cluster that operates near capacity can distribute depreciation and facility costs across far more billable output than an underutilised local installation.  The relevant average cost is ACt = FCt ÷ Qt + VCt, where FC includes depreciation, financing, buildings, networking and fixed labour; Q is quality-adjusted output; and VC includes marginal energy, cooling, maintenance and workload-specific costs. Scale reduces FC ÷ Q until congestion, grid limitations, networking complexity or management costs offset the gain. Capital intensity also changes bargaining power. Buyers able to commit to long-term volumes can secure capacity and preferential commercial terms, while students, researchers, small firms and lower-income governments generally purchase at retail or shared-service prices. Alphabet’s capital expenditure increased from $52.5 billion in 2024 to $91.4 billion in 2025, with the company stating that the expenditure primarily reflected technical infrastructure and that it expected a significant further increase in 2026. This does not identify the precise AI share or prove future price increases, but it establishes the scale of capital mobilisation available to a leading platform operator. Alphabet Annual Report on Form 10-K for the Fiscal Year Ended December 31, 2025 – Alphabet and U.S. Securities and Exchange Commission – February 2026


1.4 Compute as a Utility-Like Service

Cloud AI resembles a utility because it transforms a complex capital system into continuous, remotely delivered and metered access. Users do not need to own a fabrication plant, an accelerator cluster or a data centre; they pay for a service expressed through subscriptions, input tokens, output tokens, images, audio duration, storage, provisioned throughput or completed operations. Utility-like delivery can democratise access because it converts capital expenditure into variable expenditure and permits small buyers to consume fractions of expensive infrastructure. At the same time, it can create dependency if essential functions become inseparable from the provider’s identity, model interface, data formats, security architecture or proprietary tools. The utility analogy is therefore strongest in billing, continuity and dependency, but weaker in technical fungibility. Electricity of a specified voltage is substantially standardised; AI outputs are not. A cheaper model may be an inadequate substitute if it has lower reliability, weaker reasoning, insufficient language coverage, shorter context or no access to required tools. The effective unit is consequently not one token but one successfully completed, quality-controlled task. A useful measurement is ECPm,w = Cm,w ÷ Um,w, where ECP is the effective completion price of model m on workload w, C is total expenditure including retries and supervision, and U is the number of outputs that meet a predefined quality threshold.

The utility analogy also raises questions of tariff design, capacity reservation, universal access, switching rights and service continuity. If providers face congestion, they can respond through waiting times, usage caps, degraded model access, higher priority-tier prices or geographic rationing. Dynamic pricing is not automatically abusive: it can allocate scarce capacity and encourage off-peak demand. It becomes socially consequential when access to education, research, healthcare or competitive business functions depends on the ability to pay peak or frontier-model premiums. Cloud dependence is already broad enough to make this question material. Eurostat reported that 52.74% of EU enterprises purchased cloud-computing services in 2025, while 40.89% of all surveyed enterprises purchased at least one service classified as sophisticated. Among enterprises using paid cloud services, 77.53% were classified as highly dependent on sophisticated cloud functions. These figures concern cloud services generally rather than frontier AI, and they must not be misrepresented as an AI-dependency rate. They nevertheless establish the institutional substrate through which metered AI can diffuse rapidly. Cloud Computing: Statistics on the Use by Enterprises – Eurostat – January 2026

Utility characteristicElectricity or telecommunications analogueAI-compute equivalentCritical difference
MeteringkWh, minute, GBToken, image, accelerator-second, completed taskUnits are not quality-equivalent
Capacity reservationContracted power or bandwidthProvisioned throughput and reserved acceleratorsHardware and model availability vary
Congestion managementPeak tariffs or throttlingRate limits, queues, latency and model tiersQuality may be rationed as well as quantity
Reliability obligationUptime and continuityService-level agreement and model availabilityModel behaviour can change after updates
SwitchingChange supplier using standard connectionPort data, prompts, tools and workflowsProprietary APIs and model behaviour impede portability
Universal accessRegulated baseline servicePossible educational or research compute entitlementNo accepted minimum AI-capability standard exists
Cross-subsidyOne user class subsidises anotherAdvertising, enterprise contracts or investor capital subsidise consumer accessCross-subsidy is often opaque

1.5 Strategic Infrastructure and Compute Sovereignty

Compute becomes strategic infrastructure when its unavailability can impair economic continuity, state capacity, security or essential services. Sovereignty in this context does not require that every component be domestically manufactured. No major economy is fully autonomous across lithography, design software, advanced logic, memory, substrates, networking, servers and energy. Operational compute sovereignty is better defined as the ability of a jurisdiction or institution to obtain, govern, relocate, audit and sustain the minimum quality-adjusted capacity necessary for its critical functions under plausible disruption. This definition contains five dimensions: physical availability, jurisdictional control, operational competence, software and data portability, and supply-chain resilience. A country that owns a data centre but depends entirely on foreign accelerators, proprietary frameworks, remote updates and imported specialist labour possesses infrastructure without complete operational sovereignty. Conversely, a country using foreign-origin hardware may retain meaningful control if it holds installed capacity, diversified spares, local expertise, transferable models and legally enforceable continuity arrangements.

A sovereignty index used later in the report will take the form CSIj = Σwkzj,k, where CSI is the Compute Sovereignty Index for jurisdiction j, z represents normalised indicators and the weights will be tested under equal-weight, principal-component and expert-weight specifications. Indicators will cover installed accelerator capacity, domestic data-centre power, redundancy, cloud-region availability, network connectivity, access to replacement components, public research compute, workforce capability, legal control, model portability and exposure to foreign export restrictions. Europe’s policy response demonstrates that governments no longer treat frontier compute as an ordinary import. The European Commission describes AI Gigafactories as facilities bringing together more than 100,000 advanced AI processors, reliable power, advanced networking and secure supply chains. In July 2026, it launched a call intended to support up to seven such facilities and unlock more than €30 billion in investment. These are policy targets rather than completed capacity and must remain classified as announced mobilisation. AI Factories – European Commission – July 2026 EU Launches AI Gigafactories Call to Boost Europe’s Computing Capacity and Unlock More Than €30 Billion – European Commission – July 2026

Sovereignty dimensionOperational questionProposed indicatorFailure condition
Physical capacityIs sufficient compute installed or contractually guaranteed?Quality-adjusted accelerator capacity per million inhabitants and per unit of GDPCritical workloads cannot obtain capacity
Jurisdictional controlCan foreign law or supplier action interrupt access?Share of critical capacity under domestic or allied legal controlProvider or exporting state can terminate service
Energy resilienceCan the grid support sustained workloads?Firm MW, redundancy and expected outage-adjusted availabilityCapacity exists but cannot be powered reliably
Supply-chain resilienceCan failed systems be repaired or expanded?Supplier diversity, replacement lead time and spare inventorySingle upstream disruption immobilises capacity
Software portabilityCan workloads move between providers or architectures?Migration time, compatible frameworks and open formatsSwitching destroys functionality or requires prohibitive redevelopment
Data governanceCan sensitive data remain controlled and auditable?Local processing, encryption, audit and contractual enforceabilityCompute access requires unacceptable data exposure
Human capabilityCan local institutions operate and optimise the system?Engineers, administrators and AI researchers per installed unitHardware remains underused or foreign-operated
Research accessCan universities and independent researchers obtain meaningful capacity?Public compute hours per researcherFrontier research becomes institutionally exclusive

1.6 Compute as a Geopolitical Resource

Compute is geopolitical because advanced AI capability is territorially concentrated, supply chains cross multiple jurisdictions and access can be altered by export licensing, sanctions, investment controls and alliance relationships. The relevant unit of geopolitical power is not the semiconductor alone but the controlled combination of design, fabrication, packaging, memory, equipment, software and cloud deployment. A state may possess mineral inputs yet lack fabrication technology; it may host fabrication while depending on foreign tools; it may operate data centres while lacking advanced accelerators; or it may develop models that cannot be deployed economically without foreign cloud infrastructure. This interdependence produces bargaining power at chokepoints and vulnerability where substitution is slow. It also creates a distinction between global aggregate capacity and politically accessible capacity. Additional accelerators produced worldwide do not reduce scarcity for every buyer if export controls, allocation agreements or security rules exclude particular jurisdictions.

The semiconductor chain demonstrates this concentration of technical functions. TSMC reported that advanced technologies defined as 7 nanometres and below accounted for 77% of wafer revenue in the second quarter of 2026, while its high-performance-computing platform accounted for 66% of quarterly revenue. These are revenue shares within TSMC, not global market shares, but they show the company’s increasing exposure to leading-edge and HPC demand. TSMC Second Quarter 2026 Earnings Conference Transcript – Taiwan Semiconductor Manufacturing Company – July 2026 NVIDIA’s fiscal-year 2026 filing states that the company depends on foundries to manufacture its semiconductor wafers and presents itself as a full-stack computing-infrastructure company rather than a component supplier. This combination of external manufacturing dependence and downstream platform breadth illustrates why power in the AI chain is distributed but asymmetric: a design firm may control the architecture and software ecosystem without owning the fabrication plant, while the foundry controls a scarce production process without independently controlling final demand. NVIDIA Annual Report on Form 10-K for the Fiscal Year Ended January 25, 2026 – NVIDIA and U.S. Securities and Exchange Commission – February 2026


1.7 Potential Financialisation: What the Bitcoin Analogy Explains—and What It Distorts

AI compute may become more financialised without becoming a cryptocurrency or a conventionally traded commodity. Financialisation occurs when rights to future capacity, priority, revenue or infrastructure are separated from immediate physical use and become objects of contracting, financing or speculation. Long-term accelerator orders, take-or-pay cloud contracts, reserved data-centre capacity, power-purchase agreements, equipment-backed debt and capacity options can transfer scarcity expectations into present asset values. A model provider may secure future compute through a strategic investment from a hyperscaler; the hyperscaler may receive preferred access, distribution rights or revenue participation; and smaller buyers may then face a residual spot market. The Federal Trade Commission’s January 2025 study of major AI partnerships identified agreements involving substantial cloud commitments, exclusivity-related provisions, access to sensitive information and potential switching implications. The report does not establish that every partnership is anticompetitive, but it confirms that investment, cloud purchasing and competitive alignment can be contractually interconnected. FTC Staff Report on AI Partnerships and Investments 6(b) Study – Federal Trade Commission – January 2025

The Bitcoin analogy is valid only at the level of scarcity narratives, divisible consumption units and expectation-driven pricing. It fails in five fundamental respects. First, compute supply is endogenous: companies can manufacture more accelerators and build more data centres, although with delays. Second, compute is heterogeneous: an hour on one architecture is not necessarily substitutable for an hour on another. Third, compute depreciates as newer architectures improve price-performance. Fourth, compute is geographically and electrically constrained. Fifth, it is productive capacity rather than a bearer asset; value arises from the workloads it executes. A more accurate analogy is a hybrid of electricity, cloud bandwidth, industrial machinery and reserved transport capacity. The report will test financialisation through the ratio FCIt = Vreserved,t ÷ Vtotal,t, where Vreserved is the estimated value of capacity committed under forward or preferential arrangements and Vtotal is total available capacity. A high ratio would indicate that quoted retail prices describe only a residual market. Because contractual disclosure is incomplete, the measure will require intervals and must never be presented with false precision.

AttributeBitcoinElectricityIndustrial machineryAI compute
Supply ruleProtocol-limitedCapacity and fuel constrainedInvestment-drivenFabrication, packaging, memory, power and investment constrained
HomogeneityRelatively fungible unitsStandardised within technical limitsHighly heterogeneousHighly heterogeneous
StorageDigital bearer assetLimited and costlyAsset can remain idleCapacity flow is perishable; hardware is durable but depreciating
DepreciationNo physical depreciationNot applicable to consumed energyPhysical and technological depreciationRapid technological and economic depreciation
Geographic dependenceNetwork accessGrid-specificInstallation-specificData-centre, grid, jurisdiction and network-specific
Productive useIndirectUniversal energy inputDirect production assetDirect cognitive and computational input
Dynamic pricing potentialExchange priceSpot and peak tariffsRental or leasing ratesToken, latency, model, workload and capacity-tier pricing
Appropriate analogy scoreLowModerateHighHybrid category

1.8 The AI Compute Value Chain

The AI value chain is best analysed as a set of complementary layers rather than a simple linear sequence. A shortage at any indispensable layer can reduce the output of the entire system, while control of a technically replaceable layer may generate little bargaining power if switching is inexpensive. The relevant measure is therefore bottleneck-adjusted capacity. If effective capacity is constrained by the minimum available complementary input, then Ceffective = min(Clogic, Cmemory, Cpackaging, Cnetwork, Cpower, Csoftware). This Leontief-style representation is deliberately stringent: in practice, some substitution is possible through lower precision, slower networks, different models or reduced service quality. The model will later be expanded into a constant-elasticity-of-substitution structure. For Chapter 1, it captures the central point that surplus accelerator-design demand is economically irrelevant if HBM or packaging prevents finished systems from being delivered.

Table 1.1 — AI Compute Value Chain, Scarcity Nodes and Bargaining Power

LayerCore economic functionCapital intensityShort-run substitutabilityPotential scarcity indicatorSource of bargaining powerPrimary later chapter
Semiconductor equipmentEnables leading-edge fabricationExtremeVery lowTool backlog and delivery lead timeProprietary process technology and installed baseChapter 3
Design software and IPConverts architecture into manufacturable designHigh knowledge intensityLowLicence concentration and switching timeStandards, libraries and engineering integrationChapter 3
Accelerator architectureExecutes AI workloadsHigh R&DMedium over longer periodsQuality-adjusted shipments and order backlogPerformance, software ecosystem and developer baseChapters 2–3
Leading-edge foundryManufactures logic diesExtremeVery low in the short runAdvanced-node utilisation and capacityYield, process leadership and scaleChapter 3
HBM and advanced memorySupplies model weights and data at high bandwidthExtremeLowContracted supply, bit output and priceManufacturing know-how and qualificationChapters 2–3
Advanced packagingIntegrates logic and memoryExtremeVery lowCoWoS-equivalent capacity and lead timeProcess qualification and limited capacityChapter 3
Substrates and componentsConnect and support packaged systemsHighLow to mediumLead times and supplier concentrationQualification and production constraintsChapter 3
Server integrationConverts components into deployable systemsHighMediumRack delivery and liquid-cooling availabilityIntegration, warranty and customer qualificationChapters 2–4
NetworkingConnects accelerators at cluster scaleHighLow for frontier trainingBandwidth, topology and port availabilityScale efficiency and architecture compatibilityChapters 3–5
Data-centre facilitiesHouses and cools systemsExtremeLow locallyMW under construction and connection queueLand, permits, cooling and grid accessChapter 5
ElectricityOperates the infrastructureExtreme system capitalLow during local grid constraintFirm MW, prices and connection delayLocation-specific availabilityChapters 5 and 9
Cloud orchestrationAllocates and meters resourcesExtremeMedium, but migration is costlyUtilisation, reserved capacity and availabilityScale, customer base and integrated servicesChapters 3 and 5
Foundation modelsConverts infrastructure into capabilityExtreme training and R&DTask-dependentBenchmark-adjusted access and priceModel quality, brand, data and distributionChapters 5 and 7
Application layerEmbeds models in workflowsVariableMedium to highCustomer retention and integration depthData, workflow and distribution lock-inChapters 6–7
User organisationConverts output into economic valueVariableNot substitutableSkills, governance and complementary capitalProprietary data and domain competenceChapters 6–8

Three conclusions follow from this map. First, market concentration must be measured separately at every layer; a single “AI market share” is analytically meaningless. Second, high concentration does not necessarily identify the binding bottleneck. A highly concentrated layer with ample capacity and credible substitution may exercise less power than a somewhat less concentrated layer with long qualification periods and no inventory. Third, bargaining power can move over time. When accelerators are scarce, the designer or cloud allocator may dominate; when accelerators become abundant but electricity connections are delayed, power availability and data-centre permits become the constraint. The U.S. Department of Energy reported that American data centres consumed approximately 176 TWh in 2023, or 4.4% of total U.S. electricity, and projected a range of 325–580 TWh by 2028, equivalent to approximately 6.7–12% of national consumption. These are scenarios rather than fixed outcomes, but they demonstrate why power must be treated as part of the compute chain rather than an external operating expense. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – U.S. Department of Energy – December 2024


1.9 Where Value, Scarcity and Bargaining Power Accumulate

Economic value is captured where an actor controls a scarce complement, creates a difficult-to-replicate performance advantage, owns distribution or raises the cost of switching. These conditions do not always coincide. A foundry may possess strong technical scarcity but remain exposed to a few large customers; a hyperscaler may have abundant infrastructure but possess bargaining power through customer relationships and bundled services; a model provider may have superior capability but depend on a cloud partner for capital and distribution.  The balance can be represented through a bargaining-power score BPi = w₁Si + w₂Li + w₃Di + w₄Ei + w₅Vi − w₆Bi, where S is scarcity, L lock-in, D differentiation, E ecosystem control, V vertical reach and B buyer concentration. This is a research index, not an established law. Each component will be normalised and tested with alternative weights.

Table 1.2 — Preliminary Bargaining-Power Assessment

NodePhysical scarcityEcosystem lock-inBuyer concentrationReplication timePreliminary power assessmentEvidentiary status
Leading-edge fabricationHighMediumHighVery longHigh but constrained by major buyers and capital cyclesFramework inference
Advanced packagingHighMediumHighLongHigh during AI-capacity expansionSupported bottleneck hypothesis
HBMHighMediumHighLongHigh if contracted output leaves a small residual marketTo be quantified
Accelerator designHighVery high for established software ecosystemsHighLongVery high where performance and software reinforce one anotherTo be quantified
Data-centre power connectionLocally highLowVariableLongPotentially decisive in constrained regionsSupported by DOE scenarios
Hyperscale cloudHigh capacity barrierVery highLow buyer power for retail usersVery longHigh through bundling, scale and allocationTo be quantified
Frontier modelHigh R&D barrierHighMixedUncertainHigh if task quality is not reproducible by alternativesBenchmark-dependent
Open-weight modelLower legal access barrierLowerLowShorterLimited direct pricing power, but infrastructure remains necessaryWorkload-dependent
Enterprise applicationUsually lower physical barrierPotentially high workflow lock-inFragmentedMediumCan capture value through integration and proprietary dataSector-dependent
End userLow individuallyOften locked inVery low individuallyNot applicablePrice taker unless organised purchasing existsStrong prior expectation

NVIDIA’s fiscal 2026 results illustrate value capture without resolving its cause. Fiscal-year revenue was $215.9 billion, 65% above the preceding year, and fiscal-year GAAP gross margin was 71.1%. In the first quarter of fiscal 2027, Data Center revenue reached $75.2 billion, 92% above the corresponding prior-year quarter. These data demonstrate exceptional commercial performance, not unlawful conduct. Later chapters will compare margins with historical baselines, invested capital, R&D, supply commitments, competitive alternatives and changes in product mix before estimating economic rent. NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026 – NVIDIA – February 2026 NVIDIA Announces Financial Results for First Quarter Fiscal 2027 – NVIDIA – May 2026


1.10 Core Concepts and Operational Definitions

Table 1.3 — Conceptual Dictionary

ConceptOperational definition for this reportProposed measurementWhat must not be inferred automatically
Compute scarcityQuality-adjusted demand exceeds capacity available at the prevailing price and service conditionsUtilisation, queues, lead times, stock-outs, premiums and unmet demandA high price alone does not prove scarcity
Physical scarcityInsufficient deployable hardware, power or facilitiesUnits, MW, delivery times and availabilityIt does not necessarily imply market abuse
Economic scarcityCapacity exists but is unaffordable to a defined user groupCost-to-income or cost-to-revenue ratioIt is not equivalent to physical shortage
Artificial scarcitySupply is strategically withheld below the profit-maximising competitive benchmarkCapacity, output, inventories, contracts and counterfactual supplyIt cannot be inferred from concentration alone
Compute sovereigntyAbility to sustain governed access to essential compute under disruptionComposite resilience and control indexDomestic ownership alone is insufficient
Access inequalityUnequal distribution of quality-adjusted AI capabilityGini, Theil, percentile ratios and affordability indicesSubscription ownership is not capability equality
Digital exclusionInability to obtain meaningful AI use because of connectivity, devices, skills, language or affordabilityMultidimensional deprivation rateBeing online does not imply meaningful AI access
Economic rentReturn above the minimum required to keep capital and capability in their current useExcess ROIC, residual income and margin decompositionAccounting profit is not identical to rent
Network effectsProduct value increases as users, developers, tools or complementary data increaseDeveloper base, integrations and retentionScale alone is not proof of a network effect
Economies of scaleAverage cost falls as output or utilisation risesCost elasticity relative to outputLarge size is not proof of continuing scale economies
Economies of scopeJoint production costs less than separate productionMulti-product cost functionBundling is not always efficient or harmful
Vertical integrationOne firm controls multiple adjacent layersRevenue, ownership and contractual control mapIntegration can create efficiency or foreclosure
Switching costEconomic loss caused by moving provider or architectureMigration expenditure, time, performance loss and retrainingContract termination fees capture only one component
Data lock-inUser data, embeddings, histories or workflows cannot move without lossPortability completeness and migration costData possession alone does not create lock-in
Ecosystem lock-inComplementary software and skills make alternatives costlyCompatibility, developer tools and retrainingPopularity alone is insufficient
Technological dependenceCritical capability requires a supplier or jurisdiction without timely substituteReplacement time and criticality-adjusted exposureImport dependence is not always vulnerability

Compute scarcity will be decomposed into five observable forms: quantity scarcity, quality scarcity, temporal scarcity, geographic scarcity and institutional scarcity. Quantity scarcity exists when total units are inadequate. Quality scarcity exists when lower-performance substitutes are available but cannot complete the required workload. Temporal scarcity appears as waiting time, delivery delay or queueing. Geographic scarcity arises when capacity is available globally but not within the required jurisdiction or latency boundary. Institutional scarcity exists when capacity is accessible to hyperscalers or elite universities but not to small companies, independent researchers or public institutions. This decomposition prevents a common analytical error: concluding that scarcity has disappeared because some form of compute remains purchasable. A user able to run a small, highly quantised local model may possess quantity access but remain quality-excluded from workloads requiring long context, multimodality, high reliability or rapid concurrent inference.


1.11 Network Effects, Scale, Scope and Vertical Integration

AI markets can exhibit several mutually reinforcing scale mechanisms. Infrastructure scale lowers average fixed cost by raising utilisation and spreading data-centre, networking and engineering expenditure across more workloads. Model-development scale allows training, safety, evaluation and post-training costs to be amortised across more users. Data scale can improve products when legally and technically usable interactions generate feedback. Distribution scale lowers customer-acquisition costs. Developer-network effects increase the value of an architecture as more libraries, optimisation tools and trained personnel become available. These mechanisms can generate genuine efficiency while simultaneously raising entry barriers. The analytical task is not to label scale as harmful, but to determine whether efficiency gains are passed to users through lower quality-adjusted prices or retained as durable rent.

Economies of scale will be tested through ln(ACi,t) = α + βln(Qi,t) + γXi,t + εi,t. A statistically negative β would be consistent with falling average cost as output expands, subject to data quality and identification limitations. Economies of scope will be examined through SC = C(Y₁,0) + C(0,Y₂) − C(Y₁,Y₂). A positive value indicates a cost advantage from joint production. Vertical integration will be evaluated through both efficiency and foreclosure channels. Integration between accelerators, networking, cloud, models and applications can reduce coordination costs and optimise the stack. It can also make competing models or hardware less attractive through preferential access, bundling, interoperability restrictions or discriminatory commercial terms. The FTC’s AI-partnership report provides a relevant empirical starting point because it examined how investments and partnerships could affect access to cloud resources, switching, exclusivity and sensitive information. It does not support a blanket conclusion that integration is anticompetitive; it establishes the contractual mechanisms that must be tested.. FTC Staff Report on AI Partnerships and Investments 6(b) Study – Federal Trade Commission – January 2025

Table 1.4 — Efficiency and Exclusion Tests

MechanismEfficiency hypothesisExclusion hypothesisRequired evidence
Hardware–software integrationImproves performance and lowers workload costLocks developers into one architectureCross-platform benchmark, migration cost and price comparison
Cloud–model integrationReduces deployment cost and latencyPreferentially allocates capacity to affiliated modelCapacity terms, pricing and availability by model
Model–application integrationImproves user experience and workflow automationMakes rival models technically or commercially inaccessibleInteroperability, default settings and contractual terms
Data–model integrationImproves relevance and personalisationEntrenches incumbent through inaccessible feedback dataData portability, learning effects and incremental performance
Bundled enterprise servicesReduces procurement and security costUses strength in one market to foreclose anotherStand-alone versus bundled prices and customer switching
Long-term capacity contractFinances new infrastructure and stabilises demandRemoves capacity from competitors or residual buyersContract duration, reserved share and alternative capacity
Developer ecosystemReduces implementation costRaises retraining and redevelopment barriersLabour-market skills, library compatibility and porting time

1.12 Market Concentration, Dominance and Economic Rent

Concentration is an indicator, not a verdict. The report will calculate concentration at narrowly defined functional layers using CR₄ = s₁ + s₂ + s₃ + s₄ and HHI = Σsi², with shares expressed as percentages. Market definition will precede calculation: data-centre GPUs, consumer GPUs, HBM, leading-edge foundry services, general cloud infrastructure and frontier-model subscriptions cannot be placed in a single denominator.  The U.S. 2023 Merger Guidelines state that a merger producing or increasing a highly concentrated market may raise a presumption of illegality under specified conditions, but the guidelines are an enforcement framework, not a declaration that every concentrated market is unlawful. 2023 Merger Guidelines – U.S. Department of Justice and Federal Trade Commission – December 2023 The report will therefore supplement HHI with entry conditions, capacity constraints, price-cost margins, buyer power, innovation, multi-homing, switching and credible substitution.

Economic rent will be estimated through multiple methods because no single accounting ratio isolates market power. The first measure will be excess return: ERi,t = ROICi,t − WACCi,t. The second will be residual income: RIi,t = NOPATi,t − WACCi,t × ICi,t. The third will decompose gross-margin changes into selling price, product mix, unit cost, utilisation and accounting effects. A positive excess return may represent innovation rent, temporary scarcity rent, intangible capital omitted from the balance sheet, superior management or durable market power. The classification will depend on persistence, entry, price response to capacity expansion and whether returns remain above benchmarks after risk and intangible investment are accounted for.

Table 1.5 — Rent Classification Framework

Rent categoryEconomic originExpected durationCompetition implicationEmpirical signature
Innovation rentSuperior technology or productTemporary unless continuously renewedOften pro-competitiveHigh R&D, rapid performance improvement and eventual entry
Scarcity rentDemand temporarily exceeds capacityCyclical or capacity-dependentNot inherently anticompetitivePrices and margins fall after capacity expansion
Quasi-rentRecovery on specialised sunk investmentAsset-life dependentUsually normalReturn falls as capital depreciates or contracts expire
Risk premiumCompensation for uncertain investmentLinked to riskNormal if proportionateReturn correlates with investment risk
Ecosystem rentComplementary tools and users reinforce incumbentPotentially persistentAmbiguousHigh switching costs and developer dependence
Monopoly or dominance rentDurable market power restricts competitive pressurePersistentPotential concernHigh margins, weak entry, limited substitution and strategic exclusion
Regulatory rentRules or controls restrict entryPolicy-dependentMay be intended or distortiveProfit tied to licences, quotas or protected access
Geopolitical rentJurisdictional access or controls create privileged supplyEvent-dependentSecurity and trade issueRegional price and availability divergence

1.13 Access Inequality and Digital Exclusion

AI access inequality must be measured as a distribution of capabilities rather than a distribution of accounts. A free model, a paid frontier subscription, a dedicated enterprise instance and a locally controlled high-memory system are not equivalent resources. Quality-adjusted access will therefore incorporate model performance, usage limits, context capacity, latency, privacy, availability, language coverage and tool access. The proposed individual measure is QAAi,t = ΣwkAi,k,t, where A represents normalised access dimensions. Affordability will be measured as AAIi,t = Cadequate,i,t ÷ Ydisposable,i,t. For companies, the denominator becomes revenue, operating expenditure or labour cost. For universities, it becomes research expenditure or funding per researcher. A system will be deemed affordable only relative to a defined workload; a low-cost model that cannot complete the task is not an affordable substitute.

The baseline digital divide remains severe. The ITU estimated that 6.0 billion people were online in 2025, while 2.2 billion remained offline, mainly in low- and middle-income economies. It reported internet use of 94% in high-income countries and 23% in low-income countries. Measuring Digital Development: Facts and Figures 2025 – International Telecommunication Union – November 2025 These figures do not measure AI access, but they define its lower boundary: users without meaningful connectivity cannot participate fully in cloud AI, while users with weak connectivity may be unable to sustain multimodal or interactive workloads. The ITU’s 2025 connectivity analysis frames meaningful connectivity through quality, availability, affordability, devices, skills and security, a multidimensional structure that this report will extend to AI compute. Global Connectivity Report 2025 – International Telecommunication Union – 2025

Table 1.6 — AI Access Ladder

LevelAccessible capabilityTypical constraintEconomic consequence
0 — ExcludedNo reliable internet or compatible deviceConnectivity, electricity, incomeNo regular AI participation
1 — Nominal accessFree, severely limited or intermittent serviceCaps, queues and older modelsBasic assistance but weak continuity
2 — Standard consumer accessPaid general-purpose modelSubscription affordability and privacyIndividual productivity improvement
3 — Advanced professional accessHigher limits, tools, larger context and multimodalityHigher recurring costMaterial professional advantage
4 — Enterprise controlled accessSecurity, integration, governance and supportOrganisational scale and implementation costWorkflow-level productivity and data integration
5 — Local high-performance accessPrivate inference with substantial memory and throughputHardware, energy and technical expertiseSovereignty and predictable availability
6 — Frontier research accessLarge clusters, training and large-scale experimentationExtreme capital and institutional concentrationAbility to create rather than merely consume frontier systems

1.14 Falsifiable Research Hypotheses

H1 — AI-Relevant Hardware Prices Have Risen Faster Than General Consumer-Electronics Prices

H1 will be supported only if a representative, transaction-weighted AI-hardware price index rises faster than the appropriate consumer-electronics comparator over the defined period. The test cannot rely on selected flagship products or isolated resale listings. Chapter 2 will construct matched and hedonic indices for GPUs, memory, workstations, laptops and relevant Apple and Samsung configurations. The principal regression will be ln(Pi,t) = α + βXi,t + γt + δb + εi,t, where X captures memory, bandwidth, compute performance, energy, condition and other measurable characteristics; γ is the time effect and δ the brand or product-family effect. H1 will be rejected if nominal prices rise but quality-adjusted prices decline at a rate consistent with or faster than the comparator index.

H2 — The Inflation-Adjusted Cost of a Fixed Level of AI Capability Has Increased

H2 is stricter than H1. The unit of analysis will be a reproducible workload, not hardware. Cost per successful task will include capital cost, electricity, software, time, retries and human correction. The fixed-capability basket will cover reasoning, coding, long-document analysis, multilingual work and multimodal processing. H2 will be supported when the real minimum cost of reaching a constant benchmark threshold rises significantly across multiple workloads and deployment modes. It will be rejected if improved algorithms, quantisation or hardware efficiency reduce real workload cost despite higher purchase prices.

H3 — Used-Market Prices Indicate Persistent Scarcity Rather Than Temporary Speculation

H3 will compare used prices with launch MSRP, contemporaneous new prices, depreciation benchmarks and product availability. The scarcity premium will be USPi,t = (Pused,i,t − Pfundamental,i,t) ÷ Pfundamental,i,t. Persistence will require premiums across multiple months, regions and product classes, accompanied by low availability or delivery delays. H3 will be rejected if premiums are confined to brief launches, collectible products, unreliable listings, tax distortions or speculative episodes without transaction evidence.

H4 — Concentration Enables Suppliers to Retain Above-Normal Economic Rents

H4 will require a relationship between concentration and persistent excess returns after controlling for R&D, risk, intangible capital, scale and product mix. The panel specification will be ERi,t = α + βHHIm,t + γZi,t + μi + τt + εi,t. A positive β alone will not prove causality. Event studies around capacity expansion, entry or supply shocks will test whether margins respond competitively. H4 will be rejected if high returns dissipate with entry or are explained by innovation and risk-adjusted investment.

H5 — Frontier AI Access Produces Measurable Economic and Educational Advantages

H5 will be tested through matched or randomised comparisons where credible data exist. Outcomes will include task accuracy, completion time, grades, research output, coding productivity, revenue and error rates. Treatment must represent quality-adjusted access rather than any AI use. The preferred estimand is ATE = E[Y(1) − Y(0)], with controls for prior ability, institution, income and task. H5 will be rejected for workloads where smaller or cheaper systems achieve statistically equivalent results.

H6 — Lower-Income Users and Countries Are Increasingly Restricted to Lower-Capability Systems

H6 will compare quality-adjusted access, not internet penetration alone, across income groups. The tests will use affordability, available model tiers, cloud regions, hardware-wage ratios, electricity reliability, language support and public research compute. Increasing restriction requires a widening gap through time. H6 will be rejected if low-cost access, open models and shared infrastructure cause capability convergence after purchasing-power adjustment.

H7 — Current AI-Service Pricing Does Not Fully Reflect Long-Run Infrastructure Cost and May Rise Materially

H7 will estimate provider break-even prices using depreciation, financing, energy, facilities, networking, labour, training and utilisation. The cost-recovery ratio will be CRRp,t = RevenueAI,p,t ÷ FullyAllocatedCostAI,p,t. Because public segment data are incomplete, estimates will be ranges. H7 will be supported if recurring prices remain below long-run cost under plausible utilisation and cross-subsidy assumptions. It will be rejected if current pricing covers cost and required capital return or if unit costs decline faster than tariff pressure.

H8 — AI Access Is Evolving into a Dynamically Priced, Metered Resource

H8 concerns market design rather than simply rising prices. It will be supported by documented movement toward multiple metering dimensions: model class, reasoning intensity, latency, context length, time of demand, geography, capacity reservation and service priority. A Dynamic Pricing Complexity Index will count and weight independently priced dimensions. H8 will be rejected if markets converge toward simple, stable, interoperable and progressively cheaper tariffs.


1.15 Hypothesis Test Matrix

Table 1.7 — Falsification Criteria

HypothesisPrimary dependent variableCore comparatorSupport thresholdRejection conditionPrincipal risk
H1AI-hardware price indexConsumer-electronics indexStatistically higher cumulative growthQuality-adjusted growth is equal or lowerProduct-selection bias
H2Real cost per fixed-capability taskConstant benchmark thresholdSignificant real increase across robust specificationsReal task cost declinesBenchmark drift
H3Used-market premium and durationDepreciation-adjusted fair valuePersistent multi-market premium with confirmed transactionsShort-lived or listing-only premiumSelection and fake listings
H4Risk-adjusted excess returnCompetitive sector benchmarkPersistent excess return associated with concentration and weak entryReturns explained by innovation, risk or cycleMarket-definition error
H5Productivity or educational outcomeComparable non-frontier accessPositive statistically and economically significant treatment effectEquivalent outcomes from lower-cost systemsUser self-selection
H6Quality-adjusted access gapIncome and regional groupsGap widens through 2026–2031Convergence after PPP and quality adjustmentMissing lower-income data
H7Cost-recovery ratioFully allocated long-run costRatio below 1 under central assumptionsCurrent prices cover cost and capital returnOpaque segment accounts
H8Pricing-complexity indexHistorical tariff structureMore independently priced dimensions and capacity tiersSimplification and commoditisationContract confidentiality

1.16 Research Map: Hypotheses, Data and Later Chapters

Table 1.8 — Full Research Architecture

HypothesisDataset requiredMinimum frequencyMethodRobustness testDeveloped in
H1Launch MSRP, official current prices, verified transaction prices, specifications and electronics indicesMonthly or quarterly, 2020–2031Hedonic price regression and matched-model indexAlternative quality measures and geography fixed effectsChapter 2
H2Hardware TCO, API prices, benchmarks, electricity and human-review costQuarterly or annualCost-per-successful-task frontierAlternative benchmarks, quantisation and utilisationChapters 2, 4 and 5
H3Verified second-hand transactions, age, condition, warranty and new availabilityWeekly or monthlyPremium-duration and event-study modelsListings versus completed transactionsChapter 2
H4Segment shares, CR₄, HHI, margins, ROIC, WACC, entry and capacityQuarterly or annualPanel regression, residual income and event studyIntangible-capital adjustment and market definitionsChapter 3
H5User-level access, task outcomes, time, errors and background variablesExperimental or panelRCT, matching or difference-in-differencesPlacebo tasks and heterogeneous effectsChapter 7
H6Income, hardware affordability, connectivity, cloud access, language and public computeAnnualComposite index, Gini, Theil and decompositionPPP versus market FX and alternative weightsChapter 8
H7Provider capex, depreciation, energy, utilisation, revenue and tariffsQuarterlyFully allocated cost and break-even modelMonte Carlo utilisation and asset-life assumptionsChapter 5
H8Public API tariffs, enterprise contracts, priority products and reserved capacityMonthly or event-basedTariff taxonomy and complexity indexProvider and regional comparisonChapters 5 and 9

Table 1.9 — Cross-Chapter Dependency Map

Later chapterContribution to Chapter 1 hypothesesPrincipal output
Chapter 2 — Hardware InflationTests H1 and H3; supplies hardware inputs for H2Nominal, real and quality-adjusted price indices
Chapter 3 — Industrial ConcentrationTests H4 and maps upstream dependenceCR₄, HHI, entry barriers and rent decomposition
Chapter 4 — Real Cost of Local AITests the local component of H2 and H6Workload-specific five-year TCO
Chapter 5 — Cloud AI EconomicsTests H7 and H8 and cloud component of H2Cost per successful task and break-even tariffs
Chapter 6 — Enterprise ExposureTranslates H2, H7 and H8 into firm-level effectsAI cost burden, ROI and sector vulnerability
Chapter 7 — Capability DivideTests H5 and individual component of H6Educational and productivity treatment effects
Chapter 8 — Global InequalityTests country-level H6Global AI Access and Compute Sovereignty indices
Chapter 9 — ForecastsProjects all supported relationships to 2031Scenario ranges, prediction intervals and warning indicators
Chapter 10 — PolicyResponds only to empirically supported failuresCosted interventions and measurable KPIs

1.17 Preliminary Evidence Register

Table 1.10 — What Chapter 1 Establishes and What Remains Unproven

PropositionCurrent statusEvidence presently availableRequired next test
AI compute depends on multiple complementary physical and software layersEstablished conceptual factEuropean Commission supply-chain and bottleneck analysisQuantify each layer
AI infrastructure investment is exceptionally largeEstablishedAlphabet 2025 capex and European Gigafactory mobilisationCompare firms, regions and historical baseline
Advanced fabrication and HPC demand are strategically importantEstablishedTSMC revenue compositionGlobal shares and alternative capacity
Energy can become a binding compute constraintStrongly supportedDOE data-centre consumption scenariosRegional power-price and grid-connection model
AI adoption differs substantially by enterprise sizeEstablished correlationEurostat 2025 surveyIdentify causal contributions of cost, skills and scale
Global connectivity remains highly unequalEstablishedITU 2025 indicatorsExtend from connectivity to quality-adjusted AI access
AI hardware has become universally more expensive in real termsUnprovenNo complete quality-adjusted price index yetChapter 2
Used AI hardware systematically trades at two or three times launch priceUnprovenRequires transaction-level evidenceChapter 2
Supplier concentration causes unlawful pricingUnprovenConcentration and high margins are insufficientChapter 3
Current token prices are unsustainableUnprovenProvider-specific cost allocation is incompleteChapter 5
Frontier access improves all educational and professional outcomesUnproven and likely task-dependentRequires causal outcome evidenceChapter 7
Lower-income populations are trapped permanently in inferior AIUnprovenConnectivity and enterprise-adoption gaps support investigationChapter 8
Compute will be traded like BitcoinEconomically overstatedSome capacity can be reserved and dynamically pricedChapters 5 and 9

1.18 Chapter Conclusion

AI compute should be understood as a hybrid economic institution: an industrial input when consumed, a capital good when owned, a utility-like service when delivered through the cloud, strategic infrastructure when essential functions depend upon it, a geopolitical resource when access is jurisdictionally controlled and a potentially financialised capacity when future availability is reserved or monetised. This hybrid status explains why conventional consumer-electronics analysis is insufficient. A laptop price cannot reveal the full cost of local AI; a token tariff cannot reveal long-run infrastructure economics; a global shipment figure cannot reveal politically accessible supply; and an internet-penetration rate cannot reveal access to frontier capability.

The initial evidence establishes four conditions that justify the report. First, capital mobilisation is occurring at a scale that can advantage a small number of firms and states able to finance infrastructure before demand is fully realised. Second, the chain contains complementary nodes—advanced logic, HBM, packaging, networking and power—where substitution can be slow. Third, enterprise adoption already varies sharply by company size, although causality has not yet been assigned. Fourth, global connectivity and income inequalities form a pre-existing structure through which AI inequality may be transmitted. None of these conditions proves artificial scarcity, monopolistic abuse or inevitable price escalation. The central scientific obligation of the remaining chapters is to determine whether the observed system reflects a temporary investment cycle followed by falling quality-adjusted costs, or a persistent architecture in which concentrated control and unequal purchasing power transform intelligence into a rationed economic resource.

The report’s eight hypotheses are deliberately falsifiable. H1 and H2 separate sticker-price inflation from the cost of fixed capability. H3 distinguishes durable scarcity from second-hand speculation. H4 separates concentration from demonstrable economic rent. H5 measures advantage rather than assuming it. H6 converts the idea of an AI divide into a comparative access distribution. H7 tests whether current tariffs recover long-run costs. H8 identifies whether metering is becoming multidimensional and dynamic. Together they create a research programme capable of confirming, rejecting or qualifying the report’s central warning without converting a legitimate concern into a predetermined conclusion.

You were right about the formatting error in Chapter 1. All mathematical subscripts below use HTML <sub> tags. Hypothesis labels such as H1, H2 and H3 remain ordinary text.

Chapter 2 — Hardware Inflation: GPUs, Memory, Computers and the Used Market

2.1 Scope, Evidentiary Standard and Central Finding

This chapter tests whether AI-relevant hardware became structurally less affordable between 2020 and 2026 and whether the observed increase represents nominal inflation, quality-adjusted inflation, temporary scarcity, regional price distortion or speculative resale behaviour. The first conclusion is methodologically decisive: no single “AI hardware price” exists. A consumer GPU, professional accelerator, data-centre server, high-memory Apple workstation and AI-capable laptop provide different combinations of memory capacity, bandwidth, compute throughput, power efficiency, software compatibility, portability, warranty and useful life. A valid price comparison must therefore preserve either the product configuration or the workload capability. Comparing the launch price of a 2020 GPU with the price of a technically different 2025 GPU without controlling for capability can exaggerate or conceal inflation. The chapter consequently separates four objects: the nominal purchase price; the general-inflation-adjusted price; the specification-adjusted price; and the cost of completing a constant AI workload.

The verified flagship NVIDIA sequence illustrates the difference. The GeForce RTX 3090 launched in September 2020 at $1,499, the RTX 4090 in October 2022 at $1,599, and the RTX 5090 in January 2025 at $1,999. The nominal flagship entry price therefore rose by 33.4% between the RTX 3090 and RTX 5090. However, memory capacity increased from 24 GB to 32 GB, while official memory bandwidth rose from approximately 936 GB/s to 1,792 GB/s. Once the launch prices are converted into July 2026 dollars using the U.S. Consumer Price Index, the estimated real launch price rises by approximately 9%, while the real price per GB of accelerator memory declines by approximately 18% and the real price per GB/s of bandwidth declines by more than 40%. These calculations do not prove that AI became cheaper in practical terms because memory capacity and bandwidth alone do not measure model quality, software overhead, context length, energy cost or tokens per second. They do demonstrate that the claim “high-end GPUs became two or three times more expensive” is not representative of the verified NVIDIA Founders Edition flagship MSRP sequence. Introducing GeForce RTX 30 Series Graphics Cards – NVIDIA – September 2020 NVIDIA Delivers Quantum Leap in Performance, Introduces New Era of Neural Rendering with GeForce RTX 40 Series – NVIDIA – September 2022 NVIDIA Blackwell GeForce RTX 50 Series Opens New World of AI Computer Graphics – NVIDIA – January 2025

That result applies only to launch MSRP and selected technical denominators. It does not describe actual street prices, availability during launch windows, board-partner premiums, taxes, import costs or second-hand transactions. Under the evidence protocol governing this report, an asking price cannot be represented as a transaction, a single listing cannot establish a median, and a marketplace advertisement cannot prove that a product was sold. No official manufacturer, government statistical agency or audited corporate filing publishes a complete 2020–2026 international series of transaction-level used GPU prices. Consequently, the used-market medians requested for the United States, EU, United Kingdom, China, Japan, India, Brazil and lower-income markets are marked ND — not disclosed or not verifiable under the permitted source hierarchy. This is not a missing detail that can be responsibly filled with estimates: it is a material data limitation. The claim that ordinary used AI hardware systematically sells at two or three times original MSRP remains unproven until completed-transaction microdata are obtained, cleaned and audited.


2.2 Price Concepts and Measurement Architecture

Hardware inflation must be decomposed into at least seven components:

  1. general monetary inflation;
  2. change in technical quality;
  3. change in product position within the manufacturer’s range;
  4. changes in taxes, tariffs and exchange rates;
  5. temporary scarcity;
  6. distributor or reseller margin;
  7. speculative or exceptional resale premiums.

The observed local price can be represented as:

Pobserved,i,r,t = Pbase,i,t × FXr,t × (1 + Taxr,t) × (1 + Tariffr,t) × (1 + Distributioni,r,t) × (1 + Scarcityi,r,t) × (1 + Speculationi,r,t)

Where:

  • Pobserved,i,r,t is the observed price of product i in region r and period t;
  • Pbase,i,t is the manufacturer reference price;
  • FXr,t is the applicable exchange-rate conversion factor;
  • Taxr,t includes VAT, sales tax and other mandatory charges;
  • Tariffr,t includes customs duties and import charges;
  • Distributioni,r,t is the wholesale and retail margin;
  • Scarcityi,r,t measures the premium associated with supply-demand imbalance;
  • Speculationi,r,t is the residual premium not explained by ordinary distribution and scarcity.

This multiplicative formulation prevents taxes and currency changes from being incorrectly classified as manufacturer inflation. It also avoids treating the U.S. pre-sales-tax MSRP as directly comparable with an EU consumer price that normally includes VAT. A €2,000 EU price including 22% VAT is not economically equivalent to a $2,000 U.S. price excluding sales tax. Geographic comparisons must remove recoverable taxes for enterprise purchasers, retain taxes for household affordability analysis and use both market exchange rates and purchasing-power-parity conversions where appropriate.

Table 2.1 — Mandatory Price Fields and Evidentiary Status

FieldDefinitionAcceptable evidenceCurrent status
Launch dateFirst official commercial availabilityManufacturer announcementVerifiable
Launch MSRP or SEPManufacturer’s recommended launch priceOfficial manufacturer announcementVerifiable for many consumer products
Current official priceManufacturer’s currently displayed direct priceLive manufacturer store or official product pageVerifiable only while published
Street priceActual authorised-retailer transaction priceAudited retailer transactions or official statistical scanner dataGenerally unavailable
Used asking priceSeller’s requested amountMarketplace listingObservable but not a completed transaction
Used transaction priceAmount paid in a completed saleCompleted-sale microdataNot available within current source constraints
Used-market medianMedian of cleaned completed transactionsRepresentative transaction datasetND
Inflation-adjusted pricePrice converted into constant-period currencyOfficial national CPI or HICPCalculable
Quality-adjusted pricePrice divided by a stable capability measureVerified price and specification dataPartly calculable
Workload-adjusted priceCost per successful constant workloadReproducible benchmark and complete TCORequires Chapter 4 testing
Warranty statusRemaining transferable manufacturer coverageSerial or invoice-level evidenceCannot be inferred for aggregate used listings
Regional premiumLocal net-of-tax price relative to common baselineOfficial regional prices, FX and tax dataRequires configuration matching

2.3 Verified Consumer GPU Price History, 2020–2026

2.3.1 NVIDIA Flagship Sequence

The following table uses only manufacturer-published launch prices and technical specifications. The conversion into July 2026 dollars uses the U.S. CPI-U all-items index. September 2020, October 2022 and January 2025 CPI values are aligned with the respective launch months, while July 2026 is the common valuation month. The Bureau of Labor Statistics reports a July 2026 CPI-U level of 333.918, based on 1982–1984 = 100. Consumer Price Index Summary, July 2026 – U.S. Bureau of Labor Statistics – August 2026 The inflation-adjustment formula is:

P2026,i = Plaunch,i × CPIJuly 2026 ÷ CPIlaunch month,i

Table 2.2 — Verified NVIDIA Flagship Launch History

ProductOfficial launchU.S. launch MSRPMemoryOfficial/derived bandwidthBoard powerJuly 2026-dollar MSRPReal price per GBReal price per GB/s
GeForce RTX 309024 Sep 2020$1,49924 GB GDDR6X≈936 GB/s350 W≈$1,923≈$80.13≈$2.05
GeForce RTX 409012 Oct 2022$1,59924 GB GDDR6X1,008 GB/s450 W≈$1,792≈$74.67≈$1.78
GeForce RTX 509030 Jan 2025$1,99932 GB GDDR71,792 GB/s575 W≈$2,102≈$65.69≈$1.17

Interpretation: the nominal launch price increased by 33.4% from the RTX 3090 to RTX 5090. In common July 2026 dollars, the estimated increase is approximately 9%. The real price per installed GB declined from approximately $80 to $66, and the real price per GB/s of memory bandwidth declined from approximately $2.05 to $1.17. The absolute power envelope increased substantially, however, meaning purchase-price efficiency cannot be treated as equivalent to operating-cost efficiency. The RTX 5090 provides more memory and bandwidth but requires a more demanding power supply, creates greater heat density and may raise total system cost. Its 32 GB capacity also remains far below the memory required for many high-parameter models at higher precision or with large KV caches.

NVIDIA officially specifies the RTX 5090 with 32 GB GDDR7 and 1,792 GB/s memory bandwidth. The RTX 4090 carries 24 GB GDDR6X, 1,008 GB/s bandwidth and a 450 W total graphics power rating. GeForce RTX 5090 Specifications – NVIDIA – January 2025 GeForce RTX 4090 Specifications – NVIDIA – September 2022 The RTX 3090 launch announcement records 24 GB GDDR6X and a starting price of $1,499. Introducing GeForce RTX 30 Series Graphics Cards – NVIDIA – September 2020

Table 2.3 — Change Across NVIDIA Flagship Generations

MetricRTX 3090 → RTX 4090RTX 4090 → RTX 5090RTX 3090 → RTX 5090
Nominal launch MSRP+6.7%+25.0%+33.4%
Installed memory0.0%+33.3%+33.3%
Memory bandwidth≈+7.7%+77.8%≈+91.5%
Board power+28.6%+27.8%+64.3%
Real launch price in July 2026 dollars≈−6.8%≈+17.3%≈+9.3%
Real price per GB≈−6.8%≈−12.0%≈−18.0%
Real price per GB/s≈−13.2%≈−34.3%≈−43.0%

The RTX sequence therefore supports a nuanced rather than categorical conclusion. The cost of entering the current flagship category rose sharply in nominal dollars, but specification-adjusted affordability improved on memory-capacity and bandwidth measures. At the same time, minimum absolute expenditure rose: a buyer still needs approximately $2,000 before tax for the reference flagship GPU alone, excluding the computer, high-capacity system memory, storage, cooling and power supply. Quality-adjusted deflation can coexist with worsening access inequality because a lower price per unit of capability does not help a household that cannot finance the larger minimum purchase.


2.4 AMD Consumer GPU Counterfactual

AMD provides an essential competitive comparator because changes in the NVIDIA flagship series cannot establish market-wide inflation. The Radeon RX 7900 XTX was announced with a suggested e-tail price of $999 and 24 GB GDDR6, while the RX 9070 XT launched with a suggested e-tail price of $599 and 16 GB GDDR6. AMD lists the RX 7900 XTX at 355 W total board power and the RX 9070 XT at 304 W. The RX 9070 XT has an official memory bandwidth of up to 640 GB/s. AMD Unveils World’s Most Advanced Gaming Graphics Cards, Built on Groundbreaking AMD RDNA 3 Architecture – AMD – November 2022 AMD Unveils Next-Generation RDNA 4 Architecture with the Launch of AMD Radeon RX 9000 Series Graphics Cards – AMD Investor Relations – February 2025 Radeon RX 9070 XT Specifications – AMD – 2025

Table 2.4 — Verified AMD Reference Products

ProductLaunch availabilityU.S. SEPMemoryBandwidthBoard powerNominal price per GBNominal price per GB/s
Radeon RX 7900 XTX13 Dec 2022$99924 GB GDDR6960 GB/s355 W$41.63$1.04
Radeon RX 7900 XT13 Dec 2022$89920 GB GDDR6800 GB/s315 W$44.95$1.12
Radeon RX 9070 XT6 Mar 2025$59916 GB GDDR6640 GB/s304 W$37.44$0.94
Radeon RX 90706 Mar 2025$54916 GB GDDR6640 GB/s220 W$34.31$0.86

These specification ratios must not be interpreted as equivalent LLM performance. NVIDIA and AMD differ in software maturity, supported data types, kernel optimisation, framework integration and application compatibility. A nominally lower price per GB is not an economic saving if the required application cannot use the hardware efficiently. Conversely, excluding non-NVIDIA hardware because the dominant software ecosystem favours NVIDIA would obscure the extent to which ecosystem lock-in, rather than physical semiconductor cost, determines effective price. Chapter 4 will therefore measure both hardware-neutral physical capability and workload-specific realised capability.


2.5 Professional Workstation GPUs

Professional GPUs differ from consumer products through certified drivers, memory capacity, error protection, virtualisation support, enterprise servicing, warranties and application certification. Their prices cannot be compared directly with GeForce or Radeon gaming cards. A professional card may provide lower nominal compute per dollar yet be economically preferable where downtime, certification or memory integrity has a high expected cost. The relevant formula is:

EACi = PurchasePricei + ExpectedDowntimeLossi + FailureRiskCosti + SoftwareIncompatibilityCosti − WarrantyValuei

Where EAC is the expected adjusted cost. A consumer GPU can appear substantially cheaper until the cost of unsupported deployment, business interruption or limited memory is included.

Table 2.5 — Professional GPU Data Requirements

VariableConsumer cardProfessional cardRequired adjustment
Memory capacityUsually lowerOften materially higherPrice per usable GB
ECC or error protectionLimited or product-dependentMore commonly supportedExpected error-cost adjustment
Driver certificationConsumer applicationsProfessional application certificationWorkflow-compatibility value
VirtualisationRestricted or limitedEnterprise-orientedMulti-user utilisation value
WarrantyConsumer termsProfessional support termsExpected service value
Availability horizonShorter retail generationOften longer enterprise lifecycleReplacement and fleet standardisation
Price disclosureUsually public MSRPOften reseller or quotation basedOfficial price gaps must remain ND
LLM suitabilityStrong if software-supportedStronger for memory-intensive workloadsTokens per second and concurrency test

A complete professional-GPU price history cannot be created from audited corporate filings because manufacturers do not consistently publish configuration-specific transaction prices. Quotation-only prices must not be converted into invented MSRPs. Chapter 2 therefore treats professional price measurement as a controlled data-acquisition task rather than filling the table with unverified retailer values.


2.6 Data-Centre Accelerators and AI Servers

Data-centre accelerators pose the greatest pricing-transparency problem. Public discussion frequently assigns unit prices to products such as NVIDIA H100, H200, B200 or AMD Instinct accelerators, but manufacturer transaction prices depend on volume, system configuration, networking, support, delivery terms and customer agreements. The price of a single accelerator is economically incomplete because large training and inference systems are sold as integrated platforms containing processors, HBM, networking, switches, cooling, storage, racks and software. Under the report’s source rules, unofficial estimates cannot be presented as verified prices.

Table 2.6 — Data-Centre Accelerator Price Transparency

Product categoryOfficial technical specificationsPublic manufacturer unit MSRPObservable audited average transaction priceAnalytical treatment
Individual AI acceleratorUsually availableGenerally NDNDSpecifications only
Multi-GPU serverConfiguration availableSometimes quotation-basedNDFull-system quotation required
Integrated rackArchitecture availableUsually NDNDPrice per delivered capacity
Cloud accelerator instancePublic tariff often availableService price, not hardware pricePartly observableCost per accelerator-hour
Reserved cloud capacityContract-specificRarely fully publicNDScenario or disclosed-contract analysis
Sovereign AI clusterProcurement-specificSometimes public through tenderPotentially observableTender-by-tender comparison

For data-centre hardware, the preferred price unit is not dollars per accelerator. It is the levelised cost per successful workload:

LCWw = [CAPEX × CRF + OPEXannual] ÷ SuccessfulWorkloadsannual,w

The capital-recovery factor is:

CRF = r(1 + r)n ÷ [(1 + r)n − 1]

Where r is the annual discount rate and n the economic life in years. CAPEX includes accelerator systems, networking, facility modifications and installation. OPEX includes power, cooling, maintenance, software, specialised labour, insurance and downtime. A lower accelerator purchase price may not reduce LCW if utilisation, software support or networking performance is inferior.

NVIDIA’s fiscal 2026 annual filing identifies the company as a full-stack computing-infrastructure provider and states that it depends on foundries and other suppliers for manufacturing. It does not provide a standard realised unit price for data-centre GPUs. NVIDIA Annual Report on Form 10-K for the Fiscal Year Ended January 25, 2026 – NVIDIA and U.S. Securities and Exchange Commission – February 2026 This absence is itself economically relevant: enterprise buyers operate in a negotiated market while households purchase transparent retail products.


2.7 Memory Inflation: HBM, DRAM and NAND

Memory must be separated into three markets. Conventional DRAM supplies system memory; NAND supplies persistent storage; HBM supplies the extreme bandwidth required by high-performance accelerators. Their prices do not move identically because their production complexity, demand, qualification and end markets differ. HBM consumes more wafer capacity per delivered bit than conventional DRAM and requires advanced packaging integration. A shift in producer capacity toward HBM can therefore tighten conventional-memory supply even when total memory investment increases.

The minimum measurement structure is:

MPIm,t = Σwj,0 × Pj,t ÷ Pj,0 × 100

Where MPI is the memory price index for memory class m, P is the verified contract or transaction price and w is the base-period quantity weight. A chained index is preferable when product generations change rapidly. An unweighted average of advertised module prices would be invalid because it would overrepresent products with many listings and ignore transaction volumes.

Table 2.7 — Memory-Market Measurement Framework

Memory segmentPrincipal AI roleAppropriate unitRequired quality controlsPrimary scarcity indicator
HBM2e/HBM3/HBM3E/HBM4Accelerator bandwidth and capacityPrice per stack, GB and GB/sGeneration, stack height, bandwidth and qualificationContracted capacity and delivery lead time
Server DRAMCPU-side model hosting, preprocessing and databasesPrice per GBDDR generation, speed, ECC and module typeContract price and inventory
Consumer DRAMLocal workstation system memoryPrice per GBDDR generation, speed and latencyRetail transaction price
Enterprise NANDModel storage, datasets and checkpointingPrice per TB and endurance unitInterface, endurance, latency and warrantyContract price and utilisation
Consumer NANDLocal model storagePrice per TBInterface, endurance and performanceRetail transaction price

Micron forecast in December 2025 that the HBM total addressable market could grow from approximately $35 billion in 2025 to around $100 billion in 2028, implying an approximate 40% compound annual growth rate. This is a company forecast, not an observed market outcome. Micron Fiscal First Quarter 2026 Financial Results – Micron Technology – December 2025 Samsung announced plans to invest more than KRW 110 trillion in facilities and R&D during 2026, including HBM, foundry operations and advanced packaging. Samsung Electronics Shareholder Value Enhancement Plan – Samsung Electronics – March 2026 These disclosures support rapid demand and investment but do not supply a public, auditable HBM price-per-GB history. Accordingly, the chapter does not invent an HBM price series.

Table 2.8 — What Can Be Verified for Memory

Requested variableHBMDRAMNAND
Official product generationYesYesYes
Manufacturer capacity commentaryYesYesYes
Audited segment revenueOftenOftenOften
Public transaction-weighted unit priceGenerally noGenerally noGenerally no
Public international street-price medianNoNoNo
Completed used-market medianNot meaningful for chips; noNoNo
Quality-adjusted public indexRequires external transaction datasetRequires datasetRequires dataset
Five-year projectionScenario-based onlyScenario-based onlyScenario-based only

2.8 CPUs, NPUs and the Problem of Incomparable AI Throughput

The emergence of integrated NPUs changes the definition of an AI-capable computer. A laptop may perform certain local inference tasks without a discrete GPU, but NPU TOPS cannot be compared directly with GPU TFLOPS or accelerator AI TOPS. The numerical value depends on precision, sparsity, supported operators, memory architecture and workload. A 50-TOPS NPU does not necessarily deliver one-fiftieth of a 2,500-TOPS accelerator’s useful performance, because the larger device may use a different precision, assume structured sparsity or perform workloads unsupported by the NPU.

The general quality-adjusted CPU/NPU price is:

PCPU-NPU,i,tQA = Psystem,i,t ÷ G(QCPU, QNPU, M, B, E, W)

Where G is a workload-specific capability aggregator, M is accessible memory, B bandwidth, E energy efficiency and W software compatibility. A geometric mean can prevent one very high component from dominating:

Gi = ∏Qk,iwₖ

With Σwk = 1.

Table 2.9 — CPU/NPU Comparison Controls

ControlReason
PrecisionINT4, INT8, FP8, FP16 and FP32 values are not interchangeable
SparsitySome advertised throughput assumes structured sparsity
Sustained powerPeak throughput may not be maintained in thin laptops
Accessible memoryShared memory may be large but bandwidth-limited
Operator supportUnsupported operators may fall back to CPU or GPU
Framework maturityTheoretical hardware may lack practical software paths
Batch sizeThroughput and latency respond differently
Model typeVision, audio and language workloads use hardware differently
Thermal conditionBattery and plugged-in performance may diverge
System priceIntegrated NPU cost cannot be isolated reliably from the device

2.9 Apple Computers: Price, Memory and Configuration Drift

Apple systems require configuration-level rather than brand-level analysis. Statements that “Apple raised all prices” are not scientifically acceptable unless identical or hedonic configurations are compared by geography and date. Apple has sometimes raised prices, held starting prices stable, reduced them or increased included memory. For example, the 2023 M2 Mac mini started at $599, while the 2024 M4 iMac started at $1,299 with 16 GB unified memory. These products are not longitudinal substitutes, but they demonstrate why company-wide price claims require product-family datasets. Apple Introduces New Mac mini with M2 and M2 Pro – Apple – January 2023 Apple Introduces New iMac Supercharged by M4 and Apple Intelligence – Apple – October 2024

The 2023 14-inch MacBook Pro with M2 Pro started at $1,999, while the 16-inch version started at $2,499. The M2 Max supported up to 96 GB unified memory with 400 GB/s bandwidth. Apple Unveils MacBook Pro Featuring M2 Pro and M2 Max – Apple – January 2023 The 2025 Mac Studio with M4 Max could be configured with 128 GB unified memory, while the M3 Ultra version provided a different higher-memory architecture. Mac Studio 2025 Technical Specifications – Apple – March 2025

Table 2.10 — Apple AI-Workstation Data Schema

FieldRequired configuration detail
Product familyMac mini, Mac Studio, MacBook Pro, iMac
ChipExact generation and tier
CPU/GPU coresExact selected configuration
Unified memoryInstalled capacity
Memory bandwidthChip-specific bandwidth
StorageInstalled SSD capacity
Launch pricePrice for identical configuration
Current priceSame geography, tax treatment and configuration
AI workloadModel, quantisation, context and inference engine
Tokens per secondMedian sustained result after warm-up
EnergyWall-power measurement
WarrantyStandard and extended coverage
UpgradeabilityMemory and storage replacement constraints
Residual valueVerified completed resale transaction only

The economic attraction of high-memory Apple systems is that unified memory can permit models too large for a 24 GB or 32 GB consumer GPU. The limitation is that capacity does not establish speed, framework compatibility or enterprise manageability. Price per GB therefore favours high-memory unified systems only when the workload can exploit the architecture at an acceptable rate.


2.10 Samsung Devices and Relevant Competitors

Samsung spans memory production, mobile devices, laptops and semiconductor manufacturing, making it economically distinct from a pure device vendor. However, the company does not publish a single globally comparable AI-computer price series in audited financial reports. Model names, processors, memory configurations, launch timing, promotional bundles and taxes differ by country. A valid Samsung comparison must therefore use exact SKU-level official launch prices and exclude temporary trade-in credits unless trade-in value is separately modelled.

Table 2.11 — Required Samsung and Competitor Sample

CategorySamsung product familyRequired competitorsComparison unit
Premium AI laptopGalaxy Book Ultra/Pro classApple MacBook Pro, Dell Precision/XPS, Lenovo ThinkPad/Legion, HP ZBookCost per successful local workload
Mobile AI deviceGalaxy S Ultra classApple iPhone Pro, Google Pixel ProOn-device AI capability per real price
MemorySamsung DRAM/NAND/HBMMicron, SK hynixContract-price index by memory generation
Enterprise storageSamsung enterprise SSDMicron and other qualified vendorsCost per TB adjusted for endurance
AI workstation componentSamsung memory-equipped systemsCompeting memory configurationsCost per usable GB and bandwidth

No Samsung price increase is treated as established unless an identical configuration and geography are documented at two dates. Product-mix changes—such as more memory, different processors or bundled AI features—must be removed through hedonic adjustment.


2.11 High-Performance Laptops and Desktop Workstations

Laptop prices present additional measurement problems because the GPU name does not fully specify performance. Mobile GPU power limits vary by manufacturer and chassis; a laptop GPU with the same family designation as another may run at a different power envelope. Cooling, system memory, screen, battery, storage and warranty also affect price. The correct unit is a complete SKU, not a nominal GPU class.

Table 2.12 — Mandatory Laptop Controls

VariableMeasurement requirement
GPU modelExact mobile GPU designation
GPU powerSustained configured power, not only maximum permitted range
VRAMCapacity and type
System RAMCapacity, bandwidth and upgradeability
CPU/NPUExact model and enabled power profile
CoolingSustained benchmark after thermal equilibrium
Battery modeSeparate from plugged-in result
DisplayResolution and refresh rate as price controls
StorageCapacity and performance
WeightPortability control
WarrantyStandardised expected service value
AI benchmarkSame model, precision, context and inference engine
GeographyNet-of-tax and tax-inclusive price
PromotionsRecorded separately from standard price

For desktop workstations, the chassis, motherboard, power supply, cooling, memory, storage and professional support must be added to GPU cost. A $1,999 GPU does not create a $1,999 local AI system. The workstation price is:

Pworkstation = PGPU + PCPU + PRAM + Pstorage + Pboard + PPSU + Pcooling + Pcase + POS + Passembly + Pwarranty


2.12 Quality-Adjusted Price Measures

No single denominator captures AI capability. This chapter therefore adopts six complementary measures.

2.12.1 Price per Usable Accelerator-Memory GB

Pi,tVRAM = Pi,t ÷ Musable,i,t

Usable memory must exclude operating overhead, display reservation and unusable fragmentation. Installed memory is acceptable only for the preliminary hardware table.

2.12.2 Price per GB/s of Memory Bandwidth

Pi,tBW = Pi,t ÷ Bi,t

This is relevant for memory-bound inference but does not capture compute, caching or software.

2.12.3 Price per Comparable TFLOP

Pi,tFLOP = Pi,t ÷ Fi,t,p

The precision p must be identical. FP32 cannot be divided by an FP16 or FP8 throughput value.

2.12.4 Price per Token per Second

Pi,t,m,cTPS = Psystem,i,t ÷ TPSi,m,c

Where m is the model and c the complete benchmark configuration.

2.12.5 Price per Benchmark Unit

Pi,tBENCH = Pi,t ÷ Scorei,t

Only independently reproducible benchmarks with identical settings qualify.

2.12.6 Price per Successfully Completed Workload

Pi,wSUCCESS = TCOi,w ÷ Nsuccessful,i,w

This is the preferred economic metric because it includes reliability, retries, supervision and total cost.

Table 2.13 — Strengths and Limitations of Quality Metrics

MetricPrincipal strengthPrincipal weakness
Price per installed GBSimple and relevant to model fitIgnores speed and usable capacity
Price per usable GBBetter model-capacity measureRequires runtime-level measurement
Price per GB/sCaptures memory movementIgnores compute and software
Price per TFLOPCaptures arithmetic throughputPrecision and sparsity can mislead
Price per token/sDirect inference measureModel and settings specific
Price per benchmark pointEnables composite comparisonBenchmark design can bias results
Price per successful taskClosest to economic utilityExpensive to measure and workload-specific

2.13 Hedonic Price Model

The principal hedonic specification is:

ln(Pi,r,t) = α + Σk=1KβkXk,i,t + γt + δb + θr + λc + εi,r,t

Where:

  • Pi,r,t is the net-of-tax transaction price;
  • Xk,i,t contains technical characteristics;
  • γt is the time fixed effect;
  • δb is the brand fixed effect;
  • θr is the region fixed effect;
  • λc is the condition class: new, refurbished or used;
  • εi,r,t is the residual.

The feature vector will include:

  • usable accelerator memory;
  • memory bandwidth;
  • comparable throughput by precision;
  • power consumption;
  • age;
  • product class;
  • warranty;
  • system RAM;
  • storage;
  • portability;
  • professional certification;
  • software ecosystem;
  • condition;
  • seller type;
  • regional tax;
  • exchange rate;
  • verified availability.

The time dummy exponent produces a quality-adjusted price index:

HPIt = 100 × exp(γt − γbase)

If the model is estimated in log form, percentage interpretation must use exp(β) − 1 rather than treating large coefficients as exact percentages.

Table 2.14 — Model Diagnostics

DiagnosticTestRequired response
HeteroskedasticityBreusch–Pagan or WhiteRobust standard errors
MulticollinearityVIF and condition numberCombine or remove redundant hardware variables
Serial correlationTime-clustered residual testCluster by product family and period
Brand endogeneityFixed effects and matched productsReport sensitivity
Sample selectionAvailability and seller controlsWeight or model selection
Functional formRESET and spline comparisonUse nonlinear terms where justified
OutliersLeverage and Cook’s distanceVerify, winsorise only under disclosed rule
Missing dataMissingness mapMultiple imputation only where defensible
Product turnoverChained indexAvoid comparing unmatched generations directly
Benchmark driftFixed workload suiteFreeze model and settings within test window

2.14 Inflation Adjustment and Geographic Comparability

The real price in reference-period currency is:

Pi,t→Treal = Pi,t × CPIT ÷ CPIt

The real price change is:

πi,t→Treal = [Pi,Tnominal ÷ Pi,t→Treal − 1] × 100

For cross-border comparison:

Pi,r,tcommon = [Pgross,i,r,t ÷ (1 + VATr,t)] ÷ FXr,t

Household affordability retains VAT:

HAIi,r,t = Pgross,i,r,t ÷ MedianMonthlyDisposableIncomer,t

Enterprise affordability removes recoverable VAT where applicable:

EAIi,r,t = Pnet,i,r,t ÷ MedianMonthlyEnterpriseValueAddedr,t

Table 2.15 — Geographic Normalisation Protocol

MarketConsumer inflation sourcePrice treatmentCurrency treatmentPrincipal distortion
United StatesBLS CPI-UUsually exclude sales tax from MSRP, add local tax for affordabilityUSD baseState and local sales-tax variation
European UnionEurostat HICPRecord VAT-inclusive and net-of-VATEUR and local currency where applicableDifferent VAT and regional official pricing
United KingdomONS CPIVAT-inclusive and net-of-VATGBP, then common currencyGBP movements and VAT
ChinaOfficial national CPI requiredInclude applicable taxesCNY and common currencyLocal SKU and channel differences
JapanOfficial national CPI requiredInclude consumption taxJPY and common currencyYen volatility
IndiaOfficial national CPI requiredInclude GST and import effectsINR and PPP comparisonImport duties and income differential
BrazilOfficial national CPI requiredInclude indirect taxes and import costsBRL and PPP comparisonTax complexity and currency volatility
Lower-income sampleNational statistics or official international seriesConsumer gross priceLocal currency, USD and PPPThin formal distribution and income constraint

Eurostat defines the HICP as a comparable measure of changes in household consumer prices across European countries. Harmonised Indices of Consumer Prices – Eurostat – 2026 The U.S. Bureau of Labor Statistics maintains a dedicated index for computers, peripherals and smart-home-assistant devices and applies quality adjustment to changing computer products. How BLS Measures Price Change for Computers, Peripherals and Smart Home Assistant Devices – U.S. Bureau of Labor Statistics – February 2026


2.15 Used and Refurbished Market

The used-market premium must be calculated only from completed, comparable transactions:

UMPi,r,t = [Pused,i,r,t − Preference,i,r,t] ÷ Preference,i,r,t × 100

The reference price depends on the research question:

  • launch MSRP tests whether used price exceeds original nominal price;
  • inflation-adjusted MSRP tests whether it exceeds original real price;
  • contemporaneous new price tests whether used hardware is irrationally priced relative to new supply;
  • depreciation-adjusted fair value tests whether scarcity offsets ordinary ageing;
  • replacement-capability price tests whether the old product is cheaper than an equivalent current product.

A two-times-MSRP claim is:

MultiplierMSRP = Pused ÷ Plaunch MSRP

But this is not sufficient evidence of market inflation. A discontinued limited product may have collector value; a listing may never sell; the product may include a complete system; tax may be included in one price and excluded in the other; or the listing may be fraudulent.

Table 2.16 — Used-Market Cleaning Rules

IssueRequired treatment
Asking versus sold priceRetain only completed transactions
Duplicate listingDeduplicate by identifier, seller, image and timestamp
BundleSeparate GPU-only from complete system
ConditionNew-sealed, open-box, refurbished, used-working and parts-only
WarrantyRecord transferable remaining coverage
Fraud/anomalyExclude under a disclosed, reproducible rule
ShippingSeparate product price from delivery
TaxRecord separately
CurrencyConvert using transaction-date official FX
Board partnerRecord exact manufacturer and model
Factory overclockTreat as product attribute
LocationUse actual shipping origin and destination
Transaction volumeReport count and confidence interval
Thin marketsSuppress median below minimum observation threshold

Table 2.17 — Statistical Test for Persistent Used Scarcity

TestNull hypothesisEvidence supporting persistent scarcity
Median premium testMedian UMP ≤ 0Median UMP significantly above zero
Duration testPremium lasts no longer than launch periodPremium persists beyond defined threshold
Cross-region testPremium is geographically isolatedPremium observed across multiple tax-adjusted markets
Availability testPremium unrelated to new-product availabilityPremium rises when verified availability falls
Transaction-volume testPremium driven by thin tradesPremium remains with adequate volume
Event studyPremium unrelated to supply eventPremium responds to launch, export or supply shock
Mean-reversion testPremium is persistentSlow or incomplete reversion
Replacement testUsed product is not scarce in capability termsEquivalent new alternatives remain more expensive or unavailable

Determination

Under the current source restrictions, the statement that AI-capable used products generally sell for two or three times original MSRP cannot be verified. The available primary sources establish MSRP and specifications, but not representative used transaction medians. The claim must therefore remain unconfirmed, not accepted and not rejected. Isolated cases may exist, especially during launch shortages or in geographically constrained markets, but isolated cases cannot be generalised to the entire used market.


2.16 Warranty and Risk-Adjusted Used Price

A used product’s nominal price understates its economic cost if warranty, failure probability and remaining useful life differ from a new product. The risk-adjusted price is:

Pused,irisk = Ptransaction,i + pfailure,i × Lfailure,i + Cdowntime,i − Vwarranty,i

Where:

  • pfailure,i is expected failure probability;
  • Lfailure,i is replacement or repair loss;
  • Cdowntime,i is the economic cost of unavailability;
  • Vwarranty,i is remaining warranty value.

For a student, downtime may mean lost study time. For a hospital, defence contractor or financial institution, it may be unacceptable. The same used GPU can therefore have different risk-adjusted prices for different users.


2.17 Preliminary 2020–2026 Findings

Table 2.18 — Evidence-Based Findings

QuestionPreliminary resultConfidence
Did flagship NVIDIA MSRP rise nominally?Yes: $1,499 to $1,999, approximately +33.4%High
Did flagship NVIDIA MSRP rise after general inflation?Moderately, approximately +9% using common July 2026 dollarsMedium-high
Did real price per GB rise?No in the selected flagship sequence; it declinedHigh for specification ratio
Did real price per GB/s rise?No in the selected flagship sequence; it declined materiallyHigh for specification ratio
Did absolute minimum flagship expenditure rise?YesHigh
Did power requirements rise?Yes, substantiallyHigh
Did practical LLM cost necessarily decline?Not establishedOpen
Did AMD maintain lower reference prices in selected products?Yes, but software and workloads are not equivalentHigh
Did HBM demand and investment rise?Strong corporate evidence supports rapid expansionHigh
Is there a verified public HBM transaction-price history?No under permitted sourcesHigh
Are data-centre accelerator unit prices transparent?Generally noHigh
Are international used-market medians verified?NoHigh
Are two-times or three-times used prices representative?Not demonstratedOpen

The strongest finding is therefore not universal hardware inflation but a widening divergence between absolute access cost and quality-adjusted unit cost. A flagship GPU can deliver more memory bandwidth per real dollar while remaining inaccessible to more households because its minimum purchase price has increased. This is the threshold effect:

Accessi = 1 if Wealthi or CreditCapacityi ≥ MinimumSystemCost; otherwise 0.

A falling price per unit of capability does not reduce exclusion when the indivisible minimum system remains above the buyer’s budget.


2.18 Projection Framework, 2027–2031

A scientifically valid five-year forecast cannot be produced by extending MSRP alone. The model must forecast:

  • nominal hardware price;
  • general inflation;
  • quality growth;
  • memory capacity;
  • bandwidth;
  • power;
  • availability;
  • software compatibility;
  • workload performance;
  • depreciation;
  • used-market liquidity;
  • regional taxation and FX.

The quality-adjusted price evolves as:

Pt+1QA = PtQA × (1 + gnominal,t) ÷ [(1 + πt)(1 + gquality,t)]

A supply-demand specification is:

Δln(Pt) = α + β₁Δln(Demandt) − β₂Δln(Capacityt) + β₃EnergyShockt + β₄TradeShockt + β₅FXShockt + εt

The 2027–2031 values below are not claimed as observed or externally certified forecasts. They are transparent scenario indices with 2026 = 100, intended to show how different annual quality-adjusted price paths compound.

Scenario assumptions

  • Abundant-supply scenario: quality-adjusted price declines 10% annually.
  • Central-efficiency scenario: quality-adjusted price declines 3% annually.
  • Persistent-scarcity scenario: quality-adjusted price rises 8% annually.
  • Geopolitical-fragmentation scenario: quality-adjusted price rises 12% annually in import-dependent markets.
  • Dual-market scenario: frontier hardware rises 8% annually while mainstream capability declines 8% annually.

Table 2.19 — Quality-Adjusted Hardware Price Index, 2026 = 100

Scenario202620272028202920302031
Abundant supply100.090.081.072.965.659.0
Central efficiency100.097.094.191.388.585.9
Persistent scarcity100.0108.0116.6126.0136.0146.9
Geopolitical fragmentation100.0112.0125.4140.5157.4176.2
Mainstream side of dual market100.092.084.677.971.665.9
Frontier side of dual market100.0108.0116.6126.0136.0146.9

Table 2.20 — Scenario Interpretation

ScenarioSemiconductor conditionMemory conditionUsed marketAccess consequence
Abundant supplyCapacity expansion exceeds demandHBM and DRAM supply normalisesRapid depreciationBroadening access
Central efficiencySupply broadly tracks demandPeriodic but manageable constraintNormal depreciation with launch spikesGradual affordability improvement
Persistent scarcityDemand repeatedly outruns capacityHBM/packaging remain bindingPremiums persist for high-memory productsRising absolute and real burden
Geopolitical fragmentationCapacity divided by controls and blocsRegional allocation differs sharplyImport-dependent used markets tightenLarge geographic inequality
Dual marketMainstream improves; frontier remains scarcePremium memory concentrated in frontier systemsHigh-memory older systems retain valueNominal access broadens, capability gap widens

The dual-market scenario is the most important. Technological progress may reduce the cost of small and medium models while frontier capability becomes more expensive because it requires larger memory pools, more complex packaging, liquid cooling, advanced networking and higher power density. Under that scenario, statements that “AI hardware prices are falling” and “frontier AI is becoming unaffordable” can both be true.


2.19 Geographic Exposure, 2027–2031

Table 2.21 — Geographic Price-Transmission Mechanisms

MarketMain strengthsMain exposureExpected source of premium
United StatesLarge cloud market, direct vendor channels, capital accessLocal scarcity, energy and trade-policy effectsLaunch availability and state tax
European UnionLarge market and public compute investmentVAT, energy price, limited frontier-hardware productionVAT, FX, regional allocation and energy
United KingdomMature technology marketSterling volatility and import dependenceFX, VAT and distribution
ChinaLarge domestic technology baseAccess restrictions to highest-end foreign acceleratorsProduct substitution and controlled performance
JapanStrong electronics ecosystemYen volatility and imported accelerator dependenceFX and local official pricing
IndiaLarge skills base and growing digital marketIncome constraint, taxes and imported hardwareImport cost and affordability
BrazilLarge marketCurrency volatility, taxes and import costsTax and distribution
Lower-income marketsPotential open-model adoptionIncome, electricity, finance, thin distribution and warrantySmall volume, FX, credit and logistics

For lower-income markets, a product can become cheaper in U.S. dollars while becoming less affordable locally. The affordability change is:

ΔAffordabilityr = ΔLocalHardwarePricer − ΔMedianDisposableIncomer

If local hardware prices rise 5% but disposable income rises only 2%, affordability deteriorates by approximately 3%, even if quality improves. The later country analysis will therefore report three outcomes simultaneously: dollars per capability unit, months of median income required and access to financing.


2.20 Final Assessment of H1, H2 and H3

H1: Prices of AI-relevant hardware have risen faster than general consumer-electronics prices

Status: not yet confirmed across the full market. The selected NVIDIA flagship MSRP sequence shows a nominal increase and a smaller real increase, but it does not establish a market-wide index. AMD reference products and Apple configuration changes demonstrate that product-level paths differ. A complete H1 determination requires the hedonic dataset.

H2: The inflation-adjusted cost of a fixed level of AI capability has increased

Status: unresolved and unlikely to have a universal answer. The selected NVIDIA sequence shows lower real prices per GB and per GB/s, but these are incomplete capability measures. Fixed-workload testing may show declining cost for smaller models and rising cost for privacy-sensitive, high-memory or frontier workloads.

H3: Used-market prices indicate persistent scarcity rather than temporary speculation

Status: unproven under available primary data. No representative, verified international completed-transaction dataset has been identified within the permitted evidence hierarchy. Claims of two-times or three-times MSRP cannot be generalised.


2.21 Chapter Conclusion

The verified evidence does not support a simplistic conclusion that every form of AI hardware became two or three times more expensive between 2020 and 2026. It supports a more consequential finding: the market is splitting between improving technical price-performance and increasing minimum access costs. The NVIDIA flagship launch price rose from $1,499 to $1,999, but memory capacity and bandwidth increased sufficiently to lower simple real price-per-capability ratios. The improvement does not eliminate exclusion because a user must still finance the entire device and supporting system. The result resembles a high-speed rail ticket becoming cheaper per kilometre while the minimum fare rises beyond the budget of some passengers.

Memory and data-centre hardware present a different problem: price transparency is substantially weaker. HBM demand, capacity commitments and investment are documented, but public transaction prices are not. Data-centre accelerators are commonly sold through negotiated systems and cloud arrangements, preventing a reliable public unit-price history. Used-market analysis is weaker still because asking prices are frequently mistaken for completed sales. Under an academic evidentiary standard, those gaps must remain visible.

The principal five-year risk is therefore a dual market. Mainstream AI capability may become cheaper through efficiency, open models and better consumer hardware, while frontier, private and high-concurrency capability becomes more expensive because of HBM, packaging, networking, electricity and institutional demand. Such a market would not exclude poorer users from AI entirely. It would confine them to a lower capability tier. That distinction—between access to some AI and access to economically competitive AI—is the central bridge from hardware inflation to Chapters 4, 7 and 8.

Chapter 3 — Industrial Concentration and the Architecture of Scarcity

3.1 Analytical Objective

The AI industrial system is not a single market governed by one concentration ratio. It is a chain of technically complementary and economically interdependent markets spanning semiconductor-design software, intellectual property, lithography, fabrication, memory, packaging, substrates, accelerator design, networking, server assembly, cloud infrastructure, foundation models and enterprise distribution. Concentration at one layer can be offset by substitution or intensified by dependence on another. A company may possess strong accelerator-design capabilities while relying on an external foundry, HBM suppliers, packaging capacity and cloud customers. A foundry may control leading-edge production while remaining exposed to a small number of customers and equipment suppliers. A cloud provider can operate enormous infrastructure but depend on accelerators designed and fabricated elsewhere. The appropriate analytical object is therefore not “the AI monopoly,” but a network of differentiated bottlenecks whose control, scarcity, substitution time and contractual allocation jointly determine effective AI capacity.

This chapter makes five distinctions. First, market concentration measures how economic activity is distributed among firms. Second, technical concentration measures how many suppliers can satisfy a particular performance requirement. Third, dependency concentration measures whether downstream firms rely on the same upstream node. Fourth, allocation concentration measures whether future output is reserved by a small number of purchasers. Fifth, economic rent measures returns exceeding the amount required to retain capital and capability in their current use. None of these concepts independently establishes unlawful conduct. A firm can earn high margins because it created a superior product, assumed substantial technological risk, invested before demand materialised or temporarily controls scarce capacity. Conversely, a market can exhibit exclusionary effects even when accounting margins are not extraordinary, particularly where interoperability restrictions, licensing terms or long-term capacity commitments raise rivals’ costs.

The European Commission identifies data, AI accelerator chips, computing infrastructure, cloud capacity and technical expertise as potentially critical inputs and possible barriers to entry in generative-AI markets. Its analysis explicitly treats bottlenecks as context-dependent: a constrained input can reduce competition without every constraint constituting an anticompetitive practice. Competition Policy Brief: Competition in Generative AI and Virtual Worlds – European Commission – September 2024


3.2 Concentration Metrics and Interpretation

The n-firm concentration ratio is:

CRn = Σi=1nsi

Where si is firm i’s market share. CR1 measures the largest firm; CR3 and CR4 measure the combined share of the three or four largest firms. The Herfindahl–Hirschman Index is:

HHI = Σi=1N(100si)2

When market shares are expressed as decimal fractions. If shares are already expressed as percentage points, the equivalent formula is:

HHI = Σi=1Nsi2

The HHI ranges from close to zero in an extremely fragmented market to 10,000 in a pure single-firm market. Five equally sized firms generate an HHI of 2,000; four equal firms generate 2,500; three equal firms generate approximately 3,333; two equal firms generate 5,000.

Under the current U.S. Merger Guidelines:

  • HHI above 1,800 indicates a highly concentrated market;
  • an increase exceeding 100 points is considered significant;
  • a transaction producing both conditions can generate a structural presumption that competition may be substantially lessened.

These are merger-screening thresholds, not a legal declaration that every market above 1,800 is unlawful. Market definition, entry, innovation, buyer power, substitution and efficiencies remain essential. 2023 Merger Guidelines – U.S. Department of Justice and Federal Trade Commission – December 2023

Table 3.1 — Mathematical Meaning of HHI

Market structureEqual-share assumptionHHI
One supplier100%10,000
Two suppliers50% each5,000
Three suppliers33.33% each3,333
Four suppliers25% each2,500
Five suppliers20% each2,000
Six suppliers16.67% each1,667
Ten suppliers10% each1,000
Twenty suppliers5% each500

Equal shares produce the minimum HHI for a fixed number of firms. If one supplier is larger, HHI rises. A nominal three-firm market therefore cannot have an HHI below approximately 3,333 unless additional suppliers are included in the correctly defined market.


3.3 Why Conventional Market Shares Are Insufficient

Market shares can understate AI concentration for four reasons. First, revenue-based shares do not necessarily measure available frontier capacity. A supplier can have a modest share of total semiconductor revenue but control a much larger share of a technically indispensable advanced segment. Second, nominal competitors may not be substitutes. An accelerator that cannot run the required framework, fit the model or satisfy export restrictions does not constrain price effectively. Third, the same upstream node may support several nominal downstream competitors. Multiple cloud providers can offer different services while relying on the same foundry, lithography equipment or memory suppliers. Fourth, long-term contracts can allocate most future output before it enters any observable spot market.

The report therefore supplements CRn and HHI with four indices:

Supplier Dependency Index:

SDIj = Σi=1Nwidi,j

Where di,j is firm i’s dependence on supplier j and wi is firm i’s downstream importance.

Effective Substitutability Index:

ESIj = 1 ÷ [1 + Tswitch,j + Cswitch,j + Lperformance,j]

Where T is switching time, C is normalised switching cost and L is performance loss.

Reserved Capacity Ratio:

RCRj,t = Capacitycontracted,j,t ÷ Capacitytotal,j,t

Bottleneck-Adjusted Concentration:

BACj = HHIj × Criticalityj × [1 − ESIj]

A large HHI in a non-critical, easily substituted component may have limited systemic effect. A somewhat lower HHI in an indispensable input with multi-year qualification requirements can be more consequential.


3.4 Full AI Industrial-Control Map

Table 3.2 — Principal Firms by AI Supply-Chain Layer

LayerPrincipal firms or platformsFunctionCentral dependency
Semiconductor EDASynopsys, Cadence, Siemens EDA; specialised vendorsDesign, verification, physical implementation and simulationProprietary formats, process-design kits and engineering workflows
Semiconductor IPArm, Synopsys, Cadence, Imagination and specialised IP providersCPU, interface, memory and connectivity IPLicensing, standards and integration
EUV lithographyASMLPatterning of leading-edge semiconductor layersExtremely complex equipment and supplier network
DUV lithographyASML, Nikon, CanonMature and selected advanced process stepsTool capability, throughput and process qualification
Deposition and etchApplied Materials, Lam Research, Tokyo Electron and specialised firmsMaterial deposition and pattern transferProcess recipes, service network and qualification
Metrology and inspectionKLA, ASML and specialised firmsDefect detection, yield control and process measurementYield learning and process control
Leading-edge foundryTSMC, Samsung Foundry, Intel Foundry; selected Chinese domestic developmentAdvanced logic fabricationProcess maturity, yield, tool access and customer qualification
Mature-node foundryTSMC, GlobalFoundries, UMC, SMIC, Hua Hong, Tower and othersSupporting logic, control and interface chipsGeographic capacity and application qualification
HBMSK hynix, Samsung Electronics, MicronHigh-bandwidth accelerator memoryYield, stack technology, packaging integration and qualification
Conventional DRAMSamsung, SK hynix, Micron; additional regional suppliersSystem and server memoryCyclical capacity allocation
NANDSamsung, SK hynix/Solidigm, Kioxia, Western Digital/SanDisk, Micron and othersPersistent model and dataset storageLayer transitions, controller integration and cyclicality
Advanced packagingTSMC, ASE, Amkor, Samsung, Intel and other OSAT providersLogic-memory integration and chiplet packagingCoWoS-equivalent capacity, substrates and thermal design
SubstratesIbiden, Shinko, Unimicron, Nan Ya PCB, AT&S and othersPhysical package interconnectionQualification, yield and expansion lead time
Silicon interposersFoundries and packaging specialistsHigh-density connection between logic and HBMLarge-die processing and packaging integration
GPU designNVIDIA, AMD, Intel and regional challengersGeneral accelerated computingArchitecture, software and foundry access
Custom AI acceleratorsGoogle TPU, AWS Trainium/Inferentia, Microsoft Maia, Meta silicon programmes and othersWorkload-specific cloud accelerationInternal cloud scale and software integration
NetworkingNVIDIA, Broadcom, Marvell, Cisco, Arista and othersScale-up and scale-out AI interconnectBandwidth, latency, optics and switching
AI serversSupermicro, Dell, HPE, Lenovo, Inspur and ODMs including Quanta, Wiwynn and FoxconnIntegration of accelerators, CPUs, memory and coolingGPU allocation, rack power and liquid cooling
Public cloudAWS, Microsoft Azure, Google Cloud, Oracle and othersMetered compute and AI infrastructureCapital expenditure, installed capacity and customer ecosystem
Frontier modelsOpenAI, Google DeepMind, Anthropic, Meta, xAI and other regional developersFoundation-model creation and inferenceCompute, data, talent, capital and distribution
Enterprise AI distributionMicrosoft, Google, AWS, Salesforce, Oracle, SAP, ServiceNow, Adobe and specialist firmsEmbedding AI into enterprise workflowsInstalled software base, customer data and switching cost

The table identifies principal actors, not formal market shares. In most layers, audited corporate reports do not publish a complete common-denominator market dataset. Consequently, the presence of named firms must not be mistaken for a calculated CR4.


3.5 Lithography: The Highest-Order Equipment Bottleneck

Lithography is the most structurally concentrated upstream layer because leading-edge fabrication depends on exceptionally complex equipment with long development cycles, extensive intellectual property and thousands of specialised suppliers. ASML reported €32.7 billion in 2025 net sales, a 52.8% gross margin, €11.3 billion operating income, €9.6 billion net income, €4.7 billion R&D expenditure and a supplier network of approximately 5,100 companies. It sold 48 EUV systems, 279 DUV systems and 208 metrology and inspection systems during the year. ASML 2025 Annual Report – ASML – January 2026

The verified primary dataset identifies ASML as the supplier reporting commercial EUV-system sales, but it does not provide an audited global market-share denominator constructed under a competition-law market definition. A formal CR1 and HHI are therefore not stated as observed statistics. The conditional mathematical result is nevertheless clear: if the relevant market is defined as commercial EUV lithography systems and ASML supplies 100%, then:

CR1 = 100%

HHI = 1002 = 10,000

That would be a technical single-source market. It would not mean ASML can act without constraint. Its power is limited by semiconductor-cycle demand, customer capital budgets, export licensing, supplier capacity, technological execution and the risk that customers delay node transitions. Its exceptionally high R&D expenditure and supplier dependence also show why accounting margin cannot be classified entirely as monopoly rent.

Table 3.3 — Lithography Entry Barriers

BarrierSeverityExplanation
Optical and light-source complexityExtremeLeading-edge lithography requires capabilities accumulated over decades
Precision manufacturingExtremeSystem performance depends on sub-nanometre control and exceptionally complex integration
Supplier networkExtremeThousands of specialised suppliers contribute irreplaceable subsystems
Customer qualificationExtremeFoundries must integrate tools into complete process flows
Intellectual propertyExtremePatents, trade secrets and cumulative engineering knowledge
Capital requirementExtremeDevelopment and production require sustained multi-billion investment
Service infrastructureVery highInstalled tools require continuous support and calibration
Export licensingHighGovernment controls can restrict addressable markets
Time to credible entryVery longEntry would require technological, manufacturing and customer-validation cycles

3.6 Electronic-Design Automation and Semiconductor IP

EDA tools sit upstream of every advanced chip. They support architecture design, verification, timing, physical implementation, power analysis and sign-off. Their economic importance exceeds their direct share of semiconductor expenditure because an advanced design cannot be fabricated reliably without validated tools and foundry-compatible process-design kits. Switching costs include retraining engineers, rebuilding design flows, validating scripts, converting libraries, repeating verification and accepting schedule risk. These costs create ecosystem dependence even where alternative tools exist.

The European Commission’s review of Synopsys’ acquisition of Ansys identified competition concerns in three global software markets and required divestitures covering the relevant overlaps. The Commission described Synopsys as an EDA and semiconductor-IP supplier and Ansys as a simulation and analysis provider whose semiconductor tools overlapped with parts of Synopsys’ portfolio. Buying Ansys: A Synopsys of a Chip Design Story – European Commission – 2025

No complete primary-source revenue shares for the entire EDA market were identified under the permitted source hierarchy. Therefore:

  • CR3: ND;
  • CR4: ND;
  • HHI: ND.

A three-firm illustration must not be misreported as an estimate. If a correctly defined EDA submarket contained exactly three equal firms, its minimum HHI would be approximately 3,333. Actual concentration depends on tool category, specialised competitors, internal tools and the relevant geographic market.

Table 3.4 — EDA Lock-In Mechanisms

MechanismEconomic effectObservable test
Foundry-certified design flowLimits tools that can reach production sign-offNumber of certified alternatives by process node
Proprietary file and scripting formatsRaises migration costConversion time and defect rate
Engineer skill concentrationMakes alternative adoption expensiveLabour-market skill distribution
Integrated tool suiteCreates scope efficienciesSeparate versus integrated workflow cost
Semiconductor IP bundleReinforces tool adoptionShare of designs using bundled IP
Historical design librariesEmbeds prior investmentCost of library requalification
Long design cyclesMakes mid-project switching impracticalSwitching frequency by project stage
Sign-off liabilityFavours incumbent validated toolsCustomer acceptance and insurance requirements

3.7 Leading-Edge Foundry Capacity

Leading-edge foundry concentration results from capital intensity, process complexity, yield learning, tool access and customer qualification. TSMC, Samsung and Intel represent the principal global organisations pursuing the most advanced logic processes, although their commercially available capacity, yields and customer models differ. Mature-node capacity is more diversified and must not be included in a leading-edge denominator merely because it fabricates semiconductors.

TSMC reported 2025 revenue of approximately $122 billion, up 35.9% in U.S.-dollar terms. Its 2025 gross margin was 59.9%, operating margin 50.8% and net margin 45.1%. The company attributed the margin increase partly to higher capacity utilisation and cost improvement, while overseas fabrication expansion diluted profitability. TSMC 2025 Annual Report – Taiwan Semiconductor Manufacturing Company – April 2026 In the second quarter of 2026, process technologies at 7 nanometres and below accounted for 77% of TSMC wafer revenue, while HPC represented 66% of quarterly revenue. These are TSMC internal revenue shares, not global foundry shares. TSMC Second Quarter 2026 Earnings Conference Transcript – Taiwan Semiconductor Manufacturing Company – July 2026

Primary audited materials available in this research session do not provide a harmonised 2025 global leading-edge foundry share series. Consequently:

  • CR1: ND;
  • CR3: ND;
  • HHI: ND.

Table 3.5 — Foundry Barriers

BarrierLeading-edge effectSubstitution period
Fabrication-plant capitalRestricts entry and simultaneous expansionMulti-year
Yield learningMakes nominal process availability different from economic capacityMulti-quarter to multi-year
Customer qualificationPrevents instant design migrationOften multiple quarters
Process-design kitsTies design tools and libraries to a nodeProject-specific
Advanced equipmentDepends on lithography, deposition, etch and metrology suppliersMulti-year
Skilled workforceLimits rapid geographic replicationMulti-year
Utility requirementLarge, stable electricity and ultrapure-water demandLocation-specific
Export controlsRestricts equipment and customer accessPolicy-dependent
Packaging integrationLogic capacity is insufficient without downstream packagingMulti-quarter
Geographic riskConcentrated sites create correlated disruptionPotentially immediate impact

3.8 Advanced Packaging: The Hidden Capacity Multiplier

Advanced packaging transforms fabricated logic and HBM into a usable accelerator. For AI systems, packaging is no longer a low-value final step; it determines how many compute dies and HBM stacks can communicate at sufficient bandwidth and power density. TSMC’s CoWoS platform, Intel’s EMIB and Foveros technologies, Samsung’s advanced packaging, and services supplied by ASE, Amkor and other OSAT firms form a heterogeneous market. These offerings cannot automatically be included in one market because technical substitutability varies by design.

TSMC stated in April 2025 that it was working to double CoWoS capacity during that year. The company also announced plans for a 9.5-reticle-size CoWoS implementation in volume production during 2027, intended to support twelve or more HBM stacks. TSMC First Quarter 2025 Earnings Conference Transcript – Taiwan Semiconductor Manufacturing Company – April 2025 TSMC Unveils Next-Generation A14 Process at North America Technology Symposium – Taiwan Semiconductor Manufacturing Company – April 2025

No primary, audited capacity-share denominator permits a defensible packaging HHI. The appropriate concentration measure must use qualified AI-package output rather than total packaging revenue.

Table 3.6 — Packaging Capacity Measurement

VariableCorrect unitWhy revenue share is insufficient
Interposer capacityQualified interposer area per monthPackage sizes differ
AI packages completedQualified packages per quarterComplexity and HBM count vary
HBM stacks integratedStacks per package and total stacksStack intensity changes capacity use
YieldGood packages divided by startsNominal capacity overstates output
Lead timeOrder-to-delivery monthsReveals effective scarcity
Customer allocationShare committed under contractShows residual open capacity
Thermal capabilitySupported power densityDetermines substitutability
Reticle-equivalent areaReticle units per packageNormalises package scale

3.9 HBM and Conventional Memory

HBM is a stronger potential bottleneck than conventional DRAM because AI accelerators require high bandwidth, dense stacking and close integration with advanced packaging. Principal HBM suppliers include SK hynix, Samsung Electronics and Micron. Conventional DRAM is also concentrated around these large producers, although exact shares differ by product generation and period. NAND includes a broader set of firms.

Micron disclosed that its HBM business was supported by multi-year strategic customer agreements and forecast the HBM total addressable market to increase from approximately $35 billion in 2025 to around $100 billion in 2028. Micron Fiscal First Quarter 2026 Financial Results – Micron Technology – December 2025 Samsung announced plans to invest more than KRW 110 trillion in facilities and R&D during 2026, including HBM, foundry operations and advanced packaging. Samsung Electronics Shareholder Value Enhancement Plan – Samsung Electronics – March 2026

Official corporate disclosures do not provide a common audited global share denominator for HBM. Therefore, CR3 and HHI remain ND. If exactly three suppliers constituted the complete relevant market, the mathematical minimum HHI would be approximately 3,333; unequal shares would produce a higher figure. This is a structural illustration, not a reported market estimate.

Table 3.7 — HBM Scarcity Mechanisms

MechanismEffect
Wafer-intensity per delivered bitHBM capacity competes with conventional memory capacity
Complex stackingYield losses can compound across layers
Customer qualificationAccelerator suppliers cannot instantly change memory
Packaging co-designMemory must be integrated with logic and interposer design
Long-term agreementsFuture supply can be allocated before spot availability
Generation transitionHBM3E and HBM4 cannot be treated as identical output
Capital disciplineMemory producers may avoid uncontrolled commodity overcapacity
Export controlsEligible demand differs by jurisdiction
Thermal and power constraintsNominal capacity may not be deployable in all systems

The U.S. Bureau of Industry and Security added controls on HBM in December 2024 alongside new controls on semiconductor-manufacturing equipment and software. This policy confirms that HBM is treated as a strategic input, but it does not quantify commercial concentration. Commerce Strengthens Export Controls to Restrict China’s Capability to Produce Advanced Semiconductors for Military Applications – U.S. Bureau of Industry and Security – December 2024


3.10 Substrates, Interposers and Secondary Bottlenecks

Package substrates and interposers can constrain AI capacity even though their revenue is small relative to accelerators. They require high layer counts, fine interconnect geometry, low defect rates and qualification for large, high-power packages. A substrate shortage can immobilise completed logic and memory, creating a disproportionate output loss. This is the “low-value/high-criticality paradox”: a component can represent a small share of system cost while controlling the availability of the entire system.

Criticality-adjusted value is:

CAVj = RevenueSharej × OutputLossElasticityj

A low-revenue input with a very large output-loss elasticity can rank above a more expensive but substitutable component.

No audited global market shares for AI-qualified substrates or silicon interposers were identified. CR4 and HHI are therefore ND. Chapter monitoring should instead use:

  • substrate lead time;
  • qualified supplier count;
  • capacity-expansion announcements;
  • yield;
  • customer concentration;
  • package-area requirements;
  • geographic concentration;
  • inventory coverage.

3.11 Accelerator Design and CUDA Ecosystem Dependence

Accelerator competition occurs across discrete GPUs, integrated accelerators, custom ASICs and internally developed cloud chips. Principal suppliers include NVIDIA, AMD and Intel, while Google, Amazon and Microsoft have developed custom accelerators for internal cloud deployment. Chinese and other regional firms pursue domestic alternatives under different software and export-control conditions.

NVIDIA’s competitive position cannot be reduced to silicon. Its system includes GPU architecture, CUDA, libraries, compilers, communications software, NVLink, InfiniBand and Ethernet networking, reference systems and developer support. The economic switching cost includes code porting, kernel optimisation, model validation, workforce retraining, operational risk and performance uncertainty. A customer can theoretically purchase another accelerator but still face a large total migration cost.

The CUDA Lock-In Index proposed for later estimation is:

CLIi = w₁CodePortingi + w₂PerformanceLossi + w₃Retrainingi + w₄LibraryGapi + w₅OperationalRiski

Table 3.8 — Accelerator Competition Layers

LayerNVIDIA positionAlternative mechanismLock-in test
HardwareGPU and integrated systemsAMD, Intel, custom ASICs and regional acceleratorsConstant-workload cost
Programming modelCUDAROCm, SYCL, Open standards and proprietary cloud stacksPorting time
LibrariesCUDA-X ecosystemAlternative libraries and framework abstractionsFeature completeness
Scale-up interconnectNVLink/NVSwitchAlternative proprietary or open interconnectsCluster efficiency
Scale-out networkingInfiniBand and Ethernet portfolioBroadcom, Marvell, Cisco, Arista and othersEnd-to-end performance
SystemsDGX, HGX and rack-scale platformsOEM and custom systemsDeployment time
Cloud distributionInstances across hyperscalersCustom cloud chips and rival GPUsAvailability and tariff
Developer baseLarge installed skill baseCross-platform software and compiler abstractionLabour-market switching cost

No official authority or audited corporate filing supplies a complete 2025 accelerator market denominator that simultaneously covers merchant GPUs and captive custom chips. A formal global accelerator HHI is therefore ND. Treating custom TPU or Trainium capacity as zero because it is not sold as a discrete card would overstate merchant-market concentration while understating user dependence on individual clouds.


3.12 AI-Server Manufacturing

AI-server manufacturing includes branded OEMs, original-design manufacturers and cloud operators that design systems internally. Supermicro, Dell, HPE, Lenovo and Inspur coexist with manufacturing partners such as Quanta, Wiwynn and Foxconn. Their bargaining power depends less on processor intellectual property than on allocation, integration, liquid cooling, rack power, delivery capability and customer certification.

AI-server concentration cannot be calculated reliably from total server revenue because conventional enterprise servers are not substitutes for dense accelerator systems. A valid market denominator must include only systems meeting defined accelerator, networking and cooling criteria.

Table 3.9 — AI-Server Entry Barriers

BarrierSeverityExplanation
Accelerator allocationVery highIntegrators cannot ship systems without scarce GPUs
Thermal engineeringHighRack-scale AI systems require advanced cooling
Power densityHighFacility compatibility limits deployment
Networking integrationHighPerformance depends on topology and configuration
Firmware and validationHighEnterprise customers require tested configurations
Working capitalHighExpensive components create financing exposure
Customer supportHighDowntime costs require global servicing
Manufacturing scaleMedium-highVolume and supplier relationships affect delivery
Brand trustMedium-highCritical buyers favour established support providers

3.13 Cloud Infrastructure: A Verifiable Concentration Case

Cloud infrastructure is one segment where an official authority has published usable share ranges. The UK Competition and Markets Authority’s 2024 revenue analysis placed AWS and Microsoft individually in the 30–40% range for combined IaaS and PaaS, Google in the 5–10% range, IBM and Oracle each in the 0–5% range, and other providers collectively in the 20–30% range. The CMA stated that AWS and Microsoft’s combined share remained within 60–80% in 2024. Appendix D: Market Structure, Concentration Methodology and UK Share of Supply by Revenue – UK Competition and Markets Authority – 2025

Because exact shares are redacted into ranges and “other” aggregates many firms, a single exact HHI would be false precision. A strict concentration floor can nevertheless be calculated.

If AWS and Microsoft have a combined 60% and are divided equally, and Google has 5%, then:

HHI floor = 302 + 302 + 52 = 1,825

The remaining 35% necessarily adds a positive amount. Therefore actual HHI is greater than 1,825.

If AWS and Microsoft have a combined 80%, divided equally, and Google has 10%, then:

Partial HHI = 402 + 402 + 102 = 3,300

Again, remaining firms add to the value. The result establishes that the UK IaaS/PaaS market is highly concentrated under the U.S. 1,800 threshold even at the lowest mathematically favourable endpoint of the CMA ranges.

Table 3.10 — UK Cloud Concentration

MeasureOfficial range or derived valueInterpretation
AWS share, 202430–40%One of two market leaders
Microsoft share, 202430–40%One of two market leaders
Google share, 20245–10%Distant third
IBM share, 20240–5%Small
Oracle share, 20240–5%Small
Other combined20–30%Aggregate, not one firm
CR260–80%Very high two-firm concentration
CR365–90% before range reconciliationHigh; exact value redacted
HHI lower boundGreater than 1,825Highly concentrated
Plausible partial HHI at upper endpointsAt least 3,300Very highly concentrated

The CMA subsequently stated that its investigation found AWS and Microsoft to possess positions of significant market power and identified data-egress charges, interoperability barriers and Microsoft software-licensing practices as limitations on customer choice. CMA Announces Package of Actions on Business Software and Cloud Services – UK Competition and Markets Authority – March 2026


3.14 Frontier-Model Development

Frontier models require capital, compute, data, specialised talent, evaluation systems, distribution and sustained inference capacity. Principal developers include OpenAI, Google DeepMind, Anthropic, Meta and xAI, alongside major Chinese and other national developers. Market definition remains unresolved. Possible markets include model training, API access, consumer assistants, enterprise models, open-weight models, multimodal systems and task-specific inference. User counts, API revenue, token volume and benchmark capability would produce different shares.

No primary regulator or audited filing provides a complete frontier-model revenue or quality-adjusted usage denominator. Therefore:

  • CR3: ND;
  • CR4: ND;
  • HHI: ND.

The absence of a calculable HHI does not imply low concentration. It means the relevant market and denominator remain empirically incomplete.

Table 3.11 — Frontier-Model Entry Barriers

BarrierMechanism
Training computeLarge upfront cluster access
Inference capacityContinuing expenditure after model launch
DataQuality, legality, cleaning and domain coverage
TalentSmall global pool of experienced researchers and infrastructure engineers
EvaluationSafety, reliability and benchmark infrastructure
DistributionConsumer platform, cloud or enterprise installed base
Brand and trustBuyers prefer recognised models for critical deployment
RegulationCompliance fixed costs favour scale
Feedback loopsUser interactions can improve products and distribution
CapitalLong investment horizon and uncertain monetisation

3.15 Vertical Integration Between Clouds and Model Developers

Vertical and quasi-vertical relationships join infrastructure providers with model developers through equity investment, cloud commitments, revenue sharing, distribution, technical cooperation and preferential access. The FTC’s January 2025 study examined partnerships involving Microsoft and OpenAI, Amazon and Anthropic, and Alphabet and Anthropic. It found that cloud partners received various equity, revenue-sharing, consultation, control or exclusivity-related rights and that model developers obtained capital, cloud resources and distribution. Partnerships Between Cloud Service Providers and AI Developers – Federal Trade Commission – January 2025

These partnerships can create efficiencies:

  • finance otherwise unaffordable model development;
  • secure infrastructure supply;
  • integrate models into enterprise products;
  • reduce deployment latency;
  • distribute model services to a large customer base;
  • align technical road maps.

They can also create competitive risks:

  • make model developers dependent on one cloud;
  • reserve scarce compute;
  • provide cloud partners access to sensitive technical information;
  • limit multi-cloud deployment;
  • favour affiliated models in distribution;
  • increase switching cost;
  • create circular revenue and procurement relationships.

Table 3.12 — Vertical Integration Test

QuestionEfficiency interpretationCompetition-risk interpretation
Is cloud capacity guaranteed?Enables investment and stable deploymentRemoves scarce capacity from rivals
Is model distribution bundled?Reduces customer acquisition costFavours affiliated model
Are commitments exclusive?Supports relationship-specific investmentPrevents multi-cloud competition
Is sensitive information shared?Improves engineering coordinationWeakens future rivalry
Are revenues shared?Aligns incentivesReinforces circular dependence
Can the developer switch cloud?Relationship remains contestableLock-in becomes structural
Can customers export workloads?Interoperability preservedData and workflow lock-in
Are rival models offered equally?Platform remains openSelf-preferencing

3.16 Long-Term Supply Agreements and Foundry Allocation

Long-term agreements can stabilise investment in fabrication, HBM, packaging and cloud infrastructure. Capital-intensive suppliers may not expand without credible future demand; purchasers may not invest in model programmes without capacity assurance. Agreements are therefore economically productive when they share risk and finance expansion.

The competitive concern arises when a small number of purchasers reserve a dominant share of output, leaving independent firms to buy from a thin residual market. The appropriate metrics are:

Customer Allocation Ratio:

CARk,j = Capacityallocated to customer k,j ÷ Capacitytotal,j

Residual Market Ratio:

RMRj = 1 − Σk=1KCARk,j

Allocation HHI:

AHHIj = Σk=1K(100CARk,j)2

NVIDIA disclosed that in fiscal 2026 one direct customer represented 22% of total revenue and another represented 14%, primarily in Compute & Networking. This is customer concentration, not supplier market share. It shows substantial buyer concentration within NVIDIA’s revenue base and complicates a simple seller-power narrative. NVIDIA Annual Report on Form 10-K for the Fiscal Year Ended January 25, 2026 – NVIDIA and U.S. Securities and Exchange Commission – February 2026


3.17 Export Controls as an Architecture of Geographic Scarcity

Export controls divide global physical capacity from legally accessible capacity. The U.S. Bureau of Industry and Security controls designated advanced computing semiconductors, semiconductor-manufacturing equipment, HBM and related software according to performance, end user and destination. In December 2024, BIS added controls on 24 types of semiconductor-manufacturing equipment, three types of software and specified HBM. Commerce Strengthens Export Controls to Restrict China’s Capability to Produce Advanced Semiconductors for Military Applications – U.S. Bureau of Industry and Security – December 2024

In January 2026, BIS revised its licensing approach for NVIDIA H200, AMD MI325X and similar products destined for approved Chinese customers, moving to case-by-case review subject to stated conditions. Department of Commerce Revises License Review Policy for Semiconductors Exported to China – U.S. Bureau of Industry and Security – January 2026

The geopolitical availability ratio is:

GARr,t = LegallyAccessibleFrontierCapacityr,t ÷ GlobalFrontierCapacityt

A country can face severe effective scarcity even while global production rises if GAR declines.

Table 3.13 — Export-Control Effects

EffectIntended functionEconomic consequence
Restrict advanced accelerator exportsNational-security controlRegional performance ceiling
Restrict manufacturing equipmentLimit indigenous advanced productionProlonged fabrication dependence
Control HBMRestrict complete AI-system performanceMemory-based capacity constraint
Control design/manufacturing softwareLimit advanced-node productivityEDA dependence
Entity-list restrictionsTarget specific organisationsDue-diligence and transaction costs
Anti-diversion rulesPrevent circumventionWider compliance burden
Licensing uncertaintyPreserve government discretionInvestment and allocation uncertainty
Domestic substitutionEncourage alternative ecosystemsDuplication cost and fragmentation

3.18 Subsidies and Industrial Policy

Governments increasingly treat semiconductor and AI capacity as strategic infrastructure. Subsidies can reduce entry barriers, accelerate domestic fabrication, diversify geography and internalise resilience benefits not rewarded by short-term commercial returns. They can also create overcapacity, political allocation, duplicated infrastructure or subsidy competition among jurisdictions.

The economic test is:

NetSocialValuep = ResilienceBenefitp + InnovationSpilloverp + SecurityBenefitp + TaxEmploymentBenefitp − FiscalCostp − DistortionCostp − DuplicationCostp

Table 3.14 — Industrial-Policy Evaluation

PolicyPotential benefitPrincipal riskRequired KPI
Fab subsidyGeographic diversificationSubsidy competitionQualified output per public dollar
R&D creditTechnology developmentSubsidising activity that would occur anywayIncremental patent and process output
Public computeResearch and SME accessLow utilisation or political allocationSuccessful workloads and user diversity
AI GigafactoryFrontier-scale capacityReinforcing large incumbentsOpen-access share and cost per task
Workforce programmeReduces skill bottleneckTraining mismatchQualified workers retained
Energy investmentEnables data-centre expansionCost socialisationFirm capacity and grid reliability
Procurement preferenceCreates demand for domestic suppliersHigher cost and lock-inPerformance-adjusted procurement cost
Open standardsReduces switching costsSlow consensusMigration cost and interoperability

3.19 Firm-Level Profitability and Economic Rent

High profitability is evidence of value capture, not proof of unlawful market power. This chapter calculates reported margin indicators for selected critical firms using official filings.

Operating margin is:

OMi,t = OperatingIncomei,t ÷ Revenuei,t × 100

Net margin is:

NMi,t = NetIncomei,t ÷ Revenuei,t × 100

Return on invested capital is:

ROICi,t = NOPATi,t ÷ AverageInvestedCapitali,t

Economic profit is:

EPi,t = [ROICi,t − WACCi,t] × InvestedCapitali,t

Free cash flow is:

FCFi,t = OperatingCashFlowi,t − CapitalExpenditurei,t

Market value added is:

MVAi,t = EnterpriseMarketValuei,t − InvestedCapitali,t

Tobin’s q is approximated as:

qi,t = MarketValueAssetsi,t ÷ ReplacementCostAssetsi,t

A credible Tobin’s q cannot be calculated from market capitalisation alone because the replacement cost of intangible and physical assets is not directly observable.

Table 3.15 — Verified Profitability Indicators

Firm and periodRevenueGross marginOperating income or marginNet income or marginEvidentiary interpretation
NVIDIA FY2026$215.938bn71.1%$130.387bn; ≈60.4%$120.067bn; ≈55.6%Exceptional value capture across accelerated-computing platform
TSMC 2025≈$122bn59.9%50.8% margin45.1% marginStrong utilisation, node leadership and scale
ASML 2025€32.7bn52.8%€11.3bn; ≈34.6%€9.6bn; ≈29.4%High-return equipment and service ecosystem
TSMC 2024Reported annual revenue56.1%45.7%40.5%Lower than 2025, supporting utilisation and mix effects
TSMC 2023$69.30bn54.4%42.6%38.8%Semiconductor downturn demonstrates cyclicality

NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026 – NVIDIA – February 2026 TSMC 2025 Annual Report – Taiwan Semiconductor Manufacturing Company – April 2026 ASML 2025 Annual Report – ASML – January 2026 TSMC 2023 Annual Report – Taiwan Semiconductor Manufacturing Company – February 2024

Table 3.16 — Profit Explanation Matrix

Observed profit sourceNVIDIATSMCASML
Innovation rentVery highVery highVery high
Intellectual-property rentVery highHigh through process know-howVery high
Scarcity rentHigh during constrained supplyHigh at leading nodesHigh for critical tools
Scale economyVery highVery highHigh
Ecosystem rentVery highMedium-highHigh through installed base
Risk compensationHighVery high due capital cycleVery high due R&D cycle
Market-power componentPlausible and requiring formal testPlausible in leading-edge segmentsStructurally plausible
Evidence of unlawful conductNot established by marginsNot established by marginsNot established by margins

NVIDIA’s fiscal 2026 operating margin of approximately 60.4% is extraordinarily high, but the gross margin declined from 75.0% in FY2025 to 71.1% in FY2026, partly reflecting the transition to full-scale Blackwell data-centre solutions and a charge associated with H20 inventory and purchase obligations. This movement is inconsistent with a simplistic assumption that market power mechanically forces margins upward every year. It supports a mixed explanation involving innovation, system mix, scarcity and platform power.


3.20 Formal Bottleneck Analysis

A bottleneck must be evaluated by more than supplier count. This chapter defines the Global AI Bottleneck Score:

GABSj = 0.25Cj + 0.20Sj + 0.20Tj + 0.15Gj + 0.10Ij + 0.10Aj

Where each variable is scored from 0 to 100:

  • Cj: concentration;
  • Sj: lack of substitutability;
  • Tj: replacement or expansion time;
  • Gj: geographic concentration;
  • Ij: impact on downstream output;
  • Aj: degree of pre-allocation.

The scores below are structured analytical assessments, not audited statistics.

Table 3.17 — Global AI Bottleneck Ranking

RankComponentConcentrationSubstitutabilityExpansion timeDisruption impactPreliminary risk
1EUV lithographyExtremeExtremely lowExtremely longSystemicCritical
2Leading-edge foundry capacityVery highVery lowVery longSystemicCritical
3HBMVery highLowLongSevereCritical
4Advanced packagingHighLowLongSevereCritical
5Accelerator software ecosystemVery highLow-mediumLong organisationallySevereCritical
6EDA and process-design integrationVery highLowLongSystemic for new designsCritical
7AI data-centre electricityRegionally highLow locallyLongSevere by regionHigh
8High-speed networking and opticsHigh in specialised productsMedium-lowMedium-longSevere at cluster scaleHigh
9Advanced substrates/interposersHigh in qualified supplyLow short termMedium-longHighHigh
10AI-server liquid coolingModerateMediumMediumHigh locallyMedium-high
11Frontier-model talentHighly concentratedLow short termLongHigh for innovationHigh
12Conventional DRAM/NANDConcentrated but cyclicalMediumMediumModerateMedium

3.21 Disruption Transmission

The output impact of a component disruption is:

ΔCapacityAI = Elasticityj × ΔSupplyj

For strict complements, elasticity can approach one over the binding range: a 10% reduction in the bottleneck may reduce deliverable systems by close to 10%. For inputs with inventories or substitutes, short-run elasticity will be lower.

Table 3.18 — Bottleneck Failure Scenarios

Disrupted nodeImmediate effectSecond-order effectRecovery condition
EUV equipment supplyDelayed new-node capacity and tool replacementReduced future accelerator outputTool delivery, service and process recovery
Leading-edge foundryLoss of logic-die productionGPU, CPU and custom-accelerator shortagesAlternative qualification or restored fab
HBMCompleted logic cannot be configured at target memoryLower accelerator shipments or reduced specificationsQualified alternative memory
Advanced packagingLogic and memory inventories accumulate unfinishedServer delays and cloud-capacity shortfallPackaging expansion and yield restoration
EDA/sign-off toolsNew designs delayedSlower architecture transitionTool restoration or revalidation
Power connectionInstalled systems cannot operate at intended scaleLower utilisation and regional queueingGrid and generation capacity
NetworkingAccelerators operate below cluster potentialHigher cost per training or inference taskAlternative fabric and retuning
Cloud platformService interruption and workload migrationEnterprise operational disruptionMulti-cloud portability and restored service
Model APIApplication failure despite available hardwareBusiness-process interruptionModel substitution and validation
Export licenceGeographic access terminated or delayedRegional price divergence and substitutionLicence, redesign or domestic alternative

3.22 Systemic Concentration: The Common-Dependency Problem

Downstream diversity can conceal upstream commonality. Multiple AI-model providers may use different models while relying on the same cloud, accelerator architecture, foundry, lithography equipment and HBM suppliers. The probability of correlated failure is therefore higher than a simple count of model providers suggests.

The Common Dependency Ratio is:

CDRj = DownstreamCapacityDependentOnNodej ÷ TotalDownstreamCapacity

The systemic loss expectation is:

E(Lossj) = ProbabilityDisruptionj × CDRj × EconomicValueAtRisk

Table 3.19 — Common-Dependency Layers

Downstream appearanceHidden common dependency
Several cloud providersSame leading-edge foundry and equipment chain
Several model providersSame cloud and accelerator infrastructure
Several accelerator brandsSame packaging, substrates or memory suppliers
Several server brandsSame accelerator allocation and networking
Several enterprise AI productsSame frontier-model API
Several national AI programmesSame imported accelerators and software stack
Several local modelsSame consumer GPU ecosystem and memory constraints

3.23 Final Concentration Register

Table 3.20 — CRn and HHI Status by Segment

SegmentCRnHHIDetermination
Commercial EUV lithographyConditional CR1 = 100%Conditional HHI = 10,000Technical single-source condition; formal denominator not independently published
EDANDNDHigh structural concentration; exact primary shares unavailable
Leading-edge foundryNDNDVery high concentration expected; no harmonised primary denominator
HBMNDNDThree principal suppliers; exact shares unavailable
Conventional DRAMNDNDHighly concentrated structure; exact current shares unavailable
NANDNDNDMore diversified than HBM/DRAM; exact shares unavailable
Advanced packagingNDNDProduct heterogeneity prevents simple revenue denominator
AI substrates/interposersNDNDQualified-capacity data unavailable
Merchant AI acceleratorsNDNDCaptive custom chips complicate market definition
AI serversNDNDConventional-server revenue is an invalid denominator
UK IaaS/PaaS, 2024CR2 = 60–80%Strictly greater than 1,825Officially verifiable highly concentrated market
Frontier modelsNDNDMarket definition and usage denominator unresolved
Enterprise AI distributionNDNDMust be separated by application category

3.24 Final Judgment

The AI supply chain is not uniformly monopolised, but it is structurally concentrated around a small number of high-criticality nodes. The strongest concentration lies upstream in EUV lithography, leading-edge fabrication, EDA integration, HBM and advanced packaging; downstream concentration is most verifiable in public cloud infrastructure. The UK cloud case demonstrates that exact market shares are not necessary to establish a minimum structural conclusion: official ranges imply an HHI above the 1,800 high-concentration threshold even under the most concentration-minimising assumptions.

The architecture of scarcity is cumulative. Accelerator designers depend on foundries; foundries depend on lithography and EDA; accelerators depend on HBM and packaging; servers depend on networking and cooling; clouds depend on energy and server allocation; model developers depend on clouds; enterprise applications depend on models and installed software ecosystems. Scarcity at one layer can therefore transmit across the entire chain.

Profitability is likewise multi-causal. NVIDIA, TSMC and ASML report exceptional margins, but the evidence supports a combination of innovation rents, intellectual-property returns, scarcity rents, scale economies, risk compensation and potential market power. High margins alone do not establish unlawful conduct. A legal or competition finding would require defined markets, exclusionary conduct, counterfactual analysis and evidence that efficiencies do not explain the observed outcome.

The greatest five-year systemic vulnerabilities are EUV lithography, leading-edge foundry capacity, HBM, advanced packaging, EDA integration and accelerator software lock-in. Electricity and networking form the next-order constraints: even if semiconductor output expands, insufficient grid capacity or interconnection can prevent physical hardware from becoming usable compute. The central strategic conclusion is therefore that global AI capacity will not be determined by the number of chips designed, but by the smallest deployable quantity among all indispensable complements.

CHAPTER 4 — The Real Cost of Local AI

4.1 Scope, definitions and evidentiary boundary

The economic cost of local artificial intelligence cannot be inferred from the retail price of a computer, the nominal parameter count of a model, or the fact that a compressed checkpoint can technically be loaded into system memory. A local deployment becomes economically meaningful only when it can complete a specified workload at an acceptable level of quality, latency, throughput, concurrency, reliability, security and operational continuity. This chapter therefore defines usable local AI as a complete socio-technical system—not merely a model file—capable of satisfying explicit service-level and quality thresholds for its intended users. Its resource envelope includes accelerator memory, system memory, storage, interconnect bandwidth, electrical supply, cooling, networking, inference software, model licences, security controls, specialist labour, maintenance, replacement capacity, evaluation procedures and incident response. The analysis separates three evidentiary categories. First, hardware capacities and power limits attributed to named products are taken from live official manufacturer documentation. Second, memory, energy and cost relationships are derived from reproducible mathematical identities. Third, prices not publicly posted by manufacturers, future operating conditions and organisation-specific expenses are presented as transparent planning scenarios, not as observed market facts. This distinction is indispensable because enterprise accelerators, integrated servers, support contracts and secure infrastructure are frequently sold by quotation, while electricity tariffs, taxes, salaries, utilisation rates and compliance costs differ dramatically across jurisdictions. A false appearance of precision would be less scientific than a documented range accompanied by sensitivity analysis.

The physical constraints are already sufficient to disprove the proposition that every “large-memory computer” is interchangeable. An NVIDIA GeForce RTX 5090 provides 32 GB of GDDR7 memory, 1,792 GB/s of memory bandwidth and an official starting price of USD 1,999, whereas an NVIDIA H200 provides 141 GB of HBM3e and 4.8 TB/s of bandwidth. An eight-GPU NVIDIA DGX B200 integrates 1,440 GB of GPU memory, 64 TB/s of aggregate memory bandwidth and a maximum system power rating of approximately 14.3 kW. These are different economic objects: a consumer accelerator, a data-centre accelerator and a rack-scale integrated system. GeForce RTX 5090 Specifications – NVIDIA – January/2025NVIDIA GeForce RTX 5090 official specifications. NVIDIA H200 Tensor Core GPU – NVIDIA – November/2023NVIDIA H200 official product specification. NVIDIA DGX B200 – NVIDIA – 2026NVIDIA DGX B200 official specification. The memory capacity of a system therefore establishes only a feasibility boundary. It does not establish achievable tokens per second, time to first token, multi-user throughput, model quality, operational resilience or total cost.


4.2 The computational anatomy of a local LLM

4.2.1 Model-weight memory

For a dense model containing Np parameters stored at b bits per parameter, the raw weight-memory requirement is:

Mw = Np × b ÷ 8

where:

  • Mw is raw model-weight memory in bytes;
  • Np is the parameter count;
  • b is the effective number of stored bits per parameter.

When Np is expressed in billions of parameters, the resulting number is conveniently expressed in decimal gigabytes:

Mw,GB = Np,B × b ÷ 8

The following table reports theoretical weight-only memory. It excludes quantisation scales, zero points, tensor metadata, KV cache, activations, execution workspaces, multimodal encoders, operating-system use and fragmentation.

Dense-model sizeFP32FP16/BF16FP8/INT8INT43-bit theoretical2-bit theoretical
7B28 GB14 GB7 GB3.5 GB2.63 GB1.75 GB
8B32 GB16 GB8 GB4 GB3 GB2 GB
13B52 GB26 GB13 GB6.5 GB4.88 GB3.25 GB
14B56 GB28 GB14 GB7 GB5.25 GB3.5 GB
30B120 GB60 GB30 GB15 GB11.25 GB7.5 GB
32B128 GB64 GB32 GB16 GB12 GB8 GB
35B140 GB70 GB35 GB17.5 GB13.13 GB8.75 GB
70B280 GB140 GB70 GB35 GB26.25 GB17.5 GB
72B288 GB144 GB72 GB36 GB27 GB18 GB
100B400 GB200 GB100 GB50 GB37.5 GB25 GB
175B700 GB350 GB175 GB87.5 GB65.63 GB43.75 GB
405B1,620 GB810 GB405 GB202.5 GB151.88 GB101.25 GB

The table uses decimal gigabytes because hardware vendors commonly advertise capacity in decimal units. Conversion to binary gibibytes is:

MGiB = MGB ÷ 1.073741824

Consequently, 35 decimal GB of weights occupy approximately 32.60 GiB. This distinction matters near a device’s capacity limit, but it does not create usable headroom: a nominally 35 GB INT4 checkpoint cannot safely be treated as a complete 35 GB runtime.

4.2.2 Precision and quantisation regimes

RepresentationNominal bits per parameterRaw bytes per parameterPrincipal advantagePrincipal limitation
FP32324.00High numerical range and reference-grade computationExcessive memory and bandwidth for ordinary inference
FP16162.00Broad accelerator support and high throughputSmaller numerical range than BF16
BF16162.00FP32-like exponent rangeRequires compatible hardware and kernels
FP881.00High accelerator throughput; major memory reductionHardware, calibration and kernel dependence
INT881.00Mature compression path with moderate quality riskSpeedup depends on native integer execution
INT440.50Makes 30–100B-class local inference materially more feasibleQuality and speed are method-, model- and kernel-dependent
3-bit30.375Further capacity reductionMetadata overhead becomes proportionally larger; quality risk rises
2-bit20.25Extreme compressionMaterial model-specific degradation and limited kernel portability
Mixed precisionVariableVariablePreserves sensitive layers at higher precisionEffective bits per parameter must be measured, not assumed

Nominal bit width does not equal effective storage cost. Group-wise quantisation stores scales and, depending on the format, zero points or codebooks. Embedding and output layers may remain at a higher precision. Alignment padding, tensor headers and duplicated buffers add further overhead. A defensible capacity calculation should therefore use:

beff = 8 × Scheckpoint ÷ Np

where Scheckpoint is the measured checkpoint size in bytes. Runtime memory must then be measured after model initialisation because some frameworks dequantise selected tensors, construct auxiliary lookup tables or reserve execution workspaces.

Quantisation is not an unconditional economic gain. It reduces memory traffic and may increase speed when the hardware and inference engine possess efficient kernels for the selected format. Conversely, a heavily compressed model can run more slowly when its kernels repeatedly dequantise weights into a higher-precision representation, when unsupported operations fall back to the CPU, or when heterogeneous offloading saturates the PCIe or system-memory interface. Quality loss is also workload-dependent: a compression method that preserves average language-model perplexity may still degrade code generation, multilingual reasoning, exact extraction, numerical work, tool use or rare-domain terminology. The correct comparison is therefore not “INT4 versus FP16” in the abstract, but the lowest-cost representation that passes a predefined task-specific evaluation suite.


4.3 Total operational memory

The correct operational identity is:

Mtotal = Mw + MKV + Mact + Mworkspace + Mmultimodal + MOS + Mfrag

This expands the requested expression by separating execution workspaces, multimodal components, operating-system reservation and fragmentation. The components cannot be represented by a universal fixed percentage because they respond differently to context length, batch size, model architecture, framework and concurrency.

4.3.1 KV-cache memory

For a decoder-only transformer, an architecture-aware approximation is:

MKV = 2 × L × HKV × Dhead × T × C × bKV ÷ 8

where:

  • 2 represents the separately stored keys and values;
  • L is the number of transformer layers;
  • HKV is the number of key-value heads;
  • Dhead is the head dimension;
  • T is the number of cached tokens per sequence;
  • C is the number of simultaneously resident sequences;
  • bKV is the cache precision in bits.

If requests have unequal context lengths, the more general form is:

MKV = 2 × L × HKV × Dhead × bKV ÷ 8 × ΣTj

where ΣTj is the sum of cached tokens across all active sequences. This formulation avoids incorrectly multiplying a maximum context length by a nominal user count when many users are idle or have shorter prompts.

Two illustrative calculations show why context and concurrency can dominate memory planning. For a hypothetical 8B-class grouped-query model with 32 layers, eight KV heads, a head dimension of 128, a 32,768-token context, one active sequence and a 16-bit cache:

MKV = 2 × 32 × 8 × 128 × 32,768 × 1 × 16 ÷ 8
MKV = 4 GiB

For a hypothetical 70B-class grouped-query model with 80 layers and the same eight KV heads and 128-dimensional heads:

Context and concurrencyKV-cache requirement
8,192 tokens; one active sequence2.5 GiB
32,768 tokens; one active sequence10 GiB
131,072 tokens; one active sequence40 GiB
32,768 tokens; four active sequences40 GiB
32,768 tokens; eight active sequences80 GiB

These are architectural illustrations, not universal values for every 8B or 70B model. Models with multi-head attention can require substantially more cache; multi-query attention can require less. KV-cache quantisation can reduce capacity consumption, but the resulting quality, latency and kernel compatibility must be tested.

4.3.2 Other operational allocations

AllocationMain determinantsPlanning treatment
Activations, MactBatch size, prompt processing, hidden width, kernels, parallelismMeasure peak prefill and decode separately
Execution workspace, MworkspaceAttention implementation, matrix kernels, graph captureEmpirical measurement is preferable to a fixed percentage
OS and display, MOSOperating system, desktop display, background services, unified memoryReserve capacity before model allocation
Fragmentation, MfragAllocator, tensor sizes, model loading/unloading, long-running serviceStress-test over repeated requests; do not rely on a clean startup
Multimodal, MmultimodalVision/audio encoders, projectors, image tokens, preprocessingAdd encoder weights and modality-dependent cache
Redundancy reserveFailover model, rolling updates, duplicate workersRequired for production continuity but often omitted from hobby calculations

A conservative deployment rule is:

Minstalled ≥ Mpeak,measured × (1 + h)

where h is an engineering headroom factor. A planning range of 0.10–0.25 may be reasonable for early design, but it is an assumption rather than a universal standard. Production acceptance should replace it with measured peak memory under the longest supported prompt, the maximum intended number of active sequences, simultaneous prefill events, tool calls and repeated model reloads.

4.3.3 Practical INT4 memory envelopes

Model classRaw INT4 weightsIndicative single-user operational envelopePrincipal constraint
7–8B3.5–4 GB8–12 GBLong-context KV cache
13–14B6.5–7 GB12–20 GBContext, workspace and limited GPU headroom
30–35B15–17.5 GB24–40 GBDoes not comfortably fit many 24 GB workloads at long context
70–72B35–36 GB48–90 GBKV cache, concurrency and multi-device communication
100B50 GB70–130 GBSustained bandwidth and operational reserve
175B87.5 GB120–240 GBMulti-accelerator topology becomes decisive
405B202.5 GB280–550 GBSharding, interconnect, reliability and energy
Large MoEBased on total parametersOften much larger than active-parameter count suggestsExpert storage and all-to-all communication

The ranges are scenario envelopes, not specifications for individual models. They deliberately widen at larger sizes because model architecture, context, cache representation and parallelism produce increasingly large differences.


4.4 Mixture-of-experts, multimodal and frontier-scale systems

A mixture-of-experts model separates stored parameters from active parameters. If a model contains E experts but routes each token through only k experts, its arithmetic cost can approach that of the active subset while its weight-memory requirement remains related to total expert capacity:

Mw,MoE ≈ Ntotal × beff ÷ 8

Ctoken ≈ Cshared + k × Cexpert + Crouter

The active-parameter count therefore must never be substituted for total parameters when sizing storage or aggregate accelerator memory. Expert parallelism also introduces network communication: tokens must be routed to the accelerators holding the selected experts, creating all-to-all traffic, load imbalance and tail-latency risk. A nominally efficient MoE model can consequently perform poorly on consumer multi-GPU systems without high-bandwidth peer-to-peer links.

Multimodal deployment adds at least five components:

Mmultimodal = Mencoder + Mprojector + Mmodality-cache + Mpreprocess + Mgenerated-media

High-resolution images may produce hundreds or thousands of modality tokens. Video produces a sequence of frames and can multiply both preprocessing cost and context consumption. Speech systems add audio encoders, feature extraction and potentially speech synthesis. Image-generation models require diffusion or autoregressive decoders that can have a memory profile substantially different from text decoding. A computer that runs a text-only 70B model acceptably may therefore fail the equivalent multimodal service objective.

Near-frontier dense systems are physically feasible on sufficiently large integrated infrastructure but economically remote from ordinary users. A 405B dense model requires approximately 810 GB for FP16 weights alone or 202.5 GB at nominal INT4. The eight-GPU DGX B200’s official 1,440 GB aggregate GPU memory can contain 405B FP16 weights with material aggregate headroom, although actual feasibility still depends on the model, cache, kernels and parallelism. Its maximum rated power of approximately 14.3 kW implies that continuous operation would consume, before external cooling losses:

Aannual = 14.3 × 8,760 = 125,268 kWh

At a power-usage effectiveness factor of 1.3, facility energy would be approximately:

Afacility = 125,268 × 1.3 = 162,848 kWh per year

This is a maximum-rating scenario, not a measured annual load. It nevertheless demonstrates why frontier local AI moves from “computer ownership” into electrical, thermal and facility engineering.


4.5 What existing personal hardware can actually do

A system containing 128 GB of system RAM and a 12 GB GPU illustrates the difference between loadability and usability.

Model classTechnical feasibilityLikely execution patternUsability judgment
7–8B INT4/INT8StrongPrimarily GPU-residentSuitable for interactive single-user work if model quality is adequate
13–14B INT4Feasible but constrainedGPU-resident at short/moderate context or minor offloadPotentially usable; context and framework determine headroom
30–35B INT4Weight file fits system RAM but not 12 GB VRAMHeavy CPU execution or CPU/GPU offloadMay be acceptable for patient single-user research; not assumed interactive
70–72B INT4Fits system RAM in weight-only termsPredominantly CPU/system-memory executionLoadable, but sustained latency may fail practical thresholds
100B INT4Raw weights fit 128 GB; full runtime may be marginalHeavy CPU/offloadExperimental rather than dependable
175B INT4Weight-only size may fit, runtime commonly does notInadequate headroomNot operationally defensible
Large multimodal/MoEModel-specificMultiple memory pools and encodersRequires exact architecture and workload testing

No universal tokens-per-second value should be assigned to “a 12 GB GPU” because GPU model, memory bandwidth, CPU, system-memory bandwidth, PCIe generation, offload fraction, inference framework, quantisation kernel, context and sampling parameters all change performance. The correct procedure is to measure representative prompts and report the complete configuration.


4.6 Workload-specific definition of usable local AI

A deployment is usable only if it passes all applicable thresholds in the following vector:

U = f(Q, TPS, TTFT, C, T, A, R, P, G, O)

where:

  • Q = task quality;
  • TPS = output tokens per second;
  • TTFT = time to first token;
  • C = supported concurrency;
  • T = supported effective context;
  • A = availability;
  • R = reliability and recovery;
  • P = privacy and security controls;
  • G = governance and auditability;
  • O = operational update burden.

The following thresholds are proposed as falsifiable report criteria rather than universal regulatory standards.

WorkloadMinimum quality conditionPerformance thresholdContext/concurrency thresholdReliability and governance threshold
Personal assistantPasses user-defined factual and instruction testsMedian ≥15 TPS; p95 TTFT ≤3 s≥8k effective tokens; one userRecoverable service; encrypted storage
Student learningVerified subject tests; citation fidelity measuredMedian ≥12 TPS; p95 TTFT ≤4 s≥16k; one active userClear uncertainty and source checking
Coding assistantRepository-specific test suite; accepted-patch rate measuredMedian ≥20 TPS; p95 TTFT ≤3 s≥16k; one or two usersSandboxed execution; rollback
Long-document analysisExtraction recall and citation-location accuracyMedian ≥8 TPS; p95 TTFT ≤8 s≥32k or validated retrieval pipelineReproducible document ingestion
Independent researchDomain benchmark and provenance testsMedian ≥10 TPS≥32k; one to four usersVersioned model, prompts and corpora
Professional practiceValidated error and abstention thresholds≥12 TPS/user; p95 TTFT ≤4 sFour or more concurrent sessionsAccess control, audit log, backups
Clinical decision supportProspective local validation; human review mandatoryLatency set by clinical workflowCapacity under peak clinical loadHigh availability, traceability, incident response
Enterprise knowledge serviceTask success and grounded-answer rate≥10 TPS/user; p95 TTFT ≤5 s50+ concurrent sessions or measured queue SLASSO, role control, monitoring, failover
Batch analyticsValidated batch accuracyCompletion within processing windowThroughput-basedCheckpointing and job recovery
Government-sensitive useMission-specific validationMission-defined worst-case latencySurge capacity documentedSegmentation, supply-chain review, full auditability

“Fits in memory” satisfies none of these criteria beyond basic physical loadability.


4.7 Capital architectures by user class

4.7.1 Reference deployment tiers

TierIndicative architectureCredible model envelopeAppropriate users
T1 — Entry local12–16 GB GPU; 64–128 GB RAM7–14B interactive; larger models through slower offloadStudent, individual professional
T2 — Advanced workstation24–32 GB GPU; 128–256 GB RAM14–35B strong; 70B offloaded or constrainedResearcher, small practice
T3 — Large unified-memory or dual-accelerator workstation64–192 GB usable accelerator/unified pool35–100B quantised, depending on contextResearch group, startup
T4 — Multi-GPU server192–640 GB aggregate accelerator memory70–405B quantised; smaller models at higher concurrencySME, laboratory, hospital
T5 — Integrated enterprise system1 TB+ aggregate accelerator memoryNear-frontier dense inference and large MoE systemsEnterprise, government
T6 — ClusterMultiple integrated systems with high-speed fabricLarge-scale serving, fine-tuning and resilienceMajor enterprise, state or frontier laboratory

Aggregate memory is not equivalent to a single coherent memory pool. Model weights must be sharded; communication can dominate latency; consumer cards may lack suitable peer-to-peer topology; failure probability rises with device count; and rolling updates can require duplicate capacity.

4.7.2 Planning ranges by institution

All monetary bands below are 2026 USD planning scenarios. They are not vendor quotations. They include broad implementation ranges because procurement geography, taxes, support, security and labour vary greatly.

User or institutionTypical useful scaleIndicative initial CAPEXFive-year cash TCO, excluding internal labourFive-year full economic TCO, including specialist labourMain economic constraint
Student7–14B; sometimes 30B offload$1,500–$5,000$2,500–$9,000$4,000–$20,000Affordability and rapid obsolescence
Independent researcher14–35B; constrained 70B$4,000–$15,000$7,000–$30,000$20,000–$100,000Time spent integrating and evaluating
Small professional practice14–70B$8,000–$35,000$15,000–$80,000$50,000–$300,000Security, support and downtime
Startup35–100B; multi-user$25,000–$200,000$70,000–$600,000$250,000–$2 millionEngineering payroll and utilisation risk
SME70–405B quantised or multiple smaller replicas$75,000–$750,000$250,000–$3 million$1–10 millionReliability, concurrency and integration
University laboratory70B to near-frontier experiments$100,000–$3 million$400,000–$10 million$2–30 millionGrants, facilities and specialist retention
Hospital/privacy-sensitive organisationValidated 35–405B services$250,000–$5 million$1–20 million$5–60 millionGovernance, validation and high availability
Large enterpriseMultiple model classes and business units$1–50 million$5–200 million$20–500 millionPlatform engineering and governance
Government agencySecure departmental system or sovereign cluster$2–100+ million$10–500+ million$50 million–$1 billion+Security accreditation, resilience and sovereign supply

These bands should not be combined into a single mean. The distributions are highly skewed, and institutional requirements create discontinuities. A hospital cannot simply purchase the student configuration multiplied by the number of clinicians; it requires access control, validated workflows, auditability, continuity, backups, change management and security monitoring. The HIPAA Security Rule, for example, requires regulated US entities to implement appropriate administrative, physical and technical safeguards for electronic protected health information. Summary of the HIPAA Security Rule – U.S. Department of Health and Human Services – August/2026HHS Summary of the HIPAA Security Rule. This does not prescribe a particular accelerator, but it changes the economic boundary of an acceptable deployment.


4.8 Five-year total cost of ownership

4.8.1 Correct discounted formulation

The requested TCO model becomes economically consistent when a residual value received in year five is discounted:

TCO5 = CAPEX + Σt=15 [(Et + Ct + Mt + St + Lt + Dt) ÷ (1+r)t] − RV5 ÷ (1+r)5

where:

  • Et = electricity and cooling;
  • Ct = connectivity and cloud overflow/support;
  • Mt = maintenance, spares and warranties;
  • St = software, licences and monitoring;
  • Lt = specialised labour;
  • Dt = downtime, failure and replacement cost;
  • RV5 = year-five resale or residual value;
  • r = real or nominal discount rate consistent with all cash flows.

If all cash flows are expressed in nominal dollars, r must be nominal and component prices should incorporate their expected inflation. If cash flows are constant-price real values, r must be real. Mixing nominal electricity escalation with a real discount rate overstates cost.

For a constant annual operating cost O and constant discount rate r:

PV(O) = O × [1 − (1+r)−5] ÷ r

At r = 7%, the five-year present-value factor is approximately 4.1002.

4.8.2 Electricity and cooling

Annual facility energy is:

KWhfacility,t = PIT,t × 8,760 × ut × PUEt

Et = KWhfacility,t × pe,t

where:

  • PIT,t is average IT power while active, in kW;
  • ut is the duty cycle;
  • PUEt is facility power divided by IT power;
  • pe,t is the electricity price per kWh.

For a personal workstation, a literal data-centre PUE may be inappropriate; cooling may already be embedded in household heating or air-conditioning. In that case, a cooling multiplier should be separately estimated. For a server facility, measured PUE is preferable.

US data centres consumed approximately 176 TWh, or 4.4% of US electricity, in 2023; the official federal assessment projected 325–580 TWh, or 6.7–12% of US electricity, by 2028. This is a system-level projection rather than a tariff forecast, but it identifies energy availability and grid connection as material cost risks. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – U.S. Department of Energy – December/2024U.S. Department of Energy data-centre electricity assessment.

4.8.3 Illustrative workstation TCO

Assume:

  • CAPEX = $5,500;
  • average active IT power = 0.65 kW;
  • utilisation = 30%;
  • cooling multiplier = 1.10;
  • electricity = $0.15/kWh;
  • maintenance and replacement reserve = $300/year;
  • software and backup = $400/year;
  • connectivity/cloud overflow = $200/year;
  • downtime allowance = $250/year;
  • specialist labour excluded from cash TCO;
  • residual value in year five = 10% of CAPEX;
  • discount rate = 7%.

Annual electricity is:

E = 0.65 × 8,760 × 0.30 × 1.10 × 0.15
E ≈ $282

Annual non-labour operating cost is:

O = 282 + 300 + 400 + 200 + 250
O = $1,432

Five-year present value is:

PV(O) = 1,432 × 4.1002
PV(O) ≈ $5,871

Present value of residual value is:

PV(RV5) = 550 ÷ 1.075
PV(RV5) ≈ $392

Therefore:

TCO5,cash = 5,500 + 5,871 − 392
TCO5,cash ≈ $10,979

If the owner spends 80 hours per year on model acquisition, quantisation compatibility, benchmarking, security, backups and troubleshooting, and time is valued at $50/hour, annual labour becomes $4,000. Its five-year present value is approximately $16,401, raising the full economic TCO to approximately $27,380. The example demonstrates that electricity is often not the dominant personal or small-organisation cost. Labour and underutilised capital can exceed it.

4.8.4 Sensitivity matrix

Using the same workstation, five-year discounted electricity cost varies as follows:

Active utilisation$0.10/kWh$0.20/kWh$0.40/kWh
10%$102$204$408
30%$305$610$1,220
60%$610$1,220$2,440
90%$915$1,830$3,660

Assumptions: 0.65 kW active power, 1.10 cooling multiplier, 7% discount rate and constant real electricity prices. The matrix shows why consumer inference economics are often dominated by hardware and labour, while continuously utilised server economics become substantially more sensitive to energy and cooling.


4.9 Cost per million tokens

4.9.1 Energy intensity

If output is generated continuously at Rtok tokens per second:

KWhMtok = PIT × PUE × 1,000,000 ÷ (Rtok × 3,600)

For a 0.5 kW workstation producing 20 output tokens per second with a cooling multiplier of 1.10:

KWhMtok = 0.5 × 1.10 × 1,000,000 ÷ (20 × 3,600)
KWhMtok ≈ 7.64 kWh

At $0.15/kWh:

EnergyCostMtok ≈ $1.15

This excludes prompt-prefill energy, idle power, failed generations, embeddings, retrieval, preprocessing and system administration. It also assumes sustained decoding, which a lightly used personal system rarely achieves.

4.9.2 Levelised cost

A more complete levelised measure is:

LCMtok = TCOperiod ÷ [Tokenssuccessful,period ÷ 1,000,000]

“Successful” tokens should mean tokens belonging to completed workloads that pass quality and usefulness criteria. Counting rejected, hallucinated or unusable output understates economic cost.

For the illustrative workstation TCO of $10,979:

Successful output over five yearsLevelised cash cost
100 million tokens$109.79 per million
500 million tokens$21.96 per million
1 billion tokens$10.98 per million
5 billion tokens$2.20 per million

Including the illustrative labour cost and using the $27,380 full economic TCO:

Successful output over five yearsFull economic cost
100 million tokens$273.80 per million
500 million tokens$54.76 per million
1 billion tokens$27.38 per million
5 billion tokens$5.48 per million

Local AI becomes economically competitive only when utilisation is sufficiently high, the selected local model passes the required quality threshold, and privacy or control benefits are valued. A low-quality local answer is not a cheaper substitute for a successful frontier-model answer if it requires repetition, correction or professional review.


4.10 Reliability, privacy and the hidden duplication requirement

Local deployment can reduce the transmission of prompts and documents to an external inference provider, but it does not automatically provide privacy. The threat surface includes operating-system telemetry, model download channels, malicious model artefacts, compromised dependencies, unencrypted logs, cached prompts, vector databases, administrative accounts, remote-management tools, physical theft, backup media and inference-server APIs exposed without adequate authentication. NIST’s AI Risk Management Framework treats AI risk management as an organisational lifecycle covering governance, mapping, measurement and management rather than as a property delivered by hardware location alone. Artificial Intelligence Risk Management Framework 1.0 – National Institute of Standards and Technology – January/2023NIST AI RMF 1.0.

Production reliability creates a hidden capacity multiplier. If a service must remain available while a model is updated, a second runnable copy may be necessary. If the system cannot tolerate a failed GPU, it needs spare capacity or an external fallback. The availability relation for a repairable component is approximately:

A = MTBF ÷ (MTBF + MTTR)

where MTBF is mean time between failures and MTTR is mean time to repair. End-to-end availability is lower when all serial components must function. For n independent serial components:

Asystem = ΠAi

A system containing accelerators, storage, network switches, cooling and software services can therefore be less reliable than any one component. High-availability design often duplicates the inference endpoint, storage, network path and power supply; consequently, the capital cost of a dependable service can approach 1.5–2.5 times that of the minimum system capable of running the model.


4.11 Benchmark protocol

Every local system evaluated later in the report should disclose:

CategoryMandatory fields
HardwareExact accelerator, count, memory, CPU, RAM channels, storage, interconnect
SoftwareOS, driver, inference framework, kernel/library versions
ModelExact checkpoint, parameter count, architecture, licence, quantisation format
WorkloadPrompt length, output length, context occupancy, sampling settings
PerformancePrefill tokens/s, decode tokens/s, median and p95 TTFT
ConcurrencyActive sequences, queue depth, per-user throughput, p95 latency
MemoryIdle, model-loaded, peak prefill, peak decode, fragmentation after endurance run
EnergyWall-socket IT energy, facility multiplier, idle and active power
QualityDomain benchmark, task-success rate, abstention and factual-error rate
ReliabilityFailure rate, restart time, update downtime, sustained-load duration
EconomicsCAPEX, tax, electricity, support, labour, residual value, utilisation
PrivacyData flow, logs, encryption, access control, telemetry, retention
ReproducibilitySeeds, scripts, test corpus version and measurement date

A valid comparison must report both prefill and decode performance. Long prompts may be limited by prefill compute, while conversational output is often limited by memory bandwidth. Average latency must not replace p95 latency in multi-user services because batching and queueing can conceal severe tail delays.


4.12 Decision matrix by user class

User classMinimum defensible local targetWhen local deployment is rationalWhen it is not rational
Student7–14B quantisedLearning, offline notes, coding support, sensitive draftsWhen frontier quality is essential and utilisation is low
Independent researcher14–35B; optional 70B offloadReproducibility, confidential corpora, sustained experimentsWhen integration time exceeds research value
Small practice14–70B with managed securityRepeated domain workflow with high privacy valueWhen no staff member can own updates and evaluation
StartupReplicated small models or 35–100B serverStable workload, high utilisation, proprietary dataWhen demand is uncertain and capital must remain flexible
SMEMulti-user server plus cloud overflowPredictable volume and integration with internal systemsWhen redundancy and staffing erase unit-cost savings
University labMulti-GPU research serverModel experimentation, controlled datasets, training accessWhen funding covers hardware but not operations
HospitalValidated model service with redundancySensitive data and clinically validated bounded tasksWhen local model quality or governance is inadequate
Large enterprisePortfolio of local, private-cloud and external servicesHigh volume, sovereignty, differentiated workflowsWhen a single procurement is expected to solve heterogeneous needs
Government agencySegmented secure inference infrastructureClassified, sovereign or mission-critical processingWhen supply-chain, staffing or lifecycle support is unresolved

4.13 Five-year outlook, 2027–2031

Over the next five years, raw arithmetic efficiency and quantisation quality are likely to improve, but those improvements will not eliminate local-AI inequality. The binding resource is moving from raw parameter storage toward a composite of memory bandwidth, context capacity, concurrency, multimodal processing, software optimisation, evaluation competence and operational reliability. Smaller models will become increasingly capable for bounded workloads, allowing individuals to obtain useful private assistants without frontier-scale infrastructure. At the same time, the definition of competitive capability will move upward as remote frontier systems gain stronger reasoning, tool use, multimodality and long-context performance. This creates a moving-target effect: the hardware cost of reproducing today’s capability may fall while the cost of matching the current frontier remains high.

Three trajectories should be tested:

ScenarioHardware trajectoryLocal-access effectInstitutional consequence
DiffusionMemory capacity rises; efficient models improve rapidly7–70B capability becomes broadly affordableLocal AI becomes a normal workstation function
BifurcationConsumer capacity improves, but frontier requirements rise fasterBasic access broadens; frontier access remains concentratedTwo-tier capability market persists
ConstrictionMemory, packaging, energy and export constraints remain severeUseful local systems remain expensive outside wealthy marketsSovereignty programmes and regional exclusion intensify

The central inference is that local AI is not one market. It comprises at least four economic regimes: personal offline inference, professional single-node service, institutional multi-user deployment and frontier-scale sovereign infrastructure. The marginal cost of tokens can be low in a highly utilised system, but the fixed cost of entering each successive regime rises discontinuously because memory, redundancy, security and labour requirements arrive in indivisible blocks.


4.14 Testable conclusions

  1. Weight memory is a necessary but insufficient capacity measure. A 70B INT4 dense model has approximately 35 GB of raw weights, but long-context or concurrent operation can add tens of gigabytes of KV cache.
  2. Quantisation changes the feasible model class but not automatically the economically usable model class. Quality, kernel support and offload behaviour must be benchmarked.
  3. A 128 GB RAM system with a 12 GB GPU can load models much larger than its GPU memory, but loadability does not establish interactive usability. The economically defensible envelope is generally smaller than the physical RAM envelope.
  4. For individuals, capital and specialist time commonly dominate electricity. For continuously used institutional systems, energy, cooling, maintenance and redundancy become progressively more important.
  5. The cost per million useful tokens is utilisation-sensitive. Low utilisation can leave local inference more expensive than metered external access even when marginal electricity cost is small.
  6. Privacy-sensitive deployment adds costs rather than simply removing cloud costs. Security, logs, access control, model provenance, backups and incident response remain necessary.
  7. Near-frontier local inference is technically feasible only in the institutional meaning of “local.” It may run inside an organisation’s facility but still require rack-scale power, cooling, networking and specialist operations.
  8. The most scientifically valid definition of usable local AI is workload-specific. It must include quality, p95 latency, throughput, concurrency, effective context, reliability, security and maintenance.
  9. The five-year access divide should be measured against frontier-equivalent task success, not parameter count. Cheap small models may broaden basic access while leaving economically consequential frontier capability concentrated.
  10. Every later cost comparison must report both cash TCO and full economic TCO. Excluding internal labour is appropriate for cash-budget analysis but materially understates social and organisational resource consumption.

CHAPTER 5 — Cloud AI Economics and the Future Price of Tokens

5.1 Scope, analytical objective and evidentiary boundary

Cloud AI converts an exceptionally capital-intensive production system into an apparently simple metered service. The customer sees an API price per million input or output tokens, a monthly subscription, a reserved-capacity contract or a fee for completing a task. Behind that price lies a vertically layered cost structure comprising model research, training, post-training, evaluation, accelerators, high-bandwidth memory, host servers, networking, storage, data-centre construction, land, electricity, cooling, water, maintenance, specialist labour, software engineering, customer acquisition, financing and the economic cost of idle capacity. Token prices therefore cannot be interpreted as if they represented only the electricity consumed during inference. Nor can a low public API price be treated as proof that the corresponding model is inexpensive to produce. Providers can allocate training expenditure across future usage, recover costs through enterprise contracts, bundle AI with higher-margin software, cross-subsidise adoption from advertising or cloud profits, accept a temporary return below the cost of capital, or price strategically to create ecosystem dependence. Conversely, a provider can reduce prices without subsidising usage when hardware efficiency, quantisation, speculative decoding, batching, prompt caching, model routing or utilisation improves faster than capital costs rise. The empirical task is consequently not to assert that all present AI services are loss leaders, but to determine whether disclosed revenues, margins, capital expenditure, depreciation, service tiers and pricing differentials are consistent with competitive introductory pricing, cross-subsidisation or sustainable economic production.

This chapter uses three evidence classes. Observed prices are taken from live official provider documentation as verified on 15 August 2026. Corporate economics are drawn from official investor-relations disclosures, recognising that the major public companies do not publish a complete standalone profit-and-loss account for frontier-model inference. Unit-cost estimates and projections are analytical scenarios generated from explicitly stated parameters. The absence of a separately audited AI-inference segment prevents any scientifically valid claim that a named provider is definitively losing a specified amount per token. The chapter can test whether the loss-leader hypothesis is supported, contradicted or left unresolved; it cannot manufacture missing cost accounts.


5.2 Cloud AI as a production system

The basic economic chain is:

  1. Research and model architecture.
  2. Data acquisition, cleaning, licensing and governance.
  3. Pre-training.
  4. Post-training and alignment.
  5. Safety testing, red teaming and evaluation.
  6. Model optimisation and compilation.
  7. Deployment across inference clusters.
  8. Capacity orchestration, caching and routing.
  9. API, subscription and enterprise distribution.
  10. Customer support, compliance and incident management.
  11. Model replacement and infrastructure refresh.

This sequence creates two distinct but interacting cost pools:

FC = Cresearch + Ctraining + Cpost + Csafety + Csoftware + Cfacility-fixed + Ccommercial-fixed

VC = Caccelerator-time + Cmemory-time + Cenergy + Ccooling + Cnetwork + Cstorage-variable + Csupport-variable + Ctools

The distinction is not perfectly clean. Accelerator depreciation behaves as a committed fixed cost after a cluster is purchased, but it becomes a usage-attributable cost in product accounting. Specialist labour may be fixed within a planning horizon but scalable over longer periods. Data-centre capacity can be owned, financed, leased or reserved through another cloud provider. Model training is sunk after completion, yet continuous retraining and replacement transform it into a recurring portfolio expense. The appropriate classification therefore depends on the decision being studied:

DecisionEconomically relevant cost
Whether to serve one additional request on already-idle hardwareShort-run marginal cost
Whether to discount API prices for one quarterAvoidable operating cost plus strategic objective
Whether to build another inference clusterIncremental full cost and required return
Whether a model family is sustainable over five yearsFully allocated lifecycle cost
Whether an AI company earns economic profitRevenue minus all operating costs and cost of capital
Whether a customer should use cloud or local AICustomer-facing price plus integration, risk and switching costs

A provider can rationally price above short-run marginal cost but below fully allocated average cost while spare capacity exists. Such pricing is not necessarily predatory or irrational; it may maximise contribution margin during demand ramp-up. It becomes structurally unsustainable if revenue remains insufficient to replace depreciating accelerators, finance facilities, develop successor models and compensate invested capital.


5.3 The complete cloud-AI cost stack

5.3.1 Cost-stack architecture

Cost layerPrincipal driverFixed, variable or mixedAppropriate allocation basePrincipal uncertainty
Accelerator purchaseDevice count, memory, interconnectFixed after purchaseAccelerator-hours or successful tokensEconomic life and residual value
Host serverCPU, RAM, storage, NICsFixedServer-hoursShared use across models
Accelerator depreciationPurchase cost, life, utilisationMixed in unit accountingProductive accelerator-secondsObsolescence may precede physical failure
NetworkingFabric, switches, optics, WANMixedBytes, requests, accelerator-hoursMoE and distributed-inference intensity
StorageWeights, datasets, logs, cachesMixedGB-month, read/write operationsReplication and retention policies
Data-centre constructionBuildings, electrical and cooling systemsFixedIT capacity, rack-kW, useful lifeConstruction lead time and stranded capacity
LandLocation and grid proximityFixedSite capacityConnection availability and planning law
ElectricityIT load and tariffVariablekWhRegional tariffs and peak charges
CoolingClimate, PUE, water systemVariable/mixedFacility energy and heat removedWeather and water restrictions
WaterCooling technology and climateVariableLitres per kWh or workloadSite-specific measurement
MaintenanceFailure rates, spares, warrantiesMixedInstalled assetsComponent scarcity
Specialist labourEngineers, researchers, security staffPrimarily fixedProduct, cluster or revenueShared R&D allocation
Pre-trainingCompute, data, researchersSunk per modelLifetime billable usageModel lifetime and training disclosure
Post-trainingFine-tuning, preference data, reinforcement learningRecurrent fixedModel family or usageIteration frequency
Safety and evaluationRed teams, benchmarks, governanceRecurrent fixedModel or releaseScope increases with capability
Inference softwareKernels, compilers, schedulersRecurrent fixedProduct or fleetProprietary efficiency gains
Customer acquisitionSales, credits, advertisingMixedCustomer cohort or revenuePromotional credits and channel costs
Enterprise complianceCertifications, legal, data residencyMixedContract or regionRegulatory fragmentation
FinancingDebt, leases, equity capitalFixed/required returnInvested capitalWeighted average cost of capital
Idle capacityUnused but installed infrastructureEconomic overheadBillable usageDemand volatility
Failed workErrors, retries, rejected outputsVariable wasteSuccessful tasksRarely visible in token accounts

The most frequently omitted item is idle capacity. Cloud inference must retain enough spare capacity to absorb demand variation, equipment failures and latency commitments. A fleet operated at 100% nominal utilisation would exhibit queues and unacceptable tail latency. If only 45% of theoretical accelerator time produces commercially billable tokens, each billed token must absorb more than twice the depreciation that would apply at 100% utilisation.


5.3.2 Accelerator depreciation

Let:

  • Kacc = installed accelerator and server capital;
  • RVn = expected residual value after n years;
  • n = useful economic life;
  • Hyear = 8,760 hours;
  • uproductive = share of time producing billable work.

Straight-line annual depreciation is:

Dannual = (Kacc − RVn) ÷ n

Depreciation per productive accelerator-hour is:

Dhour = Dannual ÷ (8,760 × uproductive)

Consider an analytical cluster costing USD 40 million, with a four-year economic life and a 10% residual value.

Dannual = (40,000,000 − 4,000,000) ÷ 4
Dannual = USD 9,000,000

Productive utilisationProductive hours per yearDepreciation per productive cluster-hour
25%2,190USD 4,110
40%3,504USD 2,568
60%5,256USD 1,712
75%6,570USD 1,370
90%7,884USD 1,142

Utilisation is therefore a first-order economic variable. A 20% improvement in raw accelerator speed is economically inferior to doubling productive utilisation if the faster system remains underused. Providers with large, diverse demand pools possess an important scale advantage because they can route traffic across customers, regions, models, interactive requests and batch jobs.

Microsoft reported quarterly capital expenditure of USD 31.9 billion in fiscal Q3 2026, with approximately two-thirds allocated to shorter-lived assets, principally GPUs and CPUs. The company simultaneously reported that continuing AI-infrastructure investment and growing AI usage reduced its cloud gross-margin percentage. Microsoft Fiscal Year 2026 Third Quarter Earnings Conference Call – Microsoft – April/2026Microsoft FY2026 Q3 investor disclosure. This is direct evidence that accelerator depreciation and AI usage are economically material. It does not disclose the cost of any particular model or token.


5.3.3 Data-centre construction, land, networking and finance

The total capital base supporting inference is broader than the accelerator fleet:

K = Kaccelerators + Kservers + Knetwork + Kstorage + Kelectrical + Kcooling + Kbuilding + Kland + Kconstruction-in-progress

Alphabet reported USD 35.7 billion of capital expenditure in Q1 2026, with the overwhelming majority directed to technical infrastructure supporting AI opportunities. Approximately 60% of that technical-infrastructure investment went to servers and 40% to data centres and networking equipment. Alphabet 2026 Q1 Earnings Call – Alphabet – April/2026Alphabet Q1 2026 investor disclosure. In Q2 2026, Alphabet reported USD 44.9 billion of capital expenditure and retained approximately the same 60:40 division between servers and data-centre/networking infrastructure. Alphabet 2026 Q2 Earnings Call – Alphabet – July/2026Alphabet Q2 2026 investor disclosure.

The disclosure establishes that analysing accelerator cost alone would omit roughly 40% of Alphabet’s recent technical-infrastructure allocation. It also demonstrates why a provider’s cash burden can rise before revenue is recognised: land, substations, buildings, cooling systems, networking and servers must be financed and constructed before capacity becomes billable. The resulting financing charge is:

Ccapital = WACC × Kemployed

where WACC is the weighted average cost of capital. Economic break-even must include this required return even if accounting operating profit is positive. A provider earning 4% on an infrastructure base while its risk-adjusted cost of capital is 9% is accounting-profitable but destroys economic value.


5.3.4 Electricity, cooling and water

Facility electricity attributable to a workload can be represented as:

Eworkload = PIT × t × PUE

Cenergy = Eworkload × pelectricity

where PUE is total facility energy divided by IT-equipment energy. Water use should be tracked separately:

Wworkload = EIT × WUE

where WUE is water-use effectiveness, normally expressed as litres per kWh of IT energy. Neither PUE nor WUE should be replaced by a universal global value. Climate, cooling design, water source, load, building age and accounting boundaries differ across facilities.

At the national level, US data centres consumed approximately 176 TWh, or 4.4% of total US electricity, in 2023; the official assessment projected 325–580 TWh, or 6.7–12% of US consumption, by 2028. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – U.S. Department of Energy – December/2024US Department of Energy data-centre electricity assessment. This does not prove that token prices must increase: efficiency, model compression and higher utilisation can offset electricity inflation. It does establish that grid access, power contracts and transmission capacity are becoming strategic inputs rather than negligible operating details.


5.3.5 Training, post-training and safety

The lifecycle cost attributable to a model family can be written as:

Cmodel-lifecycle = Cresearch + Cdata + Cpretrain + Cposttrain + Csafety + Cevaluation + Cdeployment + Cretirement

An allocation per successful commercial token would be:

Cmodel,token = Cmodel-lifecycle ÷ Tsuccessful,lifetime

The denominator is highly uncertain at launch. A successful model with trillions of billable tokens can amortise training expenditure widely. A model displaced within months may never recover its development cost. Frequent frontier releases consequently shorten the effective economic life of both model investments and optimised inference stacks.

Safety and evaluation cannot scientifically be treated as optional overhead. Their scope includes capability evaluations, domain testing, security reviews, red teaming, incident investigation, abuse monitoring, model-behaviour analysis and release governance. However, public provider price pages do not reveal how much of each token price is allocated to these activities.


5.4 Unit cost of inference

5.4.1 Core unit-cost equation

The requested inference-cost expression is:

Ctoken = [Ccompute + Cmemory + Cenergy + Cnetwork + Cfacility + Clabour + Ccapital] ÷ Tbillable

For commercial interpretation, this should be expanded to:

Csuccessful-token = [Cinference-direct + Callocated-model + Callocated-platform + Ccommercial + Crisk] ÷ Tsuccessful

where Tsuccessful excludes failed requests, free usage, promotional credits, internal tests and outputs rejected by the customer’s application.

A token-weighted blended price is:

Pblend = [PinTin + PcachedTcached + PoutTout + PreasonTreason + Ctools] ÷ Ttotal

A task-level cost is:

Ctask = PinTin + PoutTout + Pcache-writeTcache-write + Pcache-readTcache-read + ΣCtool,j + Cstorage + Cretrieval

This formulation demonstrates why advertised input-token rates cannot reliably predict the cost of an agentic workflow.


5.4.2 Provider break-even

The requested break-even relationship is:

PBE = [FC + VC + R × K] ÷ Q

where:

  • FC = attributable fixed cost;
  • VC = attributable variable cost;
  • K = invested capital;
  • R = required return on invested capital;
  • Q = commercially billable usage.

If promotional or internal use consumes capacity, a capacity-adjusted formulation is preferable:

PBE = [FC + VC + R × K] ÷ [Qphysical × ubillable × ysuccess]

where:

  • ubillable is the billable share of physical capacity;
  • ysuccess is the share producing commercially successful work.

Illustrative break-even calculation

Assume an inference platform with:

ParameterScenario value
Annual fixed costsUSD 240 million
Annual variable costsUSD 310 million
Invested capitalUSD 2.5 billion
Required return10%
Annual billable usage100 trillion tokens

Required annual revenue is:

RevenueBE = 240m + 310m + 0.10 × 2,500m
RevenueBE = USD 800 million

Break-even price per million tokens is:

PBE,MTok = 800,000,000 ÷ 100,000,000
PBE,MTok = USD 8.00

The 100 trillion-token denominator equals 100 million units of one million tokens. Sensitivity is substantial:

Billable annual tokensBreak-even price per million tokens
25 trillionUSD 32.00
50 trillionUSD 16.00
100 trillionUSD 8.00
200 trillionUSD 4.00
400 trillionUSD 2.00

The example does not estimate any named provider. It establishes the mathematical significance of scale: identical infrastructure economics can support either USD 32 or USD 2 per million tokens depending on billable throughput.


5.5 Verified public pricing architecture as of 15 August 2026

5.5.1 Selected direct API prices

The following comparison reports official list prices per one million tokens. It is not a quality-equivalence table: models differ in capability, tokenizer, speed, context, reliability and supported tools.

Provider and modelStandard inputCached input or cache hitStandard outputBatch inputBatch outputImportant modifier
OpenAI GPT-5.6 SolUSD 5.00USD 0.50USD 30.00USD 2.50USD 15.00Long context: USD 10 input and USD 45 output
OpenAI GPT-5.6 TerraUSD 2.00USD 0.20USD 12.00USD 1.00USD 6.00Long context: USD 4 input and USD 18 output
OpenAI GPT-5.6 LunaUSD 0.20USD 0.02USD 1.20USD 0.10USD 0.60Long context: USD 0.40 input and USD 1.80 output
OpenAI GPT-5.5USD 5.00USD 0.50USD 30.00Provider table identifies separate batch ratesProvider table identifies separate batch ratesAbove 272k input: 2× input and 1.5× output
OpenAI GPT-5.5 ProUSD 30.00No cached-input discountUSD 180.00SupportedSupportedHigher-compute professional tier
Anthropic Claude Opus 4.8USD 5.00USD 0.50 cache hitUSD 25.00USD 2.50USD 12.50Five-minute cache write USD 6.25
Anthropic Claude Sonnet 5USD 2.00USD 0.20 cache hitUSD 10.00USD 1.00USD 5.00Standard price retained after introductory period
Anthropic Claude Sonnet 4.6USD 3.00USD 0.30 cache hitUSD 15.00USD 1.50USD 7.50Regional endpoints may add 10%
Anthropic Claude Haiku 4.5USD 1.00USD 0.10 cache hitUSD 5.00USD 0.50USD 2.50One-hour cache write costs more than five-minute write
Google Gemini 3.6 Flash, 2026 standardUSD 0.75USD 0.075USD 3.75USD 0.375USD 1.875Officially scheduled doubling from 1 January 2027
Google Gemini 3.6 Flash, 2026 priorityUSD 1.35USD 0.135USD 6.75Not the priority modeNot the priority modeLatency/priority premium
Google Gemini 3.5 FlashUSD 1.50USD 0.15USD 9.00USD 0.75USD 4.50Output includes thinking tokens

Pricing – OpenAI API – August/2026Official OpenAI API pricing.
GPT-5.5 Model – OpenAI API – August/2026Official GPT-5.5 model pricing.
GPT-5.5 Pro Model – OpenAI API – August/2026Official GPT-5.5 Pro model pricing.
Pricing – Claude Platform Docs – August/2026Official Anthropic Claude pricing.
Gemini Developer API Pricing – Google – August/2026Official Gemini API pricing.

Three empirical observations follow.

First, output tokens are consistently more expensive than ordinary input tokens. The ratios in the selected standard tiers range from approximately 5:1 to 8:1. Output decoding is sequential and less readily parallelised than prompt processing; reasoning may add internal generation; and providers also price according to value and demand, not merely physical cost.

Second, cached inputs can cost approximately one-tenth of uncached inputs, while writing a cache can cost more than ordinary input. OpenAI states that GPT-5.6-family cache writes are billed at 1.25 times the uncached-input rate. Prompt Caching – OpenAI API – August/2026Official OpenAI prompt-caching documentation. Anthropic similarly distinguishes five-minute and one-hour cache writes from cheaper cache hits. The tariff therefore encourages repeated use of stable context but discourages indiscriminate creation of caches that are never reused.

Third, batch processing commonly receives a 50% discount. Anthropic and Google publish half-price batch rates for selected models, and Amazon Bedrock states that selected foundation models receive a 50% batch-inference discount relative to on-demand pricing. Amazon Bedrock Pricing – Amazon Web Services – August/2026Official Amazon Bedrock pricing. A consistent 50% discount is economically meaningful: it indicates that scheduling flexibility, smoother utilisation and relaxed latency materially reduce the provider’s opportunity cost.


5.5.2 Long-context pricing

Context length affects cost through at least four mechanisms:

  1. More input tokens are billed.
  2. Prefill computation increases.
  3. KV-cache memory remains occupied during generation.
  4. Long requests can reduce batching and scheduling efficiency.

A general context multiplier can be represented as:

Pcontext = Pbase × mlength

where mlength may be continuous, tiered or embedded in a model-specific price.

OpenAI’s official August 2026 table distinguishes short- and long-context prices for GPT-5.6 models. For GPT-5.6 Sol, the listed short-context rates are USD 5 input and USD 30 output per million tokens; the corresponding long-context rates are USD 10 and USD 45. Google and Anthropic apply different structures: Google lists model- and service-tier prices, while Anthropic states that Claude 4.6 and later models include the full one-million-token context window at standard token rates. These differences show that “long-context premium” can appear as an explicit multiplier, a higher model tier, cache-storage charges or simply the larger number of tokens processed.

Illustrative long-document task

Assume:

  • 500,000 uncached input tokens;
  • 20,000 output tokens;
  • no tool calls.
Model/tierInput costOutput costTotal
GPT-5.6 Sol long-contextUSD 5.00USD 0.90USD 5.90
GPT-5.6 Terra long-contextUSD 2.00USD 0.36USD 2.36
Claude Sonnet 5 standardUSD 1.00USD 0.20USD 1.20
Gemini 3.6 Flash 2026 standardUSD 0.375USD 0.075USD 0.45

The comparison does not establish equivalent answer quality. A lower-cost model becomes economically more expensive if it fails the task, requires repeated calls or produces an answer needing extensive human correction.


5.5.3 Subscriptions versus metered API access

Consumer subscriptions transform variable usage into an apparently fixed monthly payment:

ARPUsubscription = Feemonthly ÷ User

The provider’s contribution per subscriber is:

CMuser = Feemonthly − Cinference,user − Csupport,user − Cacquisition,user − Callocated-platform,user

A fixed subscription creates a usage-risk distribution. Light users subsidise heavy users within the subscriber pool unless rate limits, model routing or usage credits control the tail. Consequently, consumer plans generally include fair-use policies, flexible access, message limits, lower-priority service during congestion or automatic routing among models.

Subscriptions also perform strategic functions that a pure API tariff does not:

  • accelerate habit formation;
  • increase switching costs;
  • create persistent conversation and file ecosystems;
  • generate usage data for product improvement;
  • provide predictable recurring revenue;
  • bundle high-cost and low-cost functions;
  • obscure the marginal price of individual tasks;
  • segment customers by willingness to pay.

Enterprise contracts add identity management, data controls, audit functions, support, contractual commitments, regional processing, service levels and procurement integration. Their effective unit price cannot be inferred from public consumer fees.


5.5.4 Reserved, priority and batch capacity

Commercial modeCustomer receivesProvider receivesEconomic effect
On-demand standardFlexibility without commitmentVariable demand and uncertain utilisationStandard list price
BatchLower price; delayed completionScheduling freedom and higher utilisationDiscount
Flex or low-priorityLower price with variable latencyAbility to use residual capacityDiscount or variable service
Priority/fastLower tail latency and preferential schedulingHigher revenue per constrained capacity unitPremium
Reserved capacityGuaranteed or predictable throughputCommitted revenue and demand visibilityContracted capacity price
Regional processingGeographic controlReduced routing flexibility and regional duplicationGeographic premium
Dedicated deploymentIsolation and controlLong commitment and customer concentrationNegotiated price
Surge accessImmediate capacity during scarcityScarcity rentDynamic premium
Spot/interruptibleVery low price with interruption riskMonetisation of otherwise idle capacityDeep discount

Amazon Bedrock officially identifies Reserved, Priority, Standard and Flex service tiers. Service Tiers for Optimizing Performance and Cost – Amazon Web Services – August/2026Official Amazon Bedrock service-tier documentation. This represents an observable movement away from one uniform token price toward a yield-management model similar to cloud computing, telecommunications and transport.


5.6 Fine-tuning, retrieval, storage and agentic-tool costs

5.6.1 Fine-tuning

The full cost of a fine-tuned model is:

CFT,total = Ctraining-tokens + Cdata-preparation + Cevaluation + Chosting + Cinference + Cmaintenance

Fine-tuning can lower prompt length by internalising instructions or domain patterns, but it introduces dataset construction, validation, versioning and regression-testing costs. The break-even condition is:

Savingsprompt + Savingsquality + Savingshuman-review > CFT,total

A fine-tuned model that reduces each request by 10,000 input tokens but requires costly retraining after every base-model update may never achieve positive lifecycle value.

5.6.2 Retrieval and storage

Retrieval-augmented generation generates several charges beyond text generation:

CRAG = Cembedding + Cvector-storage + Cindexing + Cretrieval + Creranking + Cretrieved-input + Cgeneration

The customer can reduce model hallucination and avoid very long prompts, but only if retrieval precision is high. Poor retrieval increases both token expenditure and error rates.

5.6.3 Agentic tool use

An agent can call search, code execution, databases, browsers, external APIs or other models. Its cost is recursive:

Cagent = Σi=1n Cmodel,i + Σj=1m Ctool,j + Cstorage + Corchestration + Cverification

OpenAI’s official pricing lists web search at USD 10 per 1,000 calls, in addition to applicable model-token charges. Pricing – OpenAI API – August/2026Official OpenAI API and tool pricing.

Illustrative agentic workflow

Assume a research task uses:

  • 200,000 input tokens;
  • 20,000 output tokens;
  • 30 web-search calls;
  • GPT-5.6 Sol standard short-context pricing.

Model cost:

Cmodel = 0.2 × 5 + 0.02 × 30
Cmodel = USD 1.60

Search cost:

Csearch = 30 × 10 ÷ 1,000
Csearch = USD 0.30

Total direct provider charge:

Ctask = 1.60 + 0.30
Ctask = USD 1.90

If the agent enters an error loop and repeats the sequence five times, direct cost rises to USD 9.50 even though the final deliverable remains one task. Future enterprise pricing is therefore likely to emphasise budgets per completed task and controls on recursive tool use.


5.7 Is current AI pricing subsidised?

5.7.1 Definition of subsidisation

The term “subsidised” should be divided into four distinct propositions:

HypothesisTest
Short-run loss leaderPrice is below avoidable marginal cost
Full-cost under-recoveryPrice covers marginal cost but not allocated training, infrastructure and commercial costs
Below-required-return pricingAccounting profit exists but return on capital is below WACC
Cross-subsidisationLoss or low return in AI is financed by another product, segment or investor capital

These propositions require different evidence. Low price alone proves none of them.

5.7.2 Evidence consistent with subsidisation

Several observations are consistent with introductory or strategic underpricing:

  1. Capital expenditure has expanded before the full corresponding revenue stream is realised.
  2. AI infrastructure and usage have pressured cloud gross margins.
  3. Providers offer free tiers, credits, discounted batch processing and low-priced entry models.
  4. Subscription bundles can permit heavy users to consume more inference than their monthly fee would purchase at retail API prices.
  5. Model providers have strong incentives to acquire developers, establish default APIs and create ecosystem dependence.
  6. Training and post-training expenditure is not separately itemised in token prices.
  7. Rapid model replacement can shorten the amortisation period of research and infrastructure optimisation.

Microsoft reported a 66% Microsoft Cloud gross margin in fiscal Q3 2026, down year over year because of continued AI investment, while company capital expenditure reached USD 31.9 billion. Microsoft Fiscal Year 2026 Third Quarter Earnings Conference Call – Microsoft – April/2026Microsoft FY2026 Q3 investor disclosure. This supports the proposition that AI is capital-intensive and margin-dilutive at the disclosed cloud level.

5.7.3 Evidence against a universal loss-leader conclusion

The same evidence does not establish that every token is sold below cost:

  1. Microsoft Cloud remained strongly gross-profitable at the aggregate level.
  2. Google Cloud has reported positive operating income while expanding AI infrastructure.
  3. Batch and cached-input discounts may reflect genuine cost savings from scheduling and reuse.
  4. Output-token premiums may produce high contribution margins on selected workloads.
  5. Model routing can direct simple tasks toward materially cheaper models.
  6. Large providers can use their demand aggregation to achieve utilisation unavailable to smaller firms.
  7. Enterprise contracts and reserved capacity may recover costs not visible in consumer pricing.
  8. Advertising, productivity software and cloud-platform revenue can monetise AI indirectly without making the AI service economically irrational.

5.7.4 Formal test result

PropositionFinding as of August 2026
AI infrastructure is capital-intensiveStrongly supported
AI investment pressures some disclosed marginsSupported
Some access is intentionally discounted or freeSupported
Batch/flex prices reflect utilisation economicsStrongly supported
All consumer AI subscriptions are loss-makingNot established
All API tokens are sold below marginal costNot established
Frontier-model lifecycle costs are fully recovered by present token pricesNot verifiable from public disclosures
Strategic cross-subsidisation exists somewhere in the marketPlausible and consistent with business structures, but not quantifiable from disclosed segment accounts
Material future repricing is possibleSupported as a risk, not a certainty

The scientifically defensible conclusion is therefore narrower than the popular “AI companies lose money on every query” claim. Public filings support capital intensity, margin pressure and strategic investment ahead of demand. They do not disclose the model-level cost accounts required to prove universal loss-leading.


5.8 The emerging pricing system

5.8.1 Price per input token

Input-token pricing approximates prompt-processing demand, but it ignores semantic value. A million repetitive tokens and a million highly specialised tokens can cost the same despite producing different business outcomes. Caching will increasingly split input into:

  • uncached input;
  • cache-write input;
  • cache-hit input;
  • persistent-context storage.

5.8.2 Price per output token

Output pricing reflects sequential generation and captures willingness to pay. It can penalise verbose models and create incentives to produce concise answers. The effective output cost should be:

Ceffective-output = Cgenerated-output ÷ Taccepted-output

A model generating twice as much text for the same accepted result is economically less efficient even at the same listed token rate.

5.8.3 Price per reasoning token

Reasoning-token pricing monetises internal computation more directly. Its risks include limited customer observability, incentives for providers to allow inefficient reasoning and difficulty comparing models with different internal processes. A defensible system would disclose:

  • visible output tokens;
  • billed reasoning tokens;
  • maximum reasoning budget;
  • completed-task success;
  • refunds or caps for failed reasoning.

5.8.4 Price per completed task

Task pricing is likely for bounded, verifiable work:

Ptask = BaseFee + mcomplexity + mlatency + mrisk + Ctools

Examples include document classification, invoice extraction, code repair, translation, compliance review or customer-service resolution. Task pricing transfers execution-efficiency risk from the customer to the provider: if an agent uses excessive tokens, the provider absorbs the difference.

5.8.5 Price per unit of compute

A compute-based tariff could use accelerator-seconds, FLOP-equivalents or memory-bandwidth consumption:

Pcompute = aHaccelerator + bBmemory + cNnetwork

This is transparent for infrastructure buyers but unsuitable for non-technical users because identical compute can produce radically different quality.

5.8.6 Priority and latency multipliers

Ppriority = Pstandard × mpriority

The official Google table already distinguishes standard, batch, flex and priority prices, and Amazon Bedrock identifies Reserved, Priority, Standard and Flex tiers. This is direct evidence that latency is becoming separately monetised.

5.8.7 Context-length multipliers

Plong-context = Pbase × mcontext

Long-context pricing can be explicit or embedded in model tiers. Future systems may price according to maximum resident KV-cache allocation rather than only tokens actually processed.

5.8.8 Model-quality tiers

Providers already differentiate economical, balanced, frontier and high-reasoning systems. The market may evolve toward:

TierEconomic function
CommodityClassification, extraction, routing
ProfessionalGeneral business and coding
FrontierComplex reasoning and research
ExpertDomain-calibrated regulated work
SovereignGeographic, security and control guarantees

5.8.9 Reserved-capacity contracts

A reserved contract can be expressed as:

Creserved = Fcapacity + PoverageQoverage

The customer exchanges flexibility for guaranteed throughput; the provider reduces demand uncertainty and financing risk.

5.8.10 Surge pricing

Psurge,t = Pbase × [1 + λSt]

where St measures scarcity or queue pressure and λ is price sensitivity. Surge pricing could ration scarce frontier capacity during major product launches, financial events, elections, cyber incidents or regional outages. It would improve allocation efficiency but deepen inequality by allowing wealthy customers to purchase priority during scarcity.

5.8.11 Outcome-based pricing

Poutcome = Fbase + θVverified

where Vverified is measurable customer value and θ is the provider’s share. Outcome pricing is plausible for recovered revenue, resolved support cases, qualified sales leads or completed software tasks. It is difficult where causation, attribution or quality is disputed.

5.8.12 Auctioned compute capacity

An auction could allocate a fixed block of frontier inference:

Pclearing = bid of the marginal accepted buyer

Auctions could emerge for very large training runs, sovereign capacity, urgent simulation or guaranteed frontier-agent service. However, they would introduce volatility, strategic bidding and market-manipulation risks.


5.9 Five-year pricing scenarios, 2027–2031

5.9.1 Scenario framework

ScenarioProbability rangeCore mechanismToken-price directionAccess effect
Efficiency deflation25–35%Hardware and software efficiency outpace demandCommodity-model prices fall sharplyBroad basic access
Bifurcated market35–50%Small-model costs fall while frontier compute remains scarceLow tiers fall; frontier tiers remain high or riseCapability inequality persists
Infrastructure repricing15–30%Depreciation, energy and financing require fuller recoveryBroad price increases or tighter limitsSMEs and lower-income users pressured
Capacity shock5–15%Export control, conflict, grid shortage or supply disruptionSurge and regional premiumsSevere geographic inequality
Outcome transition20–40%Agents are sold by task rather than tokenToken prices lose relevanceGreater price discrimination

Ranges overlap because scenarios can coexist across models and regions.

5.9.2 Most likely structure

The central five-year case is not a universal increase in the nominal price of every token. It is a bifurcation:

  • small and medium models become cheaper;
  • cached and asynchronous work becomes heavily discounted;
  • high-priority frontier reasoning remains expensive;
  • long-context, multimodal and agentic workloads acquire additional meters;
  • enterprise data residency and reserved capacity carry premiums;
  • consumer subscriptions become more explicitly usage-limited;
  • complex tasks migrate toward outcome or credit-based pricing;
  • free tiers rely increasingly on lower-cost model routing;
  • regional scarcity creates geographic price differences.

The user-facing unit may therefore change from “one million tokens” to an opaque composite credit representing model quality, reasoning depth, tools, latency and context.


5.10 Is AI compute becoming like Bitcoin?

5.10.1 Where the analogy is useful

The analogy is useful in five limited respects.

First, both AI compute and Bitcoin-related activity depend on scarce specialised hardware, electricity and access to infrastructure. Second, both can exhibit scarcity premiums when demand expands faster than supply. Third, both can be geographically concentrated near favourable electricity, regulation or supply chains. Fourth, expectations about future value can accelerate present investment and speculative behaviour. Fifth, capacity can be represented in tradable contracts, reservations, futures or financial claims even when the underlying resource is physical.

An “AI compute index” could therefore measure the spot or forward price of a standardised capacity bundle:

ACIt = P(Hstandard, Mmemory, Bbandwidth, Llatency, Rregion, Aavailability)

Such an index could support budgeting, hedging and infrastructure contracts.

5.10.2 Where the analogy fails

PropertyBitcoinAI compute
Supply ruleProtocol-defined issuance and maximum supplyProduced through manufacturing and construction
FungibilityUnits are protocol-standardisedGPUs, TPUs, memory, models and locations differ
DurabilityDigital ledger asset does not physically wear outHardware depreciates and becomes obsolete
StorageCan be held without productive useIdle compute still incurs depreciation and facility cost
TransportabilityTransferable across a global ledgerBound to data centres, grids, networks and jurisdiction
Store of valueCan be held as an assetUnused past compute cannot generally be stored
QualityOne unit is equivalent to another at protocol levelOne accelerator-hour can vary radically in capability
Location dependenceOwnership transfer is largely location-independentLatency, sovereignty and export rules are location-sensitive
Production responseIssuance cannot exceed protocolSupply can expand, though with long lead times
ExpiryAsset does not expire technologicallyReserved capacity expires; hardware ages
Demand basisMonetary/speculative/network demandProductive demand for computation and intelligence
Unit standardisationNative unit existsNo universal equivalent across models and workloads

Compute is closer to a combination of electricity, cloud capacity, airline seats and industrial machinery than to Bitcoin. Like electricity, it must be generated and transmitted through infrastructure. Like cloud capacity, it is heterogeneous and service-dependent. Like airline seats, unused time expires and priority can be dynamically priced. Like machinery, it requires capital, maintenance and replacement.

The most important distinction is temporal non-storability. A Bitcoin held today remains a Bitcoin tomorrow. An unused accelerator-second at 14:00 cannot be sold at 15:00. Providers are therefore incentivised to discount batch and flexible workloads to fill otherwise idle capacity. This is yield management, not monetary scarcity.


5.11 Empirical tests for the next five years

Test IDPropositionRequired dataMethodFalsification condition
C5-T1Present prices under-recover full lifecycle costProvider cost accounts, capex, depreciation, usageFully allocated unit economicsRevenue per successful token persistently exceeds cost plus required return
C5-T2Batch discounts reflect utilisation benefitsBatch and on-demand throughput/costDifference-in-differencesNo material provider cost reduction from scheduling flexibility
C5-T3Output tokens have higher physical costMeasured prefill/decode resource useAccelerator telemetryOutput premium remains despite equalised physical cost and market power controls
C5-T4Frontier prices remain high while commodity prices fallHistorical model-quality-adjusted pricesHedonic panel modelFrontier-equivalent price falls at same rate as basic-model price
C5-T5Subscriptions cross-subsidise heavy usersUser-level usage and feesCohort contribution analysisEvery usage decile produces non-negative fully allocated contribution
C5-T6Data residency raises costRegional versus global endpointsMatched price comparisonNo premium after quality and SLA controls
C5-T7Agentic systems shift billing toward tasksProvider pricing cataloguesProduct-panel analysisToken-only billing remains dominant through 2031
C5-T8Scarcity increases price discriminationPriority, surge and reserved pricesPanel regressionTier dispersion falls as utilisation rises
C5-T9Infrastructure costs pressure marginsCapex, depreciation and cloud marginsDistributed-lag regressionHigher AI capital intensity consistently raises margins immediately
C5-T10Access inequality grows at the frontierRegional income and frontier usageIncome-elasticity analysisLower-income regions achieve convergent frontier usage per capita

A provider-level margin model can be estimated as:

GMj,t = α + β1AIKj,t + β2Uj,t + β3Ej,t + β4Dj,t + μj + τt + εj,t

where:

  • GMj,t = gross margin of provider j at time t;
  • AIKj,t = AI-related capital intensity;
  • Uj,t = productive utilisation;
  • Ej,t = energy cost exposure;
  • Dj,t = depreciation burden;
  • μj = provider fixed effects;
  • τt = period effects.

Because companies do not consistently disclose AI-only capital expenditure, the model must report measurement uncertainty and avoid interpreting cloud-wide correlations as pure inference economics.


5.12 Strategic implications by user class

UserPrincipal cloud advantagePrincipal cloud riskRecommended economic control
StudentFrontier quality without hardware CAPEXSubscription increases and usage limitsMixed free, subscription and local-small-model strategy
ResearcherAccess to multiple frontier modelsReproducibility and price volatilityLog exact model, tokens, tools and cost per experiment
StartupElastic capacity and rapid deploymentMargin dependence on provider tariffsModel routing, caching and contractual price protection
SMEAvoids specialist infrastructureLock-in and uncontrolled agentic expenditureBudget caps, multi-provider abstraction and task costing
UniversityAccess to high capabilityGrant budgets exposed to metered usageReserved research allocations and institutional procurement
HospitalManaged enterprise controlsSensitive-data and dependency risksRegional processing, contractual audit rights and local fallback
Large enterpriseScale and global availabilityConcentration and switching costReserved capacity plus portable evaluation layer
GovernmentImmediate access to frontier systemsSovereignty, export and continuity risksSovereign capacity, diversified supply and emergency local capability

5.13 Final assessment

The available evidence supports a nuanced conclusion. Cloud AI is not simply “cheap computation sold at a loss,” nor is it an ordinary mature software service with negligible marginal cost. It is a capital-intensive infrastructure business combined with an unusually rapid research cycle and a platform-competition strategy. Public filings demonstrate extraordinary infrastructure investment and observable margin pressure. Official pricing demonstrates strong segmentation by model quality, input/output direction, caching, context, latency, geography and scheduling flexibility. What public evidence does not demonstrate is that all current inference is priced below short-run marginal cost or that every subscription is loss-making.

The greatest five-year risk is not necessarily a uniform multiplication of every posted token price. Providers may continue reducing the cost of basic inference while shifting cost recovery into frontier reasoning, output generation, priority service, long context, data residency, tools, storage, agentic loops and guaranteed capacity. The market can therefore appear deflationary at the headline price-per-token level while becoming more expensive for economically valuable completed work.

The probable endpoint is a multi-dimensional tariff:

PAI = f(M, Tin, Tout, Treason, L, C, G, A, R, X, V)

where:

  • M = model-quality tier;
  • Tin = input tokens;
  • Tout = output tokens;
  • Treason = reasoning tokens or effort;
  • L = latency or priority;
  • C = context and cache allocation;
  • G = geographic processing requirement;
  • A = guaranteed availability;
  • R = reserved capacity;
  • X = external tools and retrieval;
  • V = verified task value or outcome.

AI access will consequently resemble a metered strategic utility more than a single software subscription. Commodity intelligence may become abundant, while reliable frontier intelligence—delivered with long context, low latency, tool access, privacy, geographic control and contractual guarantees—remains scarce and expensive. That division, rather than the nominal price of one million generic tokens, will determine the economic and social distribution of AI capability between 2027 and 2031.

CHAPTER 6 — Enterprise Exposure and the Cost of Mandatory AI Adoption

6.1 Analytical scope and central proposition

Enterprise AI adoption is entering a phase in which expenditure may become competitively necessary before its financial return can be demonstrated conclusively. A company may purchase licences, API capacity, cloud infrastructure, local accelerators, data-engineering services, cybersecurity controls and specialist personnel not because a particular project has already produced a positive net present value, but because customers, competitors, employees, regulators and investors increasingly expect AI-enabled speed, personalisation, analytics and automation. This produces a new category of expenditure: the defensive competitive input. Like cybersecurity, digital connectivity or enterprise software, AI may become difficult to avoid even when its attributable productivity gain remains uncertain. The economic risk is asymmetrical. A company that adopts too slowly may lose customers, talent and operational efficiency; a company that adopts too quickly may accumulate duplicated licences, fragmented pilots, proprietary lock-in, compliance liabilities and infrastructure whose utilisation remains too low to recover its cost.

Observed adoption is rising rapidly but remains highly unequal by enterprise size. In 2025, 20.0% of EU enterprises with at least ten employees reported using AI technologies, compared with 13.5% in 2024, 8.1% in 2023 and 7.7% in 2021. Large-enterprise adoption reached 55.03%, materially exceeding adoption among smaller firms. Use of Artificial Intelligence in Enterprises – Eurostat – December/2025Eurostat enterprise AI statistics. OECD data similarly indicate that reported firm-level AI use increased to 20.2% in 2025 from 14.2% in 2024 and 8.7% in 2023 across countries with available data. AI Use by Individuals Surges Across the OECD as Adoption by Firms Continues to Expand – OECD – January/2026OECD firm-level AI adoption statistics.

These figures establish diffusion, not profitability. A survey response indicating use of at least one AI technology does not disclose implementation depth, expenditure, productivity, risk-adjusted return or whether the system is embedded in a critical workflow. This chapter therefore constructs representative financial archetypes. Every monetary result is an analytical planning scenario expressed in 2026 USD unless otherwise stated; it is not an observed account for a named company.


6.2 Enterprise-size definitions and exposure

The European Commission defines an SME through staff headcount and either turnover or balance-sheet total. SMEs employ fewer than 250 people and ordinarily have annual turnover not exceeding EUR 50 million or a balance-sheet total not exceeding EUR 43 million. SME Definition – European Commission – August/2026European Commission SME definition.

For modelling, this chapter uses the following operational classes:

Enterprise classEmployeesIllustrative annual revenueAI-adoption characteristics
Microenterprise1–9USD 0.25–2 millionSubscription-led; little internal technical capacity
Small enterprise10–49USD 2–15 millionSaaS and managed-service adoption
Medium enterprise50–249USD 15–100 millionHybrid integration; emerging governance function
Large corporation250+USD 100 million–10 billion+Portfolio adoption, proprietary integration and dedicated governance
Systemically or strategically critical organisationVariableBudget or revenue can exceed USD 1 billionHigh assurance, resilience, sovereignty and regulatory obligations

The economic importance of small firms makes unequal adoption a macroeconomic problem rather than a niche technology issue. Eurostat reported that micro and small businesses represented 99% of EU enterprises in 2022, employed 77.5 million people and generated EUR 11.9 trillion of turnover. Large firms represented only 0.2% of enterprises but generated 51% of turnover. Micro and Small Businesses Make Up 99% of Enterprises in the EU – Eurostat – October/2024Eurostat enterprise-size statistics. If AI entails large fixed costs for governance, data preparation and integration, those costs will consume a larger share of small-firm revenue even where API prices are identical.


6.3 Core enterprise financial metrics

6.3.1 AI Cost Ratio

AI Cost Ratio = Annual AI Expenditure ÷ Revenue

The ratio measures the share of revenue absorbed by AI expenditure. It should be calculated both gross and net of capitalised development:

AI Cost Ratiocash = AI Cash Expenditure ÷ Revenue

AI Cost RatioP&L = AI Expense Recognised in Period ÷ Revenue

A company purchasing infrastructure may record a large cash outflow but recognise depreciation over several years. The cash ratio measures financing pressure; the income-statement ratio measures current accounting burden.

6.3.2 AI Labour Burden

AI Labour Burden = Annual AI Expenditure ÷ Annual Labour Cost

This ratio compares AI expenditure with the labour base from which productivity benefits are commonly expected. It is economically informative because an AI programme costing 15% of annual labour expense must create a very large improvement in labour productivity, revenue or risk reduction to break even.

6.3.3 Net AI Value

Net AI Value = Productivity Gains + Incremental Revenue − Risk Costs − AI Expenditure

For financial accuracy, incremental revenue should be converted into incremental contribution, not treated as entirely profit:

Net AI Value = Labour Savings + Non-labour Savings + m × Incremental Revenue + Avoided Losses − Risk Costs − AI Expenditure

where m is the contribution margin on incremental revenue.

6.3.4 Return on investment

ROIAI = [NPV(Benefits) − NPV(Costs)] ÷ NPV(Costs)

The net present value of benefits is:

NPV(Benefits) = Σt=0n Bt ÷ (1+r)t

The net present value of costs is:

NPV(Costs) = Σt=0n Ct ÷ (1+r)t

The project creates financial value only when:

NPV(Benefits) > NPV(Costs)

and:

ROIAI > 0

For a risk-adjusted investment hurdle h, acceptance requires:

ROIAI ≥ h

A positive ROI below the company’s risk-adjusted hurdle rate can still represent value destruction relative to alternative investments.


6.4 The complete enterprise AI cost stack

6.4.1 Direct and indirect costs

Cost categoryIncluded expenditurePrimary driverFrequently omitted element
User licencesPer-seat AI assistants and embedded softwareNumber of enabled usersPaid but inactive seats
API consumptionInput, output, reasoning and multimodal usageWorkload volume and model tierFailed calls and agent loops
Cloud infrastructureCompute, databases, storage and networkingData volume and availabilityEgress and regional-processing premiums
Local infrastructureServers, accelerators, storage and facilitiesModel size, concurrency and resilienceSpares, replacement and cooling
CybersecurityIdentity, monitoring, testing and incident responseThreat profile and system criticalityModel-specific red teaming
ComplianceLegal assessment, documentation and conformity workSector, jurisdiction and risk classificationContinuous evidence maintenance
Data preparationCleaning, labelling, permissions and lineageData fragmentation and qualityRights remediation
GovernancePolicies, inventory, committees and controlsNumber and criticality of systemsShadow AI discovery
Model evaluationBenchmarks, validation, bias and robustness testsUse-case riskRegression testing after model updates
Human supervisionReview, escalation and overrideError severity and regulatory exposureReviewer fatigue
IntegrationAPIs, workflow redesign, testing and migrationLegacy-system complexityProcess redesign
Specialist personnelAI engineers, data scientists, security and legal staffScope and assuranceRecruitment and retention premium
Vendor switchingRe-engineering, data migration and evaluationProprietary dependenceTemporary parallel operation
DowntimeLost output, service interruption and recoveryAvailability and fallback capacityQueue accumulation after recovery
Model errorCorrection, rework, customer harm and liabilityError probability and consequenceCorrelated systematic failure
Privacy-preserving deploymentIsolation, encryption, local inference and auditingData sensitivityLower utilisation from dedicated capacity
TrainingAI literacy and role-specific competenceWorkforce sizeRefresher training after updates
Change managementProcess redesign, communication and adoptionOrganisational complexityProductivity dip during transition
Opportunity costCapital and staff diverted from other projectsPortfolio constraintsAbandoned alternatives

Annual AI expenditure should therefore be calculated as:

AIE = Llicence + AAPI + Ccloud + Ilocal + Ssecurity + Ccompliance + Ddata + Ggovernance + Eevaluation + Hsupervision + Iintegration + Ppersonnel + Vswitching + Ddowntime + Rerror

A narrow estimate including only licences and API charges can understate full economic expenditure by several multiples.


6.4.2 Risk-adjusted error costs

The expected annual cost of model errors is:

E(Risk Cost) = Σj=1m pj × Lj

where:

  • pj is the annual probability of error class j;
  • Lj is its total loss if realised.

Loss should include:

Lj = Ccorrection + Cdowntime + Ccustomer + Clegal + Cregulatory + Creputation + Cremediation

A low-frequency error can dominate AI economics when its consequence is large. An AI system producing USD 2 million in annual process savings but carrying a 2% probability of a USD 100 million loss has an expected loss of USD 2 million before risk aversion, capital requirements or tail-risk limits are considered.

For risk-averse institutions, expected value is insufficient. A more appropriate decision rule is:

Risk-Adjusted Net AI Value = Expected Benefits − Expected Costs − λ × Tail Risk

where λ reflects the organisation’s risk tolerance and Tail Risk can be measured through value at risk, expected shortfall or scenario loss.


6.5 Representative enterprise archetypes

6.5.1 Baseline annual scenarios

The following table applies one consistent annual model. “Productivity gains” represent realised cash savings or additional productive capacity assigned a defensible monetary value. Incremental revenue is converted to contribution using the margins embedded in each scenario. Risk cost includes expected error, downtime, privacy and liability losses.

ArchetypeRevenue or operating budgetLabour costAnnual AI expenditureProductivity gainIncremental contributionRisk costNet AI value
Micro professional firmUSD 0.80mUSD 0.32mUSD 0.022mUSD 0.018mUSD 0.008mUSD 0.004mUSD 0.000m
Small enterpriseUSD 5mUSD 1.5mUSD 0.150mUSD 0.120mUSD 0.060mUSD 0.030mUSD 0.000m
Medium enterpriseUSD 30mUSD 7mUSD 0.600mUSD 0.560mUSD 0.240mUSD 0.100mUSD 0.100m
Large diversified companyUSD 5bnUSD 1.2bnUSD 65mUSD 75mUSD 40mUSD 15mUSD 35m
BankUSD 10bnUSD 3bnUSD 180mUSD 200mUSD 90mUSD 70mUSD 40m
InsurerUSD 5bnUSD 1bnUSD 90mUSD 100mUSD 50mUSD 45mUSD 15m
Healthcare organisationUSD 2bnUSD 1.1bnUSD 55mUSD 70mUSD 10mUSD 35m−USD 10m
Legal/professional firmUSD 100mUSD 55mUSD 4mUSD 5.5mUSD 1.5mUSD 1mUSD 2m
ManufacturerUSD 1bnUSD 180mUSD 18mUSD 20mUSD 12mUSD 6mUSD 8m
Media companyUSD 200mUSD 70mUSD 6mUSD 7mUSD 4mUSD 3mUSD 2m
Telecommunications operatorUSD 5bnUSD 900mUSD 100mUSD 120mUSD 80mUSD 25mUSD 75m
Defence/critical-infrastructure operatorUSD 3bnUSD 700mUSD 75mUSD 65mUSD 30mUSD 35m−USD 15m
Public administrationUSD 2bn budgetUSD 900mUSD 60mUSD 70mUSD 0mUSD 30m−USD 20m
University/research systemUSD 1bn budgetUSD 600mUSD 35mUSD 35mUSD 5mUSD 20m−USD 15m

These figures are not sector averages. They are transparent archetypes designed to identify financial thresholds. Negative results do not mean AI is socially undesirable or strategically avoidable. A hospital, defence organisation or public authority may rationally accept a negative accounting return to obtain safety, sovereignty, service quality or strategic capability. The correct conclusion is that such deployments require a broader public-value or mission-value framework rather than a false claim of immediate commercial ROI.


6.5.2 AI Cost Ratio and AI Labour Burden

ArchetypeAI Cost RatioAI Labour BurdenAnnual return on AI expenditure before discounting
Micro professional firm2.75%6.88%0.0%
Small enterprise3.00%10.00%0.0%
Medium enterprise2.00%8.57%16.7%
Large diversified company1.30%5.42%53.8%
Bank1.80%6.00%22.2%
Insurer1.80%9.00%16.7%
Healthcare organisation2.75%5.00%−18.2%
Legal/professional firm4.00%7.27%50.0%
Manufacturer1.80%10.00%44.4%
Media company3.00%8.57%33.3%
Telecommunications operator2.00%11.11%75.0%
Defence/critical infrastructure2.50%10.71%−20.0%
Public administration3.00% of budget6.67%−33.3%
University/research system3.50% of budget5.83%−42.9%

The calculation demonstrates why the same nominal expenditure has different consequences across sectors. Legal services may support a relatively high AI Cost Ratio because labour represents a large part of cost and document-intensive work is potentially augmentable. Manufacturing may have a lower AI Cost Ratio but still a high AI Labour Burden because revenue includes substantial material and energy costs. Public administration and education may create social benefits not recorded as incremental revenue, making private-sector ROI formulas incomplete.


6.6 Cost structure by organisation size

6.6.1 Microenterprise

A representative five-person professional business may purchase:

Cost itemAnnual scenario
Five AI subscriptionsUSD 3,000
Accounting, document and CRM AI featuresUSD 2,500
API and automationUSD 1,500
External integrationUSD 4,000
Data cleanup and templatesUSD 2,000
Cybersecurity and privacy controlsUSD 2,500
Training and governanceUSD 2,000
Human review and error correctionUSD 2,500
Contingency and switchingUSD 2,000
TotalUSD 22,000

The main risk is not token consumption but fixed implementation overhead. At USD 0.8 million revenue, the expenditure equals 2.75% of turnover. If net operating margin before AI is 10%, annual operating profit is USD 80,000; AI consumes 27.5% of pre-AI profit. The relevant ratio is therefore:

AI Profit Burden = Annual AI Expenditure ÷ Pre-AI Operating Profit

AI Profit Burden = 22,000 ÷ 80,000 = 27.5%

A cost equal to only 2.75% of revenue can be strategically significant when margins are thin.

6.6.2 SME

SMEs face an intermediate problem. Their workflows are complex enough to require integration and governance, but their scale may be insufficient to spread those costs widely. The OECD reported that AI adoption among firms remains lower than adoption of other digital technologies and that significant gaps persist between SMEs and large firms. AI Adoption by Small and Medium-Sized Enterprises – OECD – December/2025OECD report on SME AI adoption.

A USD 5 million-revenue small enterprise spending USD 150,000 annually on AI must achieve one or a combination of:

  • a 10% gross saving on its USD 1.5 million labour bill;
  • a 3% reduction in total non-AI operating costs if those equal USD 5 million;
  • USD 500,000 of incremental revenue at a 30% contribution margin;
  • avoided annual losses of USD 150,000;
  • a mixed benefit portfolio with equivalent value.

If the firm realises only half the expected productivity gain, AI destroys value unless adoption prevents a larger competitive loss.

6.6.3 Large corporation

Large firms possess several advantages:

  • fixed governance costs are spread across more revenue;
  • internal data and use cases are more abundant;
  • enterprise bargaining power can reduce unit prices;
  • specialist staff can be employed internally;
  • workloads can be routed across models;
  • larger usage can justify reserved capacity;
  • benefits can be diversified across many business functions.

They also face disadvantages:

  • integration with legacy systems;
  • duplicated departmental pilots;
  • larger cyberattack surfaces;
  • complex regulatory exposure;
  • slower change management;
  • high cost of correlated errors;
  • vendor concentration at enormous scale.

The large-firm advantage is therefore not simply “more money.” It is the capacity to convert fixed AI costs into a lower cost per employee, transaction and unit of revenue.


6.7 Sector-by-sector exposure

6.7.1 Banking

AI use is already widespread among significant European banks. The European Central Bank reported in June 2026 that more than 85% of banks under European banking supervision used AI. Strengthening Operational Resilience for the Age of AI – European Central Bank – June/2026ECB assessment of AI and banking resilience.

Banks can deploy AI in:

  • fraud detection;
  • transaction monitoring;
  • credit analysis;
  • customer service;
  • coding;
  • document processing;
  • regulatory reporting;
  • cybersecurity;
  • risk modelling;
  • collections;
  • knowledge management.

The expected-value equation is:

Net AI Valuebank = Operating Savings + Fraud Losses Avoided + Credit Losses Avoided + Incremental Margin − Compliance Cost − Model Risk − AI Expenditure

Banking has high benefit potential but also high tail risk. An error in an internal drafting assistant is not equivalent to a systematic credit-scoring error. Data governance, explainability, validation, monitoring and third-party concentration must be allocated by use case. The ECB has warned that AI may improve operational efficiency while increasing operational risk and third-party dependence. The Rise of Artificial Intelligence: Benefits and Risks for Financial Stability – European Central Bank – May/2024ECB financial-stability analysis.

6.7.2 Insurance

Insurance AI can improve underwriting, claims triage, fraud detection, pricing and customer interaction. Its value equation is:

Net AI Valueinsurance = Expense Savings + Claims Leakage Reduction + Pricing Improvement + Incremental Premium Margin − Conduct Risk − Model Risk − AI Expenditure

The principal danger is that historical claims data encode past exclusions or biases. A small average improvement in loss ratio can be economically valuable, but a systematic error across thousands of policies can generate correlated liability.

6.7.3 Healthcare

Healthcare benefits include clinical-documentation support, scheduling, coding, imaging assistance, research and operational planning. Costs include validation, security, clinical supervision, integration with electronic health records and high-availability infrastructure.

Net AI Valuehealth = Administrative Savings + Capacity Value + Avoided Error + Health Outcome Value − Clinical Risk − Privacy Risk − AI Expenditure

A purely financial calculation may understate public benefit. Conversely, monetising every minute “saved” is invalid if clinicians use the time for additional verification or if staffing levels do not change. Productivity capacity becomes cash savings only when it reduces overtime, avoids hiring, increases reimbursable activity or improves outcomes sufficiently to produce measurable value.

6.7.4 Legal and professional services

Legal and advisory firms have high labour intensity and document-heavy workflows, giving them strong theoretical exposure to productivity gains. However, time saved can reduce billable hours under hourly pricing.

Net AI Valuelegal = Cost Savings + New Matters + Higher Realisation + Capacity Value − Lost Billable Hours − Liability − AI Expenditure

The business-model effect is decisive:

Billing modelAI productivity effect
Hourly billingTime savings can reduce revenue unless volume rises
Fixed feeEfficiency increases margin
Subscription/retainerEfficiency can increase capacity
Outcome feeAI value depends on result quality
Internal legal departmentEfficiency reduces cost or unmet demand

The firm may need to shift pricing before productivity becomes financially valuable.

6.7.5 Manufacturing

Manufacturing use cases include predictive maintenance, quality inspection, process control, supply-chain planning, engineering support, procurement and technical documentation.

Net AI Valuemanufacturing = Downtime Avoided + Scrap Reduction + Yield Improvement + Labour Capacity + Inventory Reduction + Incremental Margin − Safety Risk − Integration Cost − AI Expenditure

The benefit should be calculated from physical variables:

Benefitscrap = Units × Scrap Reduction × Cost per Unit

Benefitdowntime = Hours Avoided × Contribution per Production Hour

Benefitinventory = Inventory Reduction × Carrying Cost Rate

Manufacturing AI can create value without reducing headcount. Yield, energy, scrap and uptime may dominate office-productivity savings.

6.7.6 Media

Media AI can reduce transcription, translation, editing, tagging, personalisation and production costs. It also creates copyright, provenance, trust and substitution risks.

Net AI Valuemedia = Production Savings + Audience Revenue + Archive Monetisation − Rights Risk − Brand Damage − Displacement Loss − AI Expenditure

The EU’s Article 50 transparency obligations became applicable from 2 August 2026 and address marking and detection of certain AI-generated content, including deepfakes and specified publications. Code of Practice on Transparency of AI-Generated Content – European Commission – July/2026European Commission AI-content transparency framework.

6.7.7 Telecommunications

Telecommunications operators can apply AI to network optimisation, predictive maintenance, customer service, fraud, churn, sales, cybersecurity and capacity planning. Their scale and recurring data flows make them strong candidates for positive AI economics. However, telecommunications systems are critical infrastructure, and dependence on external models can create availability and sovereignty risks.

Net AI Valuetelecom = Network Savings + Churn Reduction + Fraud Avoidance + Incremental ARPU + Service Automation − Outage Risk − Security Risk − AI Expenditure

A one-percentage-point reduction in churn can be worth more than large administrative savings; the correct unit of value may be retained customer lifetime value rather than labour hours.

6.7.8 Defence and critical infrastructure

Defence and critical infrastructure require higher assurance, isolated environments, sovereign control, testing, redundancy and supply-chain review. The financial return may appear negative because mission value is not fully recorded as revenue.

Mission-Adjusted AI Value = Financial Benefits + Resilience Value + Sovereignty Value + Capability Value − Mission Risk − AI Expenditure

A system that improves intelligence processing but creates an unacceptable adversarial-attack surface should not be approved merely because its expected financial return is positive.

6.7.9 Public administration

Public administration can use AI for document handling, translation, citizen support, fraud analysis and policy research. Benefits should include service quality and waiting-time reduction:

Public AI Value = Administrative Savings + Citizen Time Saved + Service Quality + Fraud Avoided − Rights Risk − Error Cost − AI Expenditure

Public bodies cannot assume that staff time saved becomes fiscal savings. If no positions, overtime or procurement costs are reduced, the benefit is improved capacity rather than cash.

6.7.10 Education and research

Education and research institutions face a dual burden: purchasing AI access while also redesigning assessment, research integrity, data governance and teaching.

Education AI Value = Teaching Capacity + Research Acceleration + Student Support + Accessibility − Integrity Risk − Inequality Cost − AI Expenditure

If affluent institutions purchase frontier models while poorer institutions rely on lower-capability systems, AI can amplify the research and education divide described in earlier chapters.


6.8 Regulatory and governance burden

The EU AI Act entered into force on 1 August 2024 and became generally applicable on 2 August 2026, subject to phased exceptions. Governance and general-purpose-model obligations became applicable earlier; certain high-risk provisions have extended timelines following the 2026 AI Omnibus. AI Act Regulatory Framework – European Commission – August/2026Official European Commission AI Act timeline.

Compliance costs depend on role and risk. A company can be:

  • a provider;
  • a deployer;
  • an importer;
  • a distributor;
  • a product manufacturer;
  • a downstream modifier.

The annual compliance burden can be represented as:

Ccompliance = Cinventory + Cclassification + Cdocumentation + Ctesting + Cmonitoring + Ctraining + Clegal + Cincident

Regulation is not the only driver. Even where a system is not legally classified as high-risk, contractual liability, professional standards, cybersecurity and customer expectations can require comparable controls.

NIST’s Generative AI Profile treats risk management as a lifecycle covering governance, mapping, measurement and management. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile – NIST – July/2024NIST Generative AI Profile. This supports a governance model in which evaluation and monitoring are recurring operating expenses rather than one-time launch tasks.


6.9 Deployment-model comparison

6.9.1 Structural comparison

Deployment modelInitial CAPEXVariable costPrivacy controlFrontier capabilitySwitching costInternal skill requirement
Full cloud proprietaryLowHigh/usage-basedMediumVery highHighMedium
Fully local proprietaryVery highModerateHighLimited by available hardware/licenceHighVery high
Fully local open-weightHighModerateVery highModel-dependentMediumVery high
HybridMedium/highMediumHigh for sensitive workloadsHighMediumHigh
Specialised small modelLow/mediumLowHigh if localHigh for bounded taskLow/mediumMedium/high
Retrieval-augmented proprietaryMediumMedium/highMediumHighMedium/highHigh
Retrieval-augmented open-weightMedium/highLow/moderateHighModel-dependentLow/mediumHigh
Shared/cooperative computeShared CAPEXSharedPotentially highMedium/highGovernance-dependentShared specialist team

6.9.2 Financial scoring

Scores range from 1, least favourable, to 5, most favourable.

ModelCost predictabilityLow entry costPrivacyScalabilityPortabilityHigh-assurance suitability
Cloud proprietary352523
Fully local415235
Hybrid434445
Open-weight local425354
Specialised small model544344
RAG with frontier API343533
Cooperative compute434344

No model dominates every dimension. Hybrid deployment is frequently economically attractive because it routes sensitive or predictable workloads locally and sends high-complexity or burst demand to the cloud.


6.10 Cloud, local and hybrid break-even

Let:

Ccloud(Q) = Fcloud + pcloudQ

Clocal(Q) = Flocal + plocalQ

Cloud and local costs are equal at:

Q* = [Flocal − Fcloud] ÷ [pcloud − plocal]

Assume:

  • cloud fixed integration cost = USD 100,000;
  • local fixed annualised cost = USD 700,000;
  • cloud variable cost = USD 12 per million successful tokens;
  • local variable cost = USD 3 per million successful tokens.

Then:

Q = (700,000 − 100,000) ÷ (12 − 3)
Q
= 66,667 million tokens
Q* ≈ 66.7 billion successful tokens annually

Below approximately 66.7 billion successful tokens, the cloud scenario is cheaper under these assumptions. Above it, local infrastructure becomes cheaper if quality, utilisation and reliability are equivalent. They frequently are not equivalent, so a quality-adjusted threshold is needed:

Q*QA = [Flocal − Fcloud] ÷ [pcloudycloud − plocalylocal]

where y represents the cost adjustment required to obtain a successful workload.


6.11 Mandatory competitive adoption

6.11.1 The adoption game

Consider two competing firms. Each can adopt or not adopt AI.

Firm A / Firm BB adoptsB does not adopt
A adoptsBoth incur AI costs; relative advantage may disappearA may gain speed, cost or product advantage
A does not adoptA risks losing share or talentNeither pays; existing equilibrium continues

If both firms adopt and obtain identical productivity gains that competition passes to customers through lower prices, neither may retain the benefit as profit. AI then becomes a Red Queen investment: each firm must invest merely to maintain its relative position.

The private decision rule becomes:

Adopt if:

Net AI Value + Competitive Loss Avoided > 0

Define:

Dcompetitive = Revenue or margin lost if the company does not adopt

Then:

Strategic Net AI Value = Net AI Value + Dcompetitive

A project with direct Net AI Value of −USD 2 million can be rational if non-adoption is expected to destroy USD 5 million of contribution margin. Its strategic value is positive USD 3 million.

6.11.2 Mandatory-AI tax

The de facto competitive burden can be expressed as:

Mandatory AI Tax = Minimum AI Expenditure Required to Maintain Competitive Parity ÷ Revenue

Company typeIllustrative minimum parity expenditureRevenueMandatory AI Tax
Micro professional firmUSD 10,000USD 0.8m1.25%
Small enterpriseUSD 75,000USD 5m1.50%
Medium enterpriseUSD 300,000USD 30m1.00%
Large enterpriseUSD 25mUSD 5bn0.50%
BankUSD 80mUSD 10bn0.80%
Media companyUSD 3mUSD 200m1.50%

The absolute burden is larger for large firms, but the proportional burden can be greater for small firms because governance, integration and security contain indivisible fixed components.


6.12 Value-destruction thresholds

6.12.1 Minimum labour-productivity gain

Let:

  • A = annual AI expenditure;
  • RC = annual risk cost;
  • L = annual labour cost;
  • ΔR = incremental revenue;
  • m = contribution margin.

Break-even requires:

gLL + mΔR ≥ A + RC

The minimum required labour-productivity gain is:

gL* = [A + RC − mΔR] ÷ L

Archetype thresholds

ArchetypeAI expenditure plus risk costIncremental contributionLabour costMinimum labour productivity gain
Micro firmUSD 26kUSD 8kUSD 320k5.63%
Small enterpriseUSD 180kUSD 60kUSD 1.5m8.00%
Medium enterpriseUSD 700kUSD 240kUSD 7m6.57%
Large companyUSD 80mUSD 40mUSD 1.2bn3.33%
BankUSD 250mUSD 90mUSD 3bn5.33%
Healthcare organisationUSD 90mUSD 10mUSD 1.1bn7.27%
Legal firmUSD 5mUSD 1.5mUSD 55m6.36%
ManufacturerUSD 24mUSD 12mUSD 180m6.67%
Telecommunications operatorUSD 125mUSD 80mUSD 900m5.00%
Defence operatorUSD 110mUSD 30mUSD 700m11.43%
Public administrationUSD 90mUSD 0mUSD 900m10.00%
University systemUSD 55mUSD 5mUSD 600m8.33%

A “productivity gain” must become economically usable. If employees save 8% of their time but workload, staffing, revenue and service output remain unchanged, the realised financial gain may be close to zero.

6.12.2 Minimum incremental revenue

If no labour saving occurs:

ΔR* = [A + RC − S] ÷ m

where S is other verified cost savings.

For a small enterprise with USD 150,000 AI expenditure, USD 30,000 risk cost, USD 50,000 other savings and a 30% contribution margin:

ΔR = (150,000 + 30,000 − 50,000) ÷ 0.30
ΔR
= USD 433,333

The company must generate more than USD 433,333 of additional annual revenue merely to break even.

6.12.3 Minimum operating margin

If annual AI expenditure is a proportion a of revenue and pre-AI operating margin is m0:

Post-AI margin before benefits = m0 − a

A company with a 4% operating margin and AI costs equal to 3% of revenue loses 75% of operating profit before benefits.

AI Profit Burden = a ÷ m0

Pre-AI operating marginAI Cost RatioShare of operating profit consumed
3%1%33.3%
3%2%66.7%
3%3%100.0%
5%2%40.0%
5%3%60.0%
10%2%20.0%
20%2%10.0%
30%2%6.7%

Low-margin industries are therefore highly exposed even when AI costs appear modest relative to revenue.


6.13 Five-year NPV scenario

Consider a medium enterprise with:

  • initial integration and data-preparation cost: USD 1.2 million;
  • recurring year-one cost: USD 600,000;
  • recurring cost growth: 8% annually;
  • year-one benefit: USD 700,000;
  • benefit growth: 15% annually;
  • discount rate: 9%;
  • terminal value excluded;
  • expected error and downtime costs included in recurring cost.
YearCostBenefitNet cash flow
0USD 1.200mUSD 0−USD 1.200m
1USD 0.600mUSD 0.700mUSD 0.100m
2USD 0.648mUSD 0.805mUSD 0.157m
3USD 0.700mUSD 0.926mUSD 0.226m
4USD 0.756mUSD 1.065mUSD 0.309m
5USD 0.816mUSD 1.225mUSD 0.409m

Approximate discounted values:

MeasureValue
NPV of benefitsUSD 3.55m
NPV of recurring costsUSD 2.69m
Initial costUSD 1.20m
Total NPV of costsUSD 3.89m
Net project NPV−USD 0.34m
ROIAI−8.7%

Although annual net cash flow is positive from year one, the project fails to recover its initial integration cost within the five-year discounted horizon. This illustrates why pilot-level operating savings do not prove enterprise value.


6.14 Sensitivity analysis

Using the medium-enterprise baseline, the five-year result is most sensitive to adoption, benefit realisation and risk.

Benefit realisationCost overrunRisk-cost multiplierLikely financial result
50%25%2.0×Severe value destruction
75%15%1.5×Negative
100%0%1.0×Near break-even
125%0%1.0×Positive
125%−10%0.75×Strongly positive
150%−15%0.50×Transformational

The realised benefit ratio is:

BRR = Realised Benefits ÷ Forecast Benefits

The project should trigger remediation when:

BRR < BRRminimum

If forecast benefits are USD 1 million and break-even requires USD 750,000, then:

BRRminimum = 750,000 ÷ 1,000,000 = 75%


6.15 Open-weight, small-model and cooperative alternatives

6.15.1 Open-weight systems

Open-weight models reduce dependence on a single inference provider and allow local inspection, fine-tuning and deployment. They do not eliminate cost. The enterprise assumes responsibility for:

  • infrastructure;
  • optimisation;
  • security;
  • model evaluation;
  • updates;
  • monitoring;
  • licensing analysis;
  • incident response;
  • specialist staffing.

The appropriate comparison is:

TCOopen versus TCOproprietary

not “free model” versus “paid API.”

6.15.2 Small-model specialisation

A smaller model can create superior enterprise value when:

Qualitysmall,task ≥ Qualityrequired

and:

TCOsmall < TCOfrontier

Task-specific evaluation may reveal that a small model with retrieval outperforms a general frontier model on controlled extraction, classification or internal terminology while costing less and exposing less data.

6.15.3 Retrieval-augmented generation

RAG is economically attractive when the value of improved grounding and shorter model context exceeds indexing and retrieval costs:

ValueRAG = Error Reduction + Prompt Savings + Update Flexibility − Retrieval Cost − Integration Cost

RAG does not solve poor data governance. It can retrieve incorrect, obsolete or unauthorised content with high confidence.

6.15.4 Shared or cooperative compute

Universities, hospitals, municipalities, professional associations and SMEs can pool infrastructure:

Cost per Member = [Shared Fixed Cost + Shared Operating Cost] ÷ Members + Member-Specific Cost

Cooperation can reduce fixed-cost inequality but introduces governance questions:

  • capacity allocation;
  • confidentiality;
  • liability;
  • model selection;
  • maintenance responsibility;
  • admission and exit;
  • priority during scarcity;
  • cost distribution;
  • data separation.

A cooperative model is particularly attractive where members share assurance requirements but do not individually possess sufficient usage to justify dedicated infrastructure.


6.16 Enterprise decision framework

An AI investment should proceed only after answering:

TestRequired evidence
Strategic necessityQuantified competitive loss from non-adoption
Task suitabilityBaseline and model performance on representative work
Financial caseFive-year NPV, ROI and payback
Benefit convertibilityMechanism translating time savings into cash or output
Risk toleranceExpected and tail-loss assessment
Data readinessRights, quality, lineage and access
Deployment choiceCloud, local, hybrid or cooperative comparison
Exit feasibilitySwitching cost and portable data/model layer
GovernanceNamed owner, inventory, monitoring and escalation
Human controlReview threshold and accountable decision-maker
ResilienceFallback, outage process and recovery target
MeasurementPost-deployment control group or credible counterfactual

The company should establish a counterfactual:

Incremental AI Value = OutcomeAI − Outcomeno-AI

Without a control group, phased rollout, matched workflow or credible baseline, ordinary business improvement may be misattributed to AI.


6.17 Five-year enterprise outlook, 2027–2031

The most probable outcome is not universal replacement of labour but the progressive conversion of AI into a mandatory layer of enterprise infrastructure. Adoption costs will move from visible subscriptions toward less visible integration, governance, evaluation and risk-management expenditure. Large firms will increasingly build model-routing platforms combining proprietary frontier models, specialised small models, retrieval systems and local infrastructure. Smaller firms will remain more dependent on bundled SaaS products because they cannot efficiently internalise governance and engineering. This will transfer bargaining power to software vendors and create a persistent risk that AI becomes an unavoidable surcharge embedded across accounting, productivity, CRM, design, legal, security and communications products.

Three enterprise equilibria are plausible:

EquilibriumDescriptionDistributional consequence
Productivity diffusionBenefits exceed costs across firm sizesBroad increase in output and lower prices
Competitive compulsionFirms adopt to maintain parity; gains pass to customersAI becomes mandatory overhead
Capability concentrationLarge firms capture scale economies and better modelsSME margins and market share decline

The second and third outcomes can coexist. Customers may receive faster or cheaper services while enterprise profits become more concentrated among AI infrastructure and platform providers.


6.18 Final findings

  1. AI adoption is accelerating, but use rates do not establish profitability. EU enterprise adoption rose to approximately 20% in 2025, while large-enterprise use exceeded 55%.
  2. The relevant enterprise cost is substantially larger than licence and API expenditure. Data, integration, cybersecurity, governance, evaluation, human supervision and switching must be included.
  3. Small firms face a disproportionate fixed-cost burden. Identical governance work consumes a much larger share of small-firm revenue and profit.
  4. Low-margin companies are exposed even at apparently small AI Cost Ratios. AI expenditure equal to 3% of revenue eliminates pre-AI operating profit when the margin is also 3%.
  5. Time saved is not automatically financial value. It becomes cash benefit only through lower cost, avoided hiring, increased output, higher revenue or improved public service.
  6. High-risk sectors require tail-risk analysis. Expected productivity gains cannot justify deployment when a low-probability failure threatens catastrophic loss.
  7. Cloud deployment minimises initial capital but can maximise dependency. Fully local deployment improves control but requires high utilisation and specialist staff.
  8. Hybrid systems are likely to dominate mature enterprise architecture. Sensitive and predictable workloads can remain local while frontier and burst demand use external capacity.
  9. Open-weight models eliminate neither infrastructure nor governance costs. Their economic advantage is control and portability, not zero TCO.
  10. Small specialised models can outperform frontier systems economically on bounded tasks. The relevant criterion is successful workload cost, not nominal parameter count.
  11. AI may become a Red Queen investment. Companies may pay merely to preserve competitive parity, with productivity gains passed to customers rather than retained as profit.
  12. The minimum productivity threshold can be calculated explicitly. It rises with AI expenditure, risk costs and low contribution margins and falls with verifiable revenue or non-labour benefits.
  13. Positive annual savings do not guarantee positive five-year NPV. Initial integration and recurring cost escalation can leave a project value-destructive.
  14. Defence, healthcare, government and education require mission-adjusted value. Commercial ROI alone omits resilience, sovereignty, safety and public benefit.
  15. The decisive divide from 2027 to 2031 will be organisational capability, not simple access to a chatbot. Firms able to govern data, evaluate models, route workloads and negotiate infrastructure will capture disproportionate value; firms purchasing fragmented AI tools without measurement may incur a new mandatory cost without receiving a durable productivity return.

CHAPTER 7 — The AI Capability Divide: Students, Researchers and Social Mobility

7.1 Scope and central research question

The emerging AI divide is not adequately described by whether an individual has ever used a chatbot. Meaningful access has at least six dimensions: availability, affordability, model capability, usage allowance, hardware and connectivity, and the user’s ability to integrate AI into productive work. A nominally free service may impose restrictive rate limits, use a lower-capability model, lack long-context document analysis, provide no persistent project environment, withhold advanced tools or become unavailable during periods of congestion. Paid consumer access can remove some constraints while remaining materially below enterprise access in throughput, governance, data integration and guaranteed availability. Local high-performance access can provide privacy and experimentation but requires substantial capital, electricity, technical competence and maintenance. The correct analytical variable is therefore quality-adjusted effective access, not account ownership.

This distinction matters because AI can augment activities that accumulate human and institutional capital: studying, coding, translating, publishing, analysing documents, preparing applications, designing businesses, conducting experiments and searching scientific literature. Small initial differences in access may compound. A student who receives higher-quality explanations, faster feedback and stronger coding assistance can complete more projects; those projects can improve employability; higher income can finance better AI access; better access can then produce further advantage. The opposite sequence can affect students and researchers confined to outdated hardware, unstable free tiers or low-capability local models.

Adoption is already strongly stratified. In 2025, 20.0% of EU enterprises with at least ten employees reported using AI, but the rate among large enterprises reached 55.03%. Use of Artificial Intelligence in Enterprises – Eurostat – December/2025Eurostat enterprise AI statistics. Although enterprise statistics do not directly measure student access, they reveal the institutional environments into which graduates enter: researchers and workers associated with wealthy organisations can obtain better models, data, tools and compute than equally capable individuals outside them.


7.2 From digital access to capability access

7.2.1 Four access groups

GroupTypical accessEffective capabilityPrincipal limitation
G0 — No regular AI accessNo stable account, device, connectivity or permissionNone or occasional indirect assistanceComplete exclusion
G1 — Free-tier accessFree consumer model on a phone or ordinary computerBasic explanation, drafting, translation and limited codingRate limits, model routing, weak continuity and uncertain availability
G2 — Paid consumer accessMonthly subscription and adequate general-purpose deviceStronger reasoning, files, longer context and more toolsFair-use limits, no guaranteed capacity and limited institutional integration
G3 — Frontier enterprise or high-performance local accessFrontier APIs, enterprise plans, specialised agents, proprietary data or high-memory local hardwareSustained research, coding, long-document processing, automation and confidential workflowsHigh monetary, technical and organisational cost

The groups should not be treated as permanent identities. A person may move between them by task: free consumer access for general questions, university enterprise access for research and no permissible AI access for confidential clinical data. Measurement should therefore record the percentage of relevant study or work time during which each access class is available.

7.2.2 Quality-adjusted AI access

Define quality-adjusted access as:

QAAu,t = Aavailability × Qtask × Uallowance × Ttools × Ddata × Sskills

where:

  • Aavailability is the probability that the system is accessible when required;
  • Qtask is task-specific quality;
  • Uallowance is usable volume relative to demand;
  • Ttools represents access to retrieval, code execution, file analysis and external tools;
  • Ddata represents lawful access to relevant data;
  • Sskills represents user competence.

Because the variables have multiplicative interaction, one near-zero component can neutralise the others. A highly capable model provides little effective access if the user can submit only a few requests, cannot upload required documents, lacks connectivity or cannot evaluate unreliable output.


7.3 Dimensions of the AI capability divide

7.3.1 Educational attainment

AI can affect educational attainment through:

  • personalised explanations;
  • repeated feedback without additional tutoring fees;
  • adaptive practice;
  • translation;
  • accessibility assistance;
  • coding and mathematical support;
  • summarisation and note generation;
  • exam preparation;
  • administrative guidance;
  • application and scholarship support.

The causal estimand is not the difference between AI users and non-users, because users may differ in prior attainment, income, motivation and institutional support. A valid model is:

Yi,t = α + βQAAi,t + γYi,t−1 + δXi,t + μinstitution + τt + εi,t

where:

  • Yi,t is attainment;
  • QAAi,t is quality-adjusted AI access;
  • Xi,t contains income, prior grades, language, disability, device and connectivity controls;
  • μinstitution captures institutional differences;
  • τt captures period effects.

A positive β would support the access-advantage hypothesis only if AI access precedes the measured outcome and confounding is credibly controlled.

7.3.2 Research productivity

Research benefits may include:

  • literature discovery;
  • document classification;
  • code generation;
  • data cleaning;
  • translation;
  • hypothesis generation;
  • simulation support;
  • manuscript revision;
  • grant preparation;
  • administrative automation.

Research productivity should be measured through multiple outcomes:

RPi,t = w1Publications + w2Citations + w3Datasets + w4Software + w5Replication + w6ResearchQuality

Raw publication count is insufficient because AI can increase text production without improving scientific validity. Retractions, irreproducible analyses, fabricated references and homogeneous research agendas are negative outputs, not productivity.

7.3.3 Employability

AI affects employability through two opposing mechanisms. It can increase a worker’s ability to code, communicate, analyse and automate. It can also reduce the market value of routine skills. The employability advantage is therefore:

EAi = Productivity Complementarity + AI Literacy + Portfolio Quality − Skill Substitution − Verification Failure

A worker becomes more employable when AI complements scarce domain expertise. A worker whose output is easily substituted may face greater competition even after acquiring basic AI skills.

7.3.4 Entrepreneurship

Entrepreneurial opportunity can expand because AI reduces the cost of:

  • market research;
  • translation;
  • software prototyping;
  • design;
  • customer service;
  • legal-document preparation;
  • marketing;
  • data analysis;
  • business planning.

However, frontier access can create a new minimum efficient scale. If competitors use integrated agents, proprietary data and reserved inference capacity, free-tier users may be able to create prototypes but not reliable production systems.

7.3.5 Software development

Coding is among the most measurable domains because performance can be evaluated through tests, issue resolution, defect rate and accepted code. Appropriate outcomes include:

  • time to accepted pull request;
  • percentage of issues resolved;
  • unit-test pass rate;
  • security defects;
  • maintainability;
  • human review time;
  • cost per successfully completed issue.

Anthropic reported a 72.5% result for Claude Opus 4 on SWE-bench under its stated evaluation configuration. Introducing Claude 4 – Anthropic – May/2025Anthropic Claude 4 evaluation disclosure. This is evidence that frontier systems can resolve a substantial share of benchmarked software-engineering tasks under specified conditions. It is not proof that every enterprise codebase will obtain the same result.

7.3.6 Language access

Multilingual AI can reduce the penalty borne by people working outside dominant scientific and commercial languages. It can translate instructions, research, technical documentation and professional correspondence. Capability remains uneven across languages, dialects and specialist terminology.

The language-access advantage can be represented as:

LAl = Qtranslation,l × Qreasoning,l × Coveragel − ErrorCostl

A system that translates fluently but reasons less reliably in the target language may create false inclusion. Anthropic’s Sonnet 4.6 system card reports evaluation across 57 academic subjects and 14 non-English languages and publishes the model’s MMMLU methodology and results. Claude Sonnet 4.6 System Card – Anthropic – February/2026Anthropic Claude Sonnet 4.6 System Card. Such results remain provider-reported and should be independently reproduced before being used for policy conclusions.


7.4 AI Affordability Index

7.4.1 Definition

AAIc,t = Annual Cost of Adequate AI Accessc,t ÷ Median Disposable Incomec,t

The numerator must represent adequate—not merely nominal—access. It should include:

Costadequate = Subscription + API + DeviceAnnualisation + Connectivity + Storage + Software + Maintenance + Tax

An affordability index based solely on a USD 20 monthly subscription understates the cost for users who need a capable computer, broadband, storage or additional API capacity.

7.4.2 Proposed interpretation

AAI valueInterpretation
Below 1%Broadly affordable for median-income users
1–3%Noticeable but generally manageable
3–5%Material household decision
5–10%Restrictive
10–20%Severely restrictive
Above 20%Economically exclusionary without subsidy

These thresholds are analytical classifications, not international standards.

7.4.3 Adequate-access baskets

BasketComponentsAnnual scenario cost
B0 — Nominal free accessExisting device, free tier, existing connectivityUSD 0 incremental, but no adequacy guarantee
B1 — Basic paid accessConsumer subscription, ordinary device annualisation, incremental connectivityUSD 660
B2 — Professional frontier accessHigher usage, frontier subscription/API, storage and capable deviceUSD 2,400
B3 — Advanced research accessHigh API use, tools, storage and workstation annualisationUSD 4,000
B4 — High-performance private accessLocal workstation/server, maintenance, electricity and model operationsUSD 8,000–25,000

The baskets are standardised analytical scenarios. Taxes, exchange rates and local hardware prices must be applied separately.


7.4.4 Income-level affordability scenarios

Median annual disposable resourcesB1: USD 660B2: USD 2,400B3: USD 4,000B4 low: USD 8,000
USD 3,00022.0%80.0%133.3%266.7%
USD 6,00011.0%40.0%66.7%133.3%
USD 12,0005.5%20.0%33.3%66.7%
USD 24,0002.75%10.0%16.7%33.3%
USD 45,0001.47%5.33%8.89%17.78%
USD 75,0000.88%3.20%5.33%10.67%

Eurostat reported median annual disposable income of 21,245 purchasing-power standards per EU inhabitant in 2024. Living Conditions in Europe: Income Distribution and Income Inequality – Eurostat – 2026Eurostat median disposable-income statistics. At an illustrative income of 21,245 units, the standardised B1, B2 and B3 baskets would absorb approximately 3.1%, 11.3% and 18.8% respectively if the basket units were price-level comparable. A rigorous country table must convert the basket into local prices rather than equating one purchasing-power standard mechanically with one US dollar.


7.5 Student AI Affordability Index

7.5.1 Definition

SAIc,t = Annual AI and Hardware Costc,t ÷ Annual Student Disposable Resourcesc,t

Student resources should include income actually available after tuition and essential housing:

Student Resources = Grants + Scholarships + Family Support + Net Employment Income − Tuition − Essential Housing − Essential Living Costs

Using household median income as the denominator can severely understate student hardship.

7.5.2 Student affordability scenarios

Annual disposable student resourcesBasic paid access: USD 660Professional access: USD 2,400Advanced research: USD 4,000
USD 1,50044.0%160.0%266.7%
USD 3,00022.0%80.0%133.3%
USD 6,00011.0%40.0%66.7%
USD 12,0005.5%20.0%33.3%
USD 24,0002.75%10.0%16.7%

This explains why a service that appears inexpensive to a professional can be exclusionary for a student. A USD 20 monthly subscription equals USD 240 annually before hardware and connectivity. For a student with only USD 1,500 of genuinely disposable resources, the subscription alone consumes 16%.


7.6 Workstation affordability

7.6.1 Standardised workstation

For cross-country modelling, define a reference AI workstation as:

  • modern multi-core CPU;
  • 128 GB system memory;
  • 24–32 GB accelerator memory;
  • 2–4 TB solid-state storage;
  • adequate power supply and cooling;
  • five-year economic life.

The reference planning price is set at:

Preference = USD 4,000 before local tax, tariff and currency adjustment

This is an analytical basket rather than a named retail product. As an official hardware anchor, NVIDIA lists the GeForce RTX 5090 with 32 GB GDDR7 memory and a starting price of USD 1,999 before the remainder of the workstation. GeForce RTX 5090 Specifications – NVIDIA – January/2025NVIDIA GeForce RTX 5090 official specifications.

7.6.2 Wage-month formula

Wage Monthsc = Local Workstation Pricec ÷ Median Monthly Wagec

The ratio must use local after-tax wages for an affordability study. Gross earnings can be reported as a secondary indicator but should not be confused with disposable income.

7.6.3 Verified and indicative calculations

CountryOfficial earnings inputWorkstation basketWage monthsEvidentiary qualification
United StatesMedian USD 1,251/week; approximately USD 5,421/monthUSD 4,0000.74Median gross full-time earnings, Q2 2026
GermanyImplied median EUR 21.48/hour from official low-wage threshold; approximately EUR 3,723/month at 173.3 hoursEUR 4,0001.07Estimated monthly median from official hourly median
ItalyAverage hourly earnings approximately EUR 16.35 in the reported 2022 structure survey; approximately EUR 2,834/month at 173.3 hoursEUR 4,0001.41Average, not median; indicative only
IndiaAverage employee earnings INR 22,220/month in 2025INR 350,000 analytical local basket15.75Average, not median; basket requires local price verification
South AfricaAverage employee earnings ZAR 7,980/month in the latest ILO profile observation, dated 2020ZAR 80,000 analytical basket10.03Stale average; unsuitable for current ranking

US median full-time earnings were USD 1,251 per week in Q2 2026. Usual Weekly Earnings of Wage and Salary Workers – U.S. Bureau of Labor Statistics – July/2026BLS Q2 2026 median earnings. Germany’s Federal Statistical Office reported a median hourly earnings benchmark of EUR 21.48 for April 2025 through its official low-wage-threshold table. Low-Wage Thresholds in Germany – Destatis – May/2025Destatis German wage thresholds. ILOSTAT reports average monthly employee earnings of INR 22,220 for India in 2025. India Country Profile – ILOSTAT – 2026ILOSTAT India labour profile.

The table must not be interpreted as a definitive cross-country league table because wage concepts, dates, taxes and workstation prices are not perfectly harmonised. Its robust finding is ordinal: hardware representing less than one month of full-time US median earnings can represent many months of employee income in lower-income markets.


7.7 Measuring the capability difference between access groups

7.7.1 Outcome matrix

CapabilityG0: no accessG1: free tierG2: paid consumerG3: frontier/enterprise/local high-performance
Basic explanationNoneModerateHighHigh
Complex reasoningNoneLimited/variableHighHighest available
CodingNoneBasic to moderateStrongStrong with tools and sustained workloads
Factual researchManual onlyVariableBetter tools and contextRetrieval, APIs, databases and auditability
Long documentsManualRestrictedModerate/highHigh throughput and persistent projects
Multilingual supportManual toolsBroad but variableHigher qualityDomain-integrated and scalable
Scientific workflowManualLimitedModerateCode, tools, private data and automation
Confidential workNo AIOften unsuitableContract-dependentEnterprise controls or local privacy
ConcurrencyNoneLowModerateHigh
ReproducibilityManualWeakModerateAPI snapshots, logging and controlled deployment
AdaptationNoneMinimalUser-levelRAG, fine-tuning, agents and local models

The table expresses expected structural access rather than fixed model performance. Free tiers can temporarily expose frontier models, while paid users may select weaker models. The classification must therefore be recorded at the level of the actual model, service tier and task.


7.8 Reproducible capability benchmark

7.8.1 Test principle

Parameter count must not be used as a proxy for utility. Performance depends on:

  • architecture;
  • total and active parameters;
  • training data;
  • post-training;
  • quantisation;
  • context;
  • retrieval;
  • tools;
  • inference framework;
  • prompt format;
  • sampling;
  • language;
  • task specialisation;
  • compute budget.

The benchmark unit should be:

Successful Workloads ÷ Total Cost

rather than raw benchmark score alone.

7.8.2 Required benchmark suite

DimensionTest constructionPrimary metricFailure metric
ReasoningNovel multi-step problems with contamination controlsExact success rateUnsupported intermediate claims
CodingRepository issues with testsResolved issue rateNew defects and insecure code
Factual reliabilityTime-stamped questions with verified corpusSupported-claim precisionHallucination rate
MultilingualParallel domain tasks across languagesQuality parity ratioPerformance degradation by language
Long contextMultiple relevant and adversarial passagesRetrieval and synthesis accuracyContext-position failure
Document processingTables, scans, footnotes and conflicting passagesExtraction F1 and citation accuracyInvented fields
Scientific tasksReproducible calculations and literature analysisCorrect result and provenanceFabricated references
EnergyWall-socket or facility energyJoules per successful taskEnergy per failed output
CostFull provider or local TCOCost per successful taskCost per generated token without quality
PrivacyData-flow and retention auditCompliance rateUnauthorised transmission

7.8.3 Controlled experimental configuration

Every run must disclose:

  • exact model and version;
  • quantisation;
  • hardware;
  • inference framework;
  • prompt;
  • temperature;
  • reasoning effort;
  • tool access;
  • retrieval corpus;
  • context length;
  • maximum output;
  • number of trials;
  • confidence interval;
  • date;
  • total cost;
  • energy measurement.

OpenAI’s official model guidance recommends comparing configurations on representative tasks instead of assuming that the highest reasoning effort always provides the best trade-off. Model Guidance – OpenAI API – August/2026Official OpenAI model-evaluation guidance.

7.8.4 Statistical analysis

For model m and task k:

Successm,k ∈ {0,1}

Mean success is:

m = ΣSuccessm,k ÷ n

The quality-adjusted cost is:

QACm = Total Costm ÷ ΣSuccessm,k

Energy per successful task is:

Esuccess,m = Total Energym ÷ ΣSuccessm,k

Hallucination rate is:

HRm = Unsupported Factual Claimsm ÷ All Factual Claimsm

Multilingual parity for language l is:

MPm,l = Scorem,l ÷ Scorem,reference

Values below one indicate lower performance than the reference language.


7.9 Do smaller local models create a measurable disadvantage?

The hypothesis is:

H7.1: After controlling for tools, task specialisation, quantisation, context and cost, smaller local models produce a lower successful-workload rate on complex general tasks than frontier cloud models.

The counter-hypothesis is:

H7.2: On bounded and specialised tasks, a smaller local model combined with retrieval can equal or exceed a general frontier model at lower cost and higher privacy.

Both can be true.

7.9.1 Expected task pattern

TaskLikely frontier advantagePotential small-model advantage
Novel multi-step reasoningHighLimited
Complex repository codingHighSpecialised code model may compete
Routine classificationLowStrong cost and latency advantage
Structured extractionLow/moderateStrong with schema and validation
Multilingual specialist workLanguage-dependentStrong only after targeted adaptation
Long-document synthesisHigh without retrievalRAG can reduce gap
Private document processingCloud quality may be higherLocal privacy and control
Scientific reasoningHighDomain-tuned model may compete narrowly
High-concurrency automationInfrastructure-dependentSmaller models may scale more economically

Google’s Gemini 3.1 Pro model card reports evaluation across reasoning, multimodal, multilingual, agentic and long-context tasks. Gemini 3.1 Pro Model Card – Google DeepMind – February/2026Google DeepMind Gemini 3.1 Pro Model Card. Such model cards demonstrate that capability is multidimensional. They do not provide a substitute for an independently controlled local-versus-cloud experiment.

7.9.2 Quantisation controls

A valid local comparison should test at least:

  • BF16 or FP16;
  • INT8;
  • INT4;
  • selected lower-bit method if documented.

The quantisation penalty is:

QPm,q,k = Scorem,reference,k − Scorem,q,k

A lower-bit model should be considered economically superior only when:

QACm,q < QACm,reference

and:

Scorem,q,k ≥ Minimum Acceptable Scorek

A model that becomes cheaper but falls below the task’s minimum quality threshold is not an adequate substitute.


7.10 Scientific-publication and research divide

Research institutions with enterprise AI access can provide:

  • bulk document ingestion;
  • licensed databases;
  • secure code execution;
  • private datasets;
  • long-context models;
  • shared prompt libraries;
  • evaluation support;
  • dedicated local compute;
  • legal and research-integrity guidance.

An unaffiliated researcher may have only a consumer subscription and personal computer. The difference is not merely model quality. It includes data rights, concurrency, reproducibility, storage, collaboration and the ability to process sensitive material.

Define the Research AI Resource Index:

RARIi = w1Compute + w2ModelQuality + w3DataAccess + w4ToolAccess + w5Support + w6UsageAllowance

A publication model can be estimated as:

Publicationsi,t+1 = α + βRARIi,t + γPriorPublicationsi,t + δFundingi,t + μfield + εi,t

Citation and quality measures should be analysed separately to detect whether AI increases quantity without validity.


7.11 Cumulative advantage

7.11.1 Requested model

At+1 = At + αQt − βCt

where:

  • At = accumulated human or institutional advantage;
  • Qt = quality-adjusted AI access;
  • Ct = access and adaptation constraints;
  • α = conversion of AI access into advantage;
  • β = damage caused by constraints.

After n periods:

An = A0 + Σt=0n−1[αQt − βCt]

The access threshold required merely to avoid declining relative advantage is:

Qt* = βCt ÷ α

If Qt < Qt*, accumulated advantage falls relative to the competitive environment.

7.11.2 Two-user simulation

Assume:

  • α = 0.8;
  • β = 0.5;
  • high-access user: Q = 1.0 and C = 0.2;
  • constrained user: Q = 0.3 and C = 0.8;
  • both begin with A0 = 10.

Annual increments are:

ΔAhigh = 0.8 × 1.0 − 0.5 × 0.2 = 0.70

ΔAconstrained = 0.8 × 0.3 − 0.5 × 0.8 = −0.16

YearHigh-access advantageConstrained-user advantageGap
010.0010.000.00
110.709.840.86
211.409.681.72
312.109.522.58
412.809.363.44
513.509.204.30

Even with equal starting capability, persistent access differences create divergence.

7.11.3 Feedback mechanism

In reality, accumulated advantage can finance future access:

Qt+1 = q0 + λAt − φPt

where:

  • λ is the degree to which advantage improves future access;
  • Pt is the effective price;
  • φ is price sensitivity.

Substitution produces:

At+1 = At + α[q0 + λAt − φPt] − βCt

At+1 = (1 + αλ)At + αq0 − αφPt − βCt

When αλ is positive, advantage compounds rather than accumulating linearly.


7.12 The intelligence poverty trap

An intelligence poverty trap exists when low income or institutional resources cause weak AI access, weak access suppresses productivity and opportunity, and reduced opportunity prevents investment in better access.

The cycle is:

Low Resources → Low-Quality Access → Lower Productivity → Lower Income or Funding → Low Resources

The trap condition can be formalised as:

Return from AI Accesslow-resource < Cost of Adequate Access

while:

Return from AI Accesshigh-resource > Cost of Adequate Access

This can occur even when the same nominal subscription price applies to both groups because:

  • the price consumes a larger income share;
  • poorer users have weaker hardware;
  • connectivity is less reliable;
  • free time for adaptation is limited;
  • schools and firms provide less support;
  • users cannot absorb experimental failure;
  • local language performance may be lower;
  • payment systems or regional availability may restrict access.

7.12.1 Trap threshold

Let income evolve as:

Yt+1 = Yt + θAt

and access be:

Qt = max[0, q − P ÷ Yt]

A low-income equilibrium exists when P ÷ Yt is large enough to suppress Q, preventing A and Y from rising. A subsidy lowering P, a university providing shared access or a specialised low-cost model increasing q can move the user above the threshold.


7.13 Interventions and their economic logic

InterventionMechanismPrincipal risk
Student AI vouchersReduces subscription costBenefits vendor without ensuring quality
University enterprise accessPools procurement and governanceInstitutional lock-in
Public compute centresProvides local capacityUnderutilisation and maintenance
Cooperative research infrastructureShares fixed costComplex allocation and confidentiality
Open-weight model supportReduces supplier dependenceTechnical skill burden
Device grantsReduces hardware barrierRapid obsolescence
Multilingual evaluationExposes language gapsBenchmark may not reflect real use
AI-literacy programmesRaises conversion coefficient αTraining without continued access has low value
Research API creditsSupports reproducible experimentationCredits expire or favour certain providers
Public-interest model routingChooses cheapest adequate modelRequires independent evaluation
Offline and low-bandwidth systemsReduces connectivity barrierCapability may remain below frontier
Outcome-based grantsFunds demonstrated research valueCan disadvantage exploratory work

The optimal policy does not give every user the most expensive frontier model for every task. It ensures access to the lowest-cost system that meets the task’s quality threshold, together with escalation to frontier capacity when smaller systems fail.


7.14 Research design for causal measurement

7.14.1 Randomised access study

Participants should be randomly assigned to:

  • no AI;
  • free-tier AI;
  • paid consumer AI;
  • frontier enterprise AI.

All groups receive the same tasks, time windows and initial training except where training itself is part of the treatment. Outcomes include:

  • test score;
  • completion time;
  • retention after four weeks;
  • error rate;
  • transfer to new tasks;
  • confidence calibration;
  • dependency;
  • cost;
  • energy.

The treatment effect is:

ATE = E[Y|Access = j] − E[Y|Access = 0]

7.14.2 Longitudinal study

A longitudinal panel should track:

  • prior attainment;
  • socioeconomic status;
  • AI access history;
  • model tier;
  • actual usage;
  • grades;
  • publications;
  • employment;
  • income;
  • entrepreneurship;
  • skills.

A fixed-effects specification is:

Yi,t = βQAAi,t + γXi,t + μi + τt + εi,t

Individual fixed effects μi control for stable unobserved differences.

7.14.3 Institutional natural experiments

Possible quasi-experiments include:

  • staggered university licence rollout;
  • temporary free-credit programmes;
  • regional connectivity changes;
  • laboratory hardware grants;
  • service-tier changes;
  • model withdrawal;
  • price increases.

A difference-in-differences estimator is:

Effect = [Ytreated,after − Ytreated,before] − [Ycontrol,after − Ycontrol,before]

Parallel pre-treatment trends must be tested.


7.15 Five-year scenarios, 2027–2031

ScenarioAccess developmentEducational and social effect
Broad diffusionCapable small models and institutional access expandDivide narrows for routine tasks
Tiered intelligenceFree models improve, but frontier models advance fasterNominal access broadens while capability gap persists
Subscription escalationHigh-end access becomes more expensive and meteredStudents and independent researchers lose relative access
Public-infrastructure responseUniversities and governments provide shared capacityGeographic divide narrows where institutions are effective
Sovereign fragmentationModel and hardware access differs by regionNationality and location become capability determinants
Agentic concentrationAdvanced agents require enterprise data and toolsOrganisational affiliation dominates individual talent

The central scenario is tiered intelligence. Basic AI becomes ubiquitous, but frontier-quality reasoning, long-context analysis, private data integration, sustained agentic work and high-volume usage remain concentrated among wealthy individuals and institutions. The resulting divide is less visible than total exclusion because both rich and poor users can claim to “have AI,” while the quality and productive capacity of that access differ substantially.


7.16 Falsifiable hypotheses

HypothesisTestable propositionRequired evidence
H7.1Paid AI access improves task completion relative to no accessRandomised user study
H7.2Frontier access produces greater gains than free-tier access on complex tasksMulti-treatment experiment
H7.3Smaller local models underperform frontier models on general complex tasksControlled benchmark
H7.4Specialised small models can match frontier systems on bounded tasksTask-specific quality-adjusted cost study
H7.5AI access increases research outputLongitudinal researcher panel
H7.6AI access increases publication quantity more than qualityPublication, citation and replication analysis
H7.7Multilingual capability gaps reduce gains outside dominant languagesParallel multilingual experiment
H7.8AI affordability is negatively associated with national incomeCountry-level AAI panel
H7.9Access effects compound over timeLongitudinal cumulative-advantage model
H7.10Institutional access mediates the effect of personal incomeMultilevel student and researcher model
H7.11High AI prices reduce entrepreneurial entryRegional or cohort study
H7.12Free-tier access does not eliminate the frontier capability divideQuality-adjusted access comparison

7.17 Final findings

  1. AI access is not binary. No access, free access, paid consumer access and frontier institutional access represent materially different productive environments.
  2. Affordability must include hardware, connectivity, storage and usage. Subscription price alone substantially understates adequate-access cost.
  3. Student affordability is structurally worse than household affordability. Students often possess far less disposable income than the national median.
  4. An AI-capable workstation can cost less than one median monthly wage in a high-income country but many months of earnings in lower-income markets. Tariffs and local scarcity can widen the difference.
  5. Smaller local models do not always create a disadvantage. They can be economically superior for bounded, specialised and privacy-sensitive tasks.
  6. Frontier models retain likely advantages in novel reasoning, difficult coding, long-context synthesis and multidisciplinary scientific work.
  7. Parameter count is not an adequate utility measure. Architecture, post-training, quantisation, tools, retrieval, context and task design must be controlled.
  8. Benchmark cost must be divided by successful workloads, not generated tokens. Cheap incorrect output is not affordable intelligence.
  9. AI can improve language access while preserving hidden quality disparities. Fluency does not guarantee equivalent reasoning or factual reliability.
  10. Institutional affiliation is becoming a determinant of effective capability. Researchers inside wealthy universities or companies can access data, tools, compute and support unavailable to independent peers.
  11. Cumulative advantage is mathematically plausible. Persistent differences in quality-adjusted access generate widening human-capital and institutional gaps even when starting ability is equal.
  12. A self-reinforcing intelligence poverty trap is possible. Low resources reduce access; weak access reduces productivity and opportunity; reduced opportunity prevents future investment.
  13. Universal free access may not eliminate the divide. If frontier systems improve faster than free systems, nominal inclusion can coexist with growing capability inequality.
  14. The correct policy target is adequate task-specific access. Public support should supply the least-cost system meeting validated quality thresholds, with escalation to frontier capacity where necessary.
  15. Between 2027 and 2031, social mobility may increasingly depend on access not merely to AI, but to sufficiently capable, sustained, verifiable and economically usable AI.

CHAPTER 8 — Countries, Development and Global AI Inequality

8.1 Scope, purpose and evidentiary limits

Global AI inequality cannot be measured through the number of chatbot accounts, national AI strategies or data centres alone. Effective access depends on a connected system of household purchasing power, hardware affordability, electricity, broadband, cloud infrastructure, advanced semiconductors, foreign-currency availability, digital skills, local-language capability, research institutions, regulation and geopolitical permission to acquire technology. A country may possess strong mobile connectivity but little research compute; inexpensive electricity but an unreliable grid; cloud access but no domestic region; capable engineers but severe foreign-currency constraints; or extensive infrastructure whose benefits remain concentrated in a small urban elite. A multidimensional index must capture these complementarities without concealing them inside an opaque ranking.

The World Bank’s 2025 framework identifies four foundations of AI development: connectivity, compute, context and competency. It concludes that high-income countries continue to dominate AI innovation, compute infrastructure and startup funding, while low- and middle-income economies face persistent gaps in connectivity, locally relevant data, skills and compute. It also identifies “Small AI”—lower-cost, task-specific systems capable of running on ordinary devices—as a potential route to wider diffusion. Digital Progress and Trends Report 2025: Strengthening AI Foundations – World Bank – November/2025World Bank AI foundations report.

This chapter constructs a Global AI Access Index, or GAI, and applies it to a pilot set of 25 country and regional cases. The numerical scores are transparent analytical estimates constructed for scenario analysis; they are not an official World Bank, ITU or national-government ranking. The index must be recalculated from a frozen country-year dataset before publication as a definitive statistical table. This qualification is necessary because several required variables—particularly local-language model quality, latency to AI-capable cloud capacity, research-compute availability and hardware street prices—are not yet available through a single harmonised international dataset.


8.2 Global connectivity and compute baseline

Almost three-quarters of the global population was online in 2025, but approximately 2.2 billion people remained offline. Internet use reached 94% in high-income economies but only 23% in low-income economies. Approximately 96% of offline people lived in low- and middle-income countries. Urban internet use reached 85%, compared with 58% in rural areas, while a typical user in a high-income country generated nearly eight times as much mobile data as a user in a low-income economy. Measuring Digital Development: Facts and Figures 2025 – International Telecommunication Union – November/2025ITU Facts and Figures 2025.

These differences matter because nominal coverage does not establish usable AI connectivity. AI use can require:

  • sustained rather than intermittent service;
  • sufficiently low latency;
  • adequate upload capacity for documents and images;
  • affordable data volume;
  • secure access;
  • compatible devices;
  • reliable payment mechanisms;
  • cloud services available in the jurisdiction;
  • continuous electricity.

ITU’s concept of universal and meaningful connectivity includes quality, availability, affordability, devices, skills and security. Global Connectivity Report 2025 – International Telecommunication Union – 2025ITU Global Connectivity Report 2025. This is more appropriate to AI-access analysis than a binary connected/not-connected variable.


8.3 Country and regional comparison

8.3.1 Structural comparison

Country or groupPrincipal strengthsPrincipal constraintsStrategic position
United StatesFrontier providers, hyperscale cloud, accelerators, capital, universitiesRegional inequality, high enterprise concentration, energy bottlenecksFrontier producer and dominant platform jurisdiction
CanadaResearch institutions, cloud access, skills, electricitySmaller domestic market and dependence on foreign platformsAdvanced adopter with selected research strengths
European UnionIndustrial base, research, regulation, multiple cloud regionsFragmentation, limited frontier-platform ownership, energy costLarge regulated market with incomplete compute sovereignty
United KingdomStrong research, finance, cloud access and language advantageDependence on imported hardware and foreign providersAdvanced service and research hub
ChinaLarge market, data centres, domestic models and industrial policyExport controls, leading-edge hardware constraintsParallel ecosystem with substantial scale
JapanHigh income, advanced industry, stable electricity and researchAgeing workforce, imported accelerator dependenceAdvanced adopter and industrial AI power
South KoreaSemiconductors, broadband, cloud, digital skillsImported frontier accelerators and geopolitical exposureHigh-readiness technology producer
IsraelAdvanced research, cybersecurity, startups and cloud investmentSmall market, geopolitical risk and import dependenceHigh-capability specialist ecosystem
IndiaLarge technical workforce, cloud regions, software sector and scaleIncome, electricity and rural connectivity inequalityRapid adopter with large internal divide
BrazilLarge market, cloud region, research and Portuguese-language scaleHardware cost, taxation, income inequality and currency riskRegional leader with affordability constraints
MexicoManufacturing integration, proximity to US, growing cloud accessResearch-compute and skills inequalityNearshore AI adopter
CroatiaEU membership, connectivity and regulatory integrationSmall market and limited research computeHigher-readiness Balkan adopter
SerbiaTechnical workforce and regional connectivityNon-EU position, smaller compute base and currency exposureEmerging regional hub
AlbaniaImproving digital government and connectivityLimited research capacity and incomeService-led emerging adopter
Bosnia and HerzegovinaEducated diaspora and regional accessInstitutional fragmentation and limited investmentCapacity-constrained adopter
MoroccoCloud and data-centre investment potential, multilingual workforceHardware affordability and research scaleLeading North African services platform
TunisiaTechnical education and nearshore potentialFinancing, currency and infrastructure constraintsTalent-rich but capital-constrained
EgyptLarge market, cables, Arabic-language demandIncome, currency and infrastructure pressurePotential regional platform with affordability barriers
South AfricaStrongest Sub-Saharan cloud footprint and research baseElectricity reliability, inequality and high hardware costContinental infrastructure leader with internal divide
KenyaDigital services, entrepreneurship and regional connectivityResearch compute, electricity and income constraintsEast African digital-services hub
NigeriaLarge population and entrepreneurshipElectricity reliability, foreign currency and infrastructureLarge demand base with severe access constraints
GhanaInstitutional stability and regional digital-services potentialSmall compute base and affordabilityEmerging public-service adopter
RwandaDigital-government orientation and policy coordinationSmall market and imported infrastructureInstitutionally agile, compute-constrained
EthiopiaLarge population and development potentialConnectivity, electricity, foreign currency and skills constraintsEarly-stage AI-access environment
BangladeshLarge workforce and digital-services potentialEnergy, hardware affordability and research-compute limitationsLabour-scale adopter with constrained infrastructure

8.4 Global AI Access Index specification

8.4.1 Core equation

GAIi = Σj=1m wjzij

where:

  • GAIi is the Global AI Access Index for country i;
  • zij is the normalised value of indicator j;
  • wj is the indicator weight;
  • Σwj = 1.

The index is scaled from 0 to 100:

GAIi,100 = 100 × GAIi

A high score represents greater affordable access, infrastructure readiness and capability.


8.2 Indicator architecture

IndicatorSymbolPreferred empirical measureDirection
Household incomeINCMedian disposable income per capita, PPP-adjustedPositive
Hardware affordabilityHWAReference workstation cost divided by median incomeNegative
Import dutiesDUTEffective tariff and tax burden on reference hardwareNegative
Electricity reliabilityRELOutage frequency/duration or reliable-access measurePositive
Electricity costEPCCommercial/household price per kWhNegative
Broadband availabilityBBAFixed/mobile meaningful-connectivity measurePositive
Data-centre latencyLATMedian latency to nearest AI-capable cloud regionNegative
Cloud-region availabilityCRANumber and capability of domestic/nearby hyperscale regionsPositive
Local-language qualityLLQReproducible model-quality parity across national languagesPositive
Digital skillsDSKAdvanced ICT-skill rate or harmonised proxyPositive
Research computeRCOQuality-adjusted public research accelerator capacityPositive
Foreign-currency accessFXAConvertibility and ability to finance technology importsPositive
Sanctions/export exposureSEELegal and practical restriction on advanced hardware/servicesNegative
Local data-centre capacityDCCOperational IT capacity per capita or GDPPositive
Regulatory environmentREGPredictability, data governance and competition balancePositive

The ITU notes that comparable ICT-skills data remain scarce: only 90 countries had submitted relevant data since 2020, and only about 40 provided broadly comparable skill-level information. ITU ICT SDG Indicators – International Telecommunication Union – 2026ITU ICT-skills data limitations. Missing data must therefore be disclosed and not silently replaced by subjective assumptions.


8.5 Normalisation

8.5.1 Min-max normalisation

For a positive indicator:

zij = [xij − min(xj)] ÷ [max(xj) − min(xj)]

For a negative indicator:

zij = [max(xj) − xij] ÷ [max(xj) − min(xj)]

Min-max normalisation is intuitive but sensitive to outliers.

8.5.2 Winsorised min-max

Values below the fifth percentile are replaced with the fifth-percentile value; those above the ninety-fifth percentile are replaced with the ninety-fifth-percentile value:

xijW = min[max(xij, P5,j), P95,j]

The min-max calculation is then applied to xijW.

8.5.3 Z-score normalisation

zij = [xij − μj] ÷ σj

Z-scores preserve distance from the mean but can produce negative index components. A cumulative normal transformation can map them to 0–1.

8.5.4 Rank normalisation

zij = [rank(xij) − 1] ÷ (n − 1)

Rank normalisation reduces outlier influence but discards cardinal differences. A country narrowly above another receives the same rank interval as a country separated by a very large absolute gap.


8.6 Weighting systems

8.6.1 Equal weights

With 15 indicators:

wj = 1 ÷ 15 = 0.0667

Equal weights are transparent but imply that every indicator has identical conceptual importance.

8.6.2 Expert weights

IndicatorExpert weight
Household income7%
Hardware affordability9%
Import duties4%
Electricity reliability8%
Electricity cost5%
Broadband availability8%
Data-centre latency5%
Cloud-region availability8%
Local-language quality7%
Digital skills8%
Research compute10%
Foreign-currency access4%
Sanctions/export exposure7%
Local data-centre capacity7%
Regulatory environment3%
Total100%

The higher weights on hardware, research compute, broadband and electricity reliability reflect their role as binding complements. Regulation receives a lower direct weight because regulatory quality should not compensate arithmetically for the absence of electricity or compute.

8.6.3 Principal-component-derived weights

PCA weights should be derived from the first components explaining a predetermined share—such as 70%—of standardised variance. If loading ajk represents indicator j on component k and λk its eigenvalue:

wj,PCA = Σk=1K|ajkk ÷ Σj=1mΣk=1K|ajkk

PCA identifies empirical covariance, not moral or developmental importance. It can assign low weight to an essential variable if that variable has little cross-country variance.


8.7 Missing-data protocol

A country should not receive a favourable score because data are unavailable. The protocol should be:

  1. Report the missing indicator.
  2. Use multiple imputation only where predictor relationships are credible.
  3. Calculate an uncertainty interval across imputed datasets.
  4. Exclude countries missing more than 30% of weighted indicators.
  5. Publish a completeness score:

Completenessi = ΣwjI(xij observed)

  1. Report:

GAIadjusted,i = GAIi × Completenessiκ

where κ is a disclosed penalty parameter.

The unpenalised and adjusted scores should both be published.


8.8 Provisional pilot index

The following pilot scores use the expert-weight framework and a 0–100 scale. They represent analytical synthesis from the official structural evidence reviewed in this report, not a frozen final statistical dataset.

Country or regional caseProvisional GAIAccess tierPrincipal binding constraint
United States89FrontierInternal affordability and geographic inequality
South Korea84AdvancedImported frontier accelerator exposure
Israel83AdvancedScale and geopolitical exposure
Canada82AdvancedForeign-platform dependence
United Kingdom82AdvancedImported hardware and platform dependence
European Union aggregate80AdvancedInternal fragmentation and frontier-provider deficit
Japan80AdvancedImported accelerator dependence
China78Advanced/parallelExport controls and leading-edge hardware
India58Emerging-scaleIncome and internal infrastructure inequality
Croatia58Emerging/advancedSmall research-compute base
Brazil55Emerging-scaleHardware affordability and taxation
South Africa52Emerging regional hubElectricity reliability and inequality
Mexico51EmergingResearch compute and skills dispersion
Morocco50Emerging regional hubResearch scale and hardware affordability
Serbia49EmergingCloud depth, market size and currency exposure
Egypt45Constrained scaleCurrency, income and infrastructure pressure
Tunisia43Constrained talent hubFinancing and foreign-currency access
Kenya43Emerging digital hubResearch compute and household affordability
Albania42EmergingSmall market and limited compute
Rwanda40Policy-led emergingScale and imported infrastructure
Bosnia and Herzegovina38ConstrainedInstitutional fragmentation and investment
Ghana38Constrained emergingCompute and affordability
Nigeria35High-potential constrainedElectricity, currency and infrastructure
Bangladesh32Constrained scaleCompute, energy and affordability
Ethiopia22Severely constrainedConnectivity, electricity and foreign currency

An aggregate EU score conceals major internal variation. Countries with domestic cloud regions, high incomes and strong research systems will score materially above lower-income member states with smaller compute ecosystems. The same applies to the United States, China, India, Brazil, South Africa and Nigeria: national averages conceal extreme urban-rural and income-based divergence.


8.9 Sensitivity analysis

8.9.1 Alternative weighting effects

Country groupEqual weightsExpert weightsExpected PCA effect
United States, Canada, UKRemain highRemain highStrong infrastructure-income component sustains rank
EU aggregateHighHighRegulatory component may matter less under PCA
ChinaHighModerately reduced by export exposureIndustrial scale may raise PCA rank
South Korea and JapanHighHighSemiconductor and connectivity strength reinforced
IsraelHighHighSmall market may reduce scale-based PCA score
IndiaMiddleMiddleDigital scale raises rank; income reduces it
Brazil and MexicoMiddleMiddleCloud and market scale offset affordability
BalkansMiddle/lowerHighly country-specificEU membership and connectivity differentiate
North AfricaLower-middleLower-middleCurrency and research compute reduce expert score
Sub-Saharan AfricaLowerLowerConnectivity and income covariance reinforce PCA penalty

8.9.2 Rank robustness standard

For each country:

RankRangei = [min Ranki,s, max Ranki,s]

where s covers:

  • equal weights;
  • expert weights;
  • PCA weights;
  • min-max;
  • winsorised min-max;
  • z-score;
  • rank normalisation;
  • alternative missing-data treatments.

A country should be assigned a robust tier only if it remains in that tier across at least 75% of specifications.

8.9.3 Monte Carlo sensitivity

Weights can be drawn from a Dirichlet distribution centred on expert weights:

w(b) ∼ Dirichlet(κwexpert)

For each of B simulations:

GAIi(b) = Σwj(b)zij

The report should publish:

  • median score;
  • fifth and ninety-fifth percentiles;
  • probability of belonging to each tier;
  • probability of outranking selected peers.

8.10 Inequality measurement

8.10.1 Gini coefficient

For access values xi:

G = [Σi=1nΣj=1n|xi − xj|] ÷ [2n2x̄]

Using the 25 provisional GAI scores:

  • n = 25;
  • mean GAI = 56.36;
  • Gini coefficient = 0.193.

The value indicates substantial but not extreme dispersion among country averages. It understates inequality because:

  • country averages weight Ethiopia and the United States equally rather than by population;
  • within-country inequality is omitted;
  • a score of zero is bounded;
  • access quality may have nonlinear economic returns.

8.10.2 Theil index

The Theil T index is:

T = (1 ÷ n)Σi=1n(xi ÷ x̄)ln(xi ÷ x̄)

For the provisional score distribution:

T = 0.0597

The Theil index is decomposable:

T = Tbetween + Twithin

Dividing the pilot cases into advanced, emerging and constrained groups produces:

ComponentTheil contributionShare
Between-group inequality0.054691.4%
Within-group inequality0.00518.6%
Total0.0597100%

This result reflects the constructed pilot grouping and should not be generalised to the world population. It nevertheless demonstrates that development-tier differences dominate this illustrative country-average distribution.

8.10.3 Percentile ratios

MeasurePilot value
P90/P102.28
P80/P202.03
Maximum/minimum4.05

The maximum pilot score, 89, is slightly more than four times the minimum, 22.

8.10.4 Palma ratio

The conventional Palma ratio is the share held by the top 10% divided by that of the bottom 40%:

Palma = Sharetop10 ÷ Sharebottom40

Applied mechanically to the equally weighted country-score distribution, the pilot Palma ratio is approximately 0.57. A value below one is possible because the bottom 40% contains four times as many country observations as the top 10%, while GAI scores are bounded and much less concentrated than income. For AI analysis, a population-weighted Palma and a compute-capacity Palma would be more informative.


8.11 Population weighting and within-country inequality

A country-level index can conceal more inequality than it reveals. Population-weighted access should use:

population = ΣPixi ÷ ΣPi

Within a country, access can be modelled by household decile, region, gender, rurality, education and institutional affiliation:

GAIi,h = f(Incomeh, Deviceh, Connectivityh, Skillsh, InstitutionalAccessh)

The national score is:

GAIi = Σh=1HphGAIi,h

A high national average can coexist with severe exclusion among rural households, informal workers, minority-language communities and underfunded institutions.


8.12 Cloud geography and latency inequality

Cloud geography matters because local regions can reduce latency, support data residency, improve availability and reduce international bandwidth dependence. The global footprint remains geographically concentrated.

AWS reported 123 Availability Zones across 39 geographic regions, with additional regions planned. AWS Global Infrastructure – Amazon Web Services – August/2026AWS global infrastructure. Microsoft reported more than 80 Azure regions and more than 500 data centres on its current global-infrastructure page. Azure Global Infrastructure – Microsoft – August/2026Microsoft Azure global infrastructure. Google Cloud reported 43 regions and 130 zones across six continents. Global Locations: Regions and Zones – Google Cloud – August/2026Google Cloud global locations.

A domestic region should not be coded simply as one. The cloud-region indicator should measure:

CRAi = Availability × AIServiceDepth × AcceleratorAvailability × Redundancy × Residency

A region lacking frontier accelerators or the required managed AI service is not equivalent to a fully provisioned US region.

Latency should be measured through repeated active tests:

LATi = median RTT from representative population nodes to nearest eligible AI endpoint

Eligibility must include:

  • model availability;
  • legal access;
  • data-residency compliance;
  • sufficient capacity;
  • payment availability.

8.13 Export controls and geopolitical access

Advanced-compute access is partly a legal and geopolitical variable. US controls apply to specified advanced computing semiconductors, semiconductor-manufacturing equipment, high-bandwidth memory and associated entities and destinations. Commerce Strengthens Restrictions on Advanced Computing Semiconductors and Semiconductor Manufacturing Equipment – Bureau of Industry and Security – December/2024US BIS advanced-semiconductor controls.

In January 2026, BIS moved to case-by-case review for specified H200-, MI325X- and similar-class exports to China subject to stated conditions. Department of Commerce Revises License Review Policy for Semiconductors Exported to China – Bureau of Industry and Security – January/2026US BIS China semiconductor licensing policy.

The sanctions/export indicator should distinguish:

Score componentQuestion
Legal eligibilityCan the country legally import the hardware?
Entity eligibilityAre principal firms or universities restricted?
Licence probabilityAre licences routinely approved, conditional or presumptively denied?
Cloud substitutabilityCan equivalent capacity be accessed remotely?
Payment accessCan customers legally and practically pay?
Diversion riskDo compliance concerns discourage suppliers?
Domestic alternativeCan local technology substitute?

China differs fundamentally from a low-income sanctioned economy: it possesses a large domestic industrial, cloud and model ecosystem capable of partial substitution. The same numerical penalty should not be applied without a domestic-capability offset.


8.14 Electricity inequality

AI readiness depends on both electricity access and quality:

EnergyReadinessi = Access × Reliability × Capacity × Affordability × Expandability

A country may report near-universal electricity access while experiencing outages that make continuous local inference unreliable. Another may offer cheap tariffs but lack generation or grid capacity for new data centres. A third may possess reliable electricity whose price makes domestic compute uncompetitive.

For local AI:

Costenergy = PIT × 8,760 × u × PUE × pelectricity

For national AI infrastructure, connection lead time and firm power availability may matter more than average tariff.


8.15 Local-language and cultural representation

8.15.1 Language-quality indicator

Local-language model quality should be measured as:

LLQi = Σl=1Lsl[Scorel ÷ Scorereference]

where sl is the population share using language l.

Tests should cover:

  • factual questions;
  • public-service terminology;
  • legal and medical language;
  • educational content;
  • dialects;
  • culturally specific reasoning;
  • toxicity and refusal;
  • speech recognition;
  • document OCR;
  • translation.

A model may support a language conversationally while performing poorly on technical, scientific or administrative tasks. Countries with large language markets—Chinese, Japanese, Korean, Arabic, Portuguese and Hindi—possess stronger commercial incentives for localisation than small-language communities.

8.15.2 Cultural dependence

Dependence on foreign platforms can shape:

  • which sources are retrieved;
  • which dialects are recognised;
  • moderation boundaries;
  • historical representation;
  • public-service terminology;
  • political and cultural assumptions;
  • data retained outside the country.

Scientific sovereignty therefore includes the ability to evaluate, adapt and govern models—not merely host them.


8.16 Development consequences

8.16.1 Economic development

AI can support development by lowering the cost of knowledge-intensive services, enabling small firms to reach international markets and improving logistics, agriculture, finance and administration. The development contribution can be represented as:

ΔYi = αAi + βKi + γSi + δDi − λMi

where:

  • Ai = AI adoption;
  • Ki = complementary capital;
  • Si = skills;
  • Di = data and institutional quality;
  • Mi = import and monopoly leakage.

If AI services are imported while little domestic value is created, productivity may rise without generating a large domestic AI industry.

8.16.2 Productivity convergence

AI promotes convergence if lower-income countries obtain capability at a lower cost than developing equivalent expertise internally:

Convergence effect > 0 when:

Imported AI Productivity Gain > Subscription Leakage + Displacement + Adaptation Cost

AI promotes divergence when leading countries combine better models with greater capital, skills and data:

Divergence effect > 0 when:

Complementarityfrontier × Existing Capitalrich > Catch-up Gainpoor

The likely result is heterogeneous. Basic administrative and translation tasks may converge, while frontier science and advanced industrial design diverge.

8.16.3 Education

Low-cost AI can distribute tutoring and translation to underserved regions. It can also create a new inequality between students using free generic systems and those with frontier access, institutional data, high-performance devices and expert supervision.

8.16.4 Public services

Governments can deploy AI for:

  • document processing;
  • fraud detection;
  • tax administration;
  • citizen support;
  • translation;
  • resource allocation;
  • health triage.

The World Bank’s 2025 GovTech Maturity Index covers 198 economies and reports a global average increase from 0.552 in 2022 to 0.589 in 2025, with widening differences between the strongest and weakest maturity groups. GovTech Maturity Index 2025 – World Bank – December/2025World Bank GovTech Maturity Index.

8.16.5 Healthcare

AI can extend scarce specialist capacity through decision support, imaging assistance, translation and administrative automation. Risks arise when imported models are not validated on local populations, diseases, languages or clinical practice.

8.16.6 Entrepreneurship

Cloud APIs reduce entry costs but expose startups to:

  • currency depreciation;
  • foreign payment restrictions;
  • sudden repricing;
  • model withdrawal;
  • data-residency limits;
  • API lock-in;
  • unequal access to enterprise discounts.

A startup paying dollar-denominated AI costs while earning local-currency revenue carries an embedded exchange-rate mismatch.

8.16.7 Skilled migration

Talent may move toward countries and institutions offering:

  • frontier compute;
  • higher salaries;
  • research datasets;
  • strong laboratories;
  • startup capital;
  • cloud capacity.

The migration feedback is:

Research Compute → Talent Attraction → Publications and Startups → Funding → More Research Compute

The reverse can become a scientific-capability trap.

8.16.8 National security

Dependence on foreign AI platforms affects:

  • intelligence analysis;
  • defence supply chains;
  • critical-infrastructure operations;
  • cyber defence;
  • government data;
  • continuity during geopolitical crisis.

Sovereignty does not require domestic production of every chip. It requires credible continuity options, diversified providers, portable systems, local evaluation and priority access during crisis.


8.17 Evidence that AI can reduce development gaps

Equalising mechanismDevelopment effect
Small AI on ordinary devicesReduces hardware requirements
Cloud accessEliminates minimum local data-centre investment
TranslationBroadens access to global knowledge
Open-weight modelsReduces licence and platform dependence
Shared public computeSpreads fixed cost
Mobile distributionReaches users without personal computers
Agricultural and health specialisationTargets high-value local problems
Coding assistanceExpands software-production capability
Digital public infrastructureEnables scalable government services
International research creditsProvides temporary frontier access

The World Bank reports that middle-income countries accounted for more than 40% of global ChatGPT traffic by mid-2025, led by countries including Brazil, India, Indonesia and Viet Nam. It also reports rapid growth in generative-AI-related vacancies in middle-income economies. Strengthening AI Foundations: Emerging Opportunities for Developing Countries – World Bank – November/2025World Bank developing-country AI factsheet.


8.18 Evidence that AI can widen development gaps

Divergence mechanismConsequence
Expensive accelerators and memoryCapital-intensive capability remains concentrated
Import taxes and currency depreciationHardware costs rise relative to local income
Electricity unreliabilityLocal inference and data centres become costly
Lack of cloud regionsHigher latency and weaker sovereignty
Export controlsAdvanced capacity differs by geopolitical alignment
Dollar-denominated subscriptionsCurrency shocks affect access
Weak local-language performanceLower productivity gain
Research-compute concentrationFrontier science remains geographically concentrated
Platform lock-inDomestic value leaks to foreign providers
Skilled migrationWeak ecosystems lose their most capable researchers
Data scarcityModels perform worse on local conditions
Regulatory uncertaintyInvestment and deployment are delayed
Free-versus-frontier tieringNominal inclusion masks capability inequality

UNCTAD reported that data centres accounted for more than one-fifth of global greenfield investment in 2025, illustrating the growing capital intensity of digital infrastructure. Data Centres Are Reshaping the Global Investment Landscape – UN Trade and Development – January/2026UNCTAD data-centre investment analysis.


8.19 Regional findings

8.19.1 North America

The United States remains the principal frontier producer. Canada benefits from geographic integration, research and cloud access but remains dependent on foreign platforms and hardware supply. Internal inequality is the principal risk: national abundance does not guarantee affordable frontier access for every student, household or small firm.

8.19.2 European Union and United Kingdom

Europe possesses strong research, industrial capacity, connectivity and regulatory institutions but lacks equivalent ownership of the dominant frontier platforms and advanced accelerator design ecosystem. The EU aggregate conceals major differences between northern and western member states and lower-income or smaller members. The UK retains strong research and financial-sector demand but shares Europe’s dependence on imported hardware and foreign hyperscalers.

8.19.3 East Asia

China, Japan and South Korea all possess high capability but different constraints. China has market scale and domestic platforms but faces export controls. Japan has strong industrial and research capacity but depends on imported leading accelerators. South Korea combines world-leading memory and semiconductor capabilities with high connectivity, although frontier-accelerator access remains geopolitically exposed.

8.19.4 Israel

Israel’s strengths in research, cybersecurity, semiconductors and startups yield high readiness. Its constraints are market size, regional security, infrastructure concentration and dependence on imported leading-edge hardware.

8.19.5 India

India demonstrates the difference between national capability and household access. It possesses world-scale software talent, domestic cloud regions and an expanding AI market, while income, language, rural connectivity and educational inequality produce a very wide internal distribution.

8.19.6 Latin America

Brazil and Mexico possess large markets and cloud access but face hardware affordability, taxation, exchange-rate and research-compute constraints. Brazil benefits from Portuguese-language scale; Mexico benefits from proximity and industrial integration with the United States.

8.19.7 Balkans

Croatia benefits from EU integration. Serbia has technical talent and potential regional importance but a smaller domestic compute base. Albania has made progress in digital government, while Bosnia and Herzegovina faces institutional fragmentation. Regional cooperative compute could be economically superior to duplicated small national clusters.

8.19.8 North Africa

Morocco has strong nearshore and multilingual potential; Tunisia possesses significant technical talent but faces capital and currency constraints; Egypt combines population scale, cable geography and Arabic-language demand with severe affordability pressures. These countries could become localisation and service hubs if electricity, cloud and research compute expand.

8.19.9 Sub-Saharan Africa

South Africa has the region’s strongest hyperscale and research position but suffers electricity and income inequality. Kenya is an East African digital-services hub. Nigeria has enormous market potential but severe power and currency constraints. Ghana and Rwanda have policy and digital-government strengths but limited compute. Ethiopia remains constrained across several complementary foundations.


8.20 Policy architecture

Policy objectiveInstrumentMeasurable outcome
Reduce hardware burdenTariff relief and pooled procurementWorkstation wage-month ratio
Improve power reliabilityGrid and dedicated research-power investmentOutage-adjusted compute availability
Expand cloud accessRegional investment and competition policyLatency and eligible service depth
Protect affordabilityStudent, SME and researcher creditsAAI and SAI reduction
Build scientific sovereigntyNational or regional research clustersResearch accelerator-hours per researcher
Improve language inclusionLocal-language datasets and evaluationsLLQ parity score
Reduce lock-inPortability and interoperability rulesSwitching time and cost
Address currency riskLocal-currency billing and hedging facilitiesAI price volatility
Expand skillsApplied AI and verification trainingSuccessful-task productivity
Support small AIGrants for local task-specific systemsCost per validated local workload
Enable regional cooperationShared compute and data infrastructuresUtilisation and member access
Protect securitySupply diversification and fallback capacityRecovery time during provider disruption

8.21 Five-year scenarios, 2027–2031

ScenarioCore developmentInequality result
Inclusive diffusionSmall AI, open models and public infrastructure scaleGlobal divide narrows
Platform dependenceCloud access expands without domestic capabilityConsumption rises; sovereignty gap widens
Frontier bifurcationBasic AI becomes cheap while frontier compute remains scarceNominal access converges; productive capability diverges
Geopolitical fragmentationExport controls and sovereign ecosystems expandAccess follows alliances and jurisdiction
Infrastructure leapfrogSelected emerging economies build energy and cloud hubsNew regional leaders emerge
Currency and debt shockImport and subscription affordability deterioratesLow-income access contracts
Regional cooperationShared Balkan, African or Latin American capacity growsSmall-country scale disadvantage falls

The most plausible central scenario is frontier bifurcation. Low-cost AI becomes widespread through mobile services, smaller models and cloud APIs, generating real benefits in education, translation, agriculture, administration and entrepreneurship. Simultaneously, frontier research, high-assurance enterprise systems, advanced multimodal agents and sovereign local deployment remain concentrated in countries possessing abundant capital, reliable electricity, leading hardware, cloud regions and research institutions.


8.22 Falsifiable hypotheses

HypothesisTest
H8.1Higher GAI predicts higher firm-level AI adoption after controlling for income
H8.2Hardware wage-month ratios predict local-model adoption
H8.3Domestic cloud regions increase enterprise AI adoption and reduce latency
H8.4Electricity unreliability reduces local compute investment
H8.5Export restrictions reduce research-compute growth in affected jurisdictions
H8.6Local-language quality predicts public and SME adoption
H8.7Foreign-currency shocks reduce API consumption in developing economies
H8.8Public research compute reduces skilled migration
H8.9Small AI produces larger proportional gains in low-income settings for bounded tasks
H8.10Frontier capability remains more concentrated than basic AI use
H8.11Between-country inequality is smaller than combined between- and within-country inequality
H8.12Regional cooperative infrastructure improves access in small economies

8.23 Final findings

  1. AI readiness is a system of complements. Connectivity without electricity, skills without compute or cloud access without foreign currency cannot produce full capability.
  2. The global digital divide remains large. Approximately 2.2 billion people remained offline in 2025, overwhelmingly in low- and middle-income economies.
  3. National averages conceal severe internal inequality. India, Brazil, China, the United States, South Africa and Nigeria contain both globally competitive centres and severely constrained communities.
  4. Cloud geography is a development variable. Domestic and nearby regions affect latency, availability, residency and bargaining power.
  5. Advanced hardware access is geopolitical. Export controls and sanctions can alter national compute capacity independently of income or technical skill.
  6. Hardware affordability should be measured in wage months, not nominal dollars. Identical equipment can represent less than one month of earnings in one country and many months in another.
  7. Local-language quality is part of infrastructure. A country cannot be considered fully AI-ready if its population receives materially weaker performance in national languages.
  8. Research compute is a sovereignty asset. Its absence can drive publication gaps, dependence on foreign platforms and skilled migration.
  9. AI can reduce development gaps. Small models, mobile access, translation, cloud APIs and shared infrastructure can distribute useful capability at low marginal cost.
  10. AI can simultaneously widen frontier gaps. Advanced accelerators, energy, data centres, research funding and enterprise tools remain highly concentrated.
  11. The provisional pilot Gini of 0.193 understates total inequality. It omits population weighting and within-country access disparities.
  12. The pilot Theil decomposition attributes 91.4% of measured dispersion to differences between broad readiness groups. This result is illustrative and depends on the constructed sample and grouping.
  13. The correct policy objective is not national ownership of every technology. It is affordable access, continuity, evaluation capacity, portability and the ability to adapt AI to local needs.
  14. Regional cooperative compute is especially important for small countries. It can spread fixed costs and reduce dependence, provided governance and capacity allocation are credible.
  15. The central risk through 2031 is a world in which basic AI becomes nearly universal while frontier-quality, private, reliable and scientifically useful AI remains concentrated. Such a system would reduce some development gaps while institutionalising a deeper global hierarchy of intelligence capability.

CHAPTER 9 — Five-Year Forecasts and Quantitative Scenarios, 2026–2031

9.1 Forecast objective, base year and evidentiary status

Forecasting AI infrastructure differs from forecasting a mature commodity. The system is undergoing simultaneous structural changes in model architecture, accelerator design, memory intensity, advanced packaging, data-centre construction, electricity procurement, inference optimisation, service pricing, regulation and geopolitical access. Historical series are short, product generations are not quality-equivalent, reported prices frequently differ from negotiated prices, and several variables—particularly billable inference tokens and advanced-packaging output—are not disclosed through complete global datasets. A single extrapolated trend would therefore create false confidence.

This chapter uses a model ensemble. Where an official physical baseline exists, such as US data-centre electricity consumption, the forecast reports physical units. Where the underlying global level is not observed consistently, it reports an index with 2026=100. Nominal-price forecasts are separated from quality-adjusted prices. Central, low and high trajectories are scenario paths rather than claims of deterministic outcomes. The 80% and 95% ranges are probabilistic simulation intervals only where a parameter distribution has been specified; they are not presented as classical time-series confidence intervals when the historical sample is inadequate.

The official evidence establishes a high-growth but capacity-constrained starting point:

  • US data-centre electricity consumption increased from 58 TWh in 2014 to 176 TWh in 2023 and was officially projected at 325–580 TWh by 2028. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – U.S. Department of Energy – December/2024US Department of Energy data-centre electricity forecast.
  • TSMC stated in April 2025 that it was working to double CoWoS capacity during 2025. TSMC First Quarter 2025 Earnings Conference – TSMC – April/2025TSMC Q1 2025 investor transcript.
  • Micron stated in March 2026 that DRAM and NAND supply-demand conditions were expected to remain tight beyond calendar 2026. Micron Fiscal Second Quarter 2026 Earnings Call – Micron – March/2026Micron Q2 FY2026 investor disclosure.
  • Microsoft disclosed USD 31.9 billion of quarterly capital expenditure in fiscal Q3 2026, with approximately two-thirds directed to shorter-lived assets, primarily GPUs and CPUs. Microsoft Fiscal Year 2026 Third Quarter Earnings Conference Call – Microsoft – April/2026Microsoft FY2026 Q3 investor disclosure.
  • Alphabet disclosed USD 44.9 billion of capital expenditure in Q2 2026, with the vast majority directed to technical infrastructure supporting AI; approximately 60% of technical-infrastructure investment was in servers and 40% in data centres and networking. Alphabet 2026 Q2 Earnings Call – Alphabet – July/2026Alphabet Q2 2026 investor disclosure.

9.2 Forecast variables and measurement units

VariableForecast unitBase-year interpretation
Consumer and professional GPU priceNominal price index, 2026=100Comparable AI-capable hardware basket
Quality-adjusted accelerator priceCost per successful workload indexControls for memory, throughput and quality
HBM pricePrice-per-GB indexBlended high-bandwidth-memory generation
DRAM pricePrice-per-GB indexBlended server and workstation DRAM
Advanced packagingCapacity indexQuality-adjusted CoWoS-equivalent capacity
Data-centre CAPEXNominal global investment indexAI and cloud technical infrastructure
Electricity demandUS data-centre TWh plus global indexPhysical US anchor
AI inference demandQuality-adjusted successful-task indexNot raw generated tokens alone
API priceCost per successful workload indexModel-, tool- and quality-adjusted
Enterprise AI expenditureReal expenditure indexLicences, compute, integration and governance
Local-AI affordabilityCost-to-income indexLower values indicate greater affordability
Global access inequalityPilot GAI GiniCountry-average quality-adjusted access
Sectoral adoption costReal total-cost indexFull implementation and governance burden

9.3 Forecasting methods

9.3.1 Trend and CAGR model

For a variable X:

CAGR = (XT ÷ X0)1/T − 1

and:

t = X0(1 + CAGR)t

CAGR is used only as a descriptive baseline. It is unreliable when supply constraints, product cycles or regulatory shocks change the growth regime.

9.3.2 Exponential smoothing

Simple exponential smoothing is:

Lt = αXt + (1−α)Lt−1

A damped-trend extension is:

t+h = Lt + (φ + φ2 + … + φh)Bt

where φ below one prevents implausibly persistent exponential growth.

9.3.3 ARIMA

An ARIMA(p,d,q) model is:

φ(B)(1−B)dXt = c + θ(B)εt

ARIMA should be used only for sufficiently long and consistent monthly or quarterly series, such as memory spot prices or electricity demand. It is unsuitable for a five-point annual series of non-comparable accelerator products.

9.3.4 Dynamic regression

For memory or accelerator prices:

ΔlnPt = α + β1ΔlnDemandt − β2ΔlnCapacityt + β3Energyt + β4FXt + β5Restrictiont + εt

This allows supply, demand, energy, currency and geopolitical variables to affect prices.

9.3.5 Panel-data model

For country i and year t:

GAIi,t = α + β1Incomei,t + β2Broadbandi,t + β3Poweri,t + β4Cloudi,t + β5Skillsi,t + μi + τt + εi,t

Country fixed effects μi control for time-invariant structural characteristics; period effects τt control for global changes.

9.3.6 Stock-flow model

Installed AI capacity evolves as:

Kt+1 = Kt + It − δKt − Rt

where:

  • Kt is productive capacity;
  • It is new installation;
  • δ is physical retirement;
  • Rt is economic obsolescence or inaccessible capacity.

Effective capacity is:

Keffective,t = Kt × Ut × At × Et

where U is utilisation, A is availability and E is software efficiency.


9.4 Learning curves

The requested learning relationship is:

Ct = C0(Qt ÷ Q0)−b

The implied learning rate is:

LR = 1 − 2−b

If cumulative inference output increases tenfold and quality-adjusted cost falls from 100 to 32:

0.32 = 10−b

b = −ln(0.32) ÷ ln(10)
b ≈ 0.495

Therefore:

LR = 1 − 2−0.495
LR ≈ 29.0%

Under this central case, each doubling of cumulative output reduces quality-adjusted cost by approximately 29%.

9.4.1 Learning-rate sensitivity

Learning exponent bImplied learning rateCost after four doublings
0.106.7%75.8
0.2012.9%57.4
0.3018.8%43.5
0.4024.2%33.0
0.5029.3%25.0
0.6034.0%18.9

Learning effects can be offset by capability escalation. If users migrate from inexpensive basic models to more compute-intensive reasoning and agentic systems, the cost of a fixed 2026 task may decline while average expenditure per user rises.


9.5 Central annual forecast, 2026–2031

9.5.1 Technology and infrastructure

All indices use 2026=100 unless otherwise specified.

Variable2026202720282029203020312026–2031 CAGR
Nominal GPU/accelerator price1001011031041051051.0%
Quality-adjusted accelerator price1008268574841−16.4%
HBM price per GB1009488837873−6.1%
Conventional DRAM price per GB1009590868278−4.9%
Advanced-packaging capacity10013016520524528523.3%
Data-centre CAPEX10012214516618419914.8%
AI inference demand1001803105007601,10061.6%
Quality-adjusted API price1007558463832−20.4%
Enterprise AI expenditure10013016520525030024.6%
Local-AI cost-to-income ratio1009589848076−5.3%

The central result is not contradictory: API prices and quality-adjusted hardware costs can fall while total enterprise expenditure rises. Demand, integration depth, model capability, multimodality, agentic tools and governance can expand faster than unit costs decline.


9.5.2 US data-centre electricity

YearCentral forecastLow trajectoryHigh trajectory
2026330 TWh280 TWh380 TWh
2027390 TWh310 TWh480 TWh
2028450 TWh325 TWh580 TWh
2029515 TWh350 TWh690 TWh
2030585 TWh380 TWh810 TWh
2031660 TWh410 TWh930 TWh

The official DOE range applies to 2028; values beyond 2028 are scenario extensions, not official federal projections. The central 2028 value of 450 TWh lies within the official 325–580 TWh range.


9.6 Price forecasts

9.6.1 GPU and accelerator prices

Nominal accelerator prices are expected to be substantially more rigid than quality-adjusted prices. Successive products may provide more memory and throughput while preserving or increasing headline prices.

Forecast202620272028202920302031
Nominal low1009591878480
Nominal central100101103104105105
Nominal high100110121133146161
Quality-adjusted central1008268574841

The high nominal path corresponds to persistent scarcity, premium product migration, memory bottlenecks, tariffs and limited competition. The low path assumes strong competition, inventory normalisation and improved alternatives.

9.6.2 HBM

HBM’s price trajectory depends on two opposing forces:

  • generational improvements and capacity expansion reduce price per bit;
  • rapid demand, complex stacking, yield and capacity trade-offs sustain premiums.

Micron stated that the expansion of HBM creates a significant trade ratio against conventional DRAM capacity and continued to describe supply-demand conditions as tight. Micron Fiscal First Quarter 2026 Earnings Call – Micron – December/2025Micron HBM and DRAM supply disclosure.

HBM price-per-GB index202620272028202920302031
Low1008269585043
Central1009488837873
High100112120125128130

9.6.3 DRAM

DRAM price-per-GB index202620272028202920302031
Low1008574665953
Central1009590868278
High100108113116118120

The high case reflects wafer allocation toward HBM and advanced products, constraining conventional memory supply despite technological progress.


9.7 Advanced-packaging forecast

TSMC’s 2025 annual report described continuing development of CoWoS, InFO, SoIC and related three-dimensional integration technologies in response to AI demand. TSMC 2025 Annual Report – TSMC – 2026TSMC 2025 Annual Report.

Advanced-packaging capacity index202620272028202920302031
Low100115135155175195
Central100130165205245285
High100150215300400515

Capacity growth does not guarantee equivalent growth in completed accelerator systems. HBM, substrates, interposers, optical networking, power equipment and final-server integration must expand simultaneously.


9.8 Data-centre CAPEX

9.8.1 Central trajectory

Data-centre CAPEX index202620272028202920302031
Low100108116122126128
Central100122145166184199
High100145195255325400

The central path assumes continued expansion but declining annual growth as grid, construction, permitting and financing constraints become binding. The high path assumes a sustained infrastructure arms race, sovereign capacity programmes and rapid agentic demand.

9.8.2 CAPEX stock-flow relation

KDC,t+1 = KDC,t + CAPEXt − Depreciationt − Impairmentt

A rise in capital expenditure does not translate immediately into productive capacity because of:

  • construction lead time;
  • grid connection;
  • accelerator delivery;
  • cooling commissioning;
  • software integration;
  • customer ramp;
  • stranded or underutilised assets.

9.9 AI inference demand

Inference demand should be measured in successful quality-adjusted tasks rather than raw tokens. Let:

Dinference,t = Userst × TasksPerUsert × ComputePerTaskt ÷ Efficiencyt

Inference-demand index202620272028202920302031
Low100140200280380500
Central1001803105007601,100
High1002505501,1002,0003,400

Efficiency can increase demand through a rebound mechanism:

Total Compute = Cost per Task × Number of Tasks

If cost per task falls by 70% while task volume increases tenfold, total expenditure and physical compute can still rise substantially.


9.10 API price forecast

9.10.1 Quality-adjusted cost per successful task

API price index202620272028202920302031
Low-price path1006543302217
Central1007558463832
High-price path100100103108115125

The central path assumes learning, routing, caching and competition. The high path reflects premium frontier migration, reasoning-token billing, priority charges, long-context multipliers and tight capacity.

9.10.2 Tariff decomposition

PAI,t = Pbase,t × Mmodel,t × Mlatency,t × Mcontext,t × Mregion,t + Ctools,t

A falling base price can coexist with a rising completed-task charge if the multipliers grow.


9.11 Enterprise expenditure

Enterprise AI expenditure index202620272028202920302031
Low100115132150168185
Central100130165205250300
High100155225320440590

The central forecast implies 24.6% annual growth. Expenditure includes:

  • licences;
  • API usage;
  • local and cloud infrastructure;
  • data preparation;
  • integration;
  • cybersecurity;
  • governance;
  • compliance;
  • human supervision;
  • evaluation;
  • switching and redundancy.

Unit-price deflation therefore does not imply budget deflation.


9.12 Sectoral adoption-cost forecasts

The table reports real total adoption-cost indices.

Sector202620272028202920302031
Microenterprises100117135153170185
SMEs100122148176205235
Large corporations100130164202243285
Banking100135170207245285
Insurance100130160195230265
Healthcare100138180225272320
Legal/professional100120140160178195
Manufacturing100128158190222255
Media100125150176200225
Telecommunications100135172212255300
Defence/critical infrastructure100142188238290345
Public administration100126155188220250
Education/research100125150178205230

High-assurance sectors rise fastest because deployment depth, security, validation and redundancy expand faster than model-unit prices decline.


9.13 Local-AI affordability

The local-AI affordability ratio is:

LAAc,t = Annualised Local AI TCOc,t ÷ Median Disposable Incomec,t

Local-AI affordability index202620272028202920302031
Improving path1008877675850
Central1009589848076
Deteriorating path100108116125133140

The central case improves slowly because quality-adjusted hardware costs fall but adequate-model requirements, memory, electricity and capability expectations rise.


9.14 Global-access inequality forecast

Using the Chapter 8 provisional GAI Gini of 0.193 as the 2026 analytical baseline:

GAI Gini202620272028202920302031
Convergence path0.1930.1850.1780.1710.1650.160
Central tiered-access path0.1930.1980.2030.2080.2120.215
Divergence path0.1930.2100.2280.2470.2680.290

The central case assumes basic access diffuses but frontier infrastructure remains concentrated.


9.15 Monte Carlo framework

For each simulation b:

  1. Draw inference-demand growth.
  2. Draw capacity growth.
  3. Draw learning exponent.
  4. Draw energy price and grid-connection delay.
  5. Draw memory and packaging constraints.
  6. Draw geopolitical regime.
  7. Calculate capacity, price, expenditure and access.
  8. Repeat at least 50,000 times.

Illustrative parameter distributions:

ParameterDistributionCentral parameter
Annual inference growthLognormal55%
Advanced-packaging growthTriangular15%, 25%, 45%
API learning exponentNormal, truncated positive0.50
Accelerator economic lifeTriangular3, 4, 6 years
Data-centre power delayDiscrete0–4 years
Export-control shockBernoulliScenario-dependent
HBM supply shockBernoulliScenario-dependent
Electricity-price growthLognormal3%
Enterprise adoptionLogistic diffusionSector-dependent

9.16 2031 forecast intervals

These intervals are model-based simulation ranges, not classical confidence intervals derived from long stationary histories.

Variable2031 central80% interval95% interval
Nominal GPU price index10582–14565–190
Quality-adjusted accelerator price4125–6515–100
HBM price-per-GB index7350–11035–150
DRAM price-per-GB index7855–11040–145
Advanced-packaging capacity285220–360170–470
Data-centre CAPEX199155–270120–380
US data-centre electricity660 TWh480–800 TWh390–930 TWh
AI inference demand1,100600–1,900350–3,200
Quality-adjusted API price3220–5012–80
Enterprise AI expenditure300220–410160–600
Local-AI affordability index7655–10540–140
GAI Gini0.2150.185–0.2500.160–0.300

9.17 Scenario probabilities

The scenario probabilities are structured judgments derived from:

  • current investment commitments;
  • concentration documented in previous chapters;
  • official packaging and memory statements;
  • power forecasts;
  • export-control developments;
  • observed pricing segmentation;
  • model-efficiency trends;
  • historical semiconductor cyclicality.

They are not objective frequencies. The scenarios are partly overlapping: financialised pricing can coexist with oligopoly, fragmentation or scarcity.

ScenarioCentral judgementPlausible range
A — Abundant Compute20%15–25%
B — Managed Oligopoly35%30–40%
C — Persistent Scarcity18%15–25%
D — Fragmented AI World17%15–25%
E — AI Utility and Financialised Access10% as dominant regime10–20%

The central point estimates sum to 100% only to construct a decision-weighted baseline. The plausible ranges do not need to sum to 100%.


9.18 Scenario A — Abundant Compute

9.18.1 Assumptions

  • advanced-packaging capacity expands rapidly;
  • HBM4 and later generations achieve strong yields;
  • foundry and memory competition increases;
  • grid and generation projects arrive on schedule;
  • open-weight models remain competitive;
  • inference software produces strong learning effects;
  • export restrictions remain limited;
  • cloud providers compete aggressively;
  • small and specialised models absorb routine demand.

9.18.2 Annual trajectory

Indicator202620272028202920302031
Quality-adjusted compute price1007558453528
HBM price-per-GB1008269585043
Packaging capacity100150215300400515
API successful-task price1006543302217
Enterprise expenditure100125155185215245
Global GAI mean56.46064687276
GAI Gini0.1930.1850.1780.1710.1650.160

9.18.3 Consequences

  • Companies obtain positive ROI from more use cases.
  • SMEs gain access to advanced functions without large CAPEX.
  • Household and student affordability improves.
  • Low- and middle-income countries benefit from mobile and small-model diffusion.
  • Frontier providers face margin pressure.
  • Total electricity demand can still rise because lower prices generate higher usage.

9.18.4 Early-warning indicators

  • packaging lead times fall below normal procurement cycles;
  • HBM contract premiums decline;
  • accelerator inventories rise;
  • API prices fall without tighter limits;
  • local models close capability gaps;
  • grid-connection queues shorten.

9.18.5 Invalidation events

  • major semiconductor disruption;
  • sustained HBM shortages;
  • severe power constraints;
  • widening export controls;
  • frontier workloads become much more compute-intensive than efficiency gains.

9.19 Scenario B — Managed Oligopoly

9.19.1 Assumptions

  • physical supply expands;
  • advanced foundry, HBM, packaging and cloud control remains concentrated;
  • providers maintain differentiated model tiers;
  • competition reduces some prices but preserves rents;
  • enterprise contracts and ecosystem lock-in increase;
  • interoperability remains incomplete;
  • basic models become cheap while frontier access remains premium.

9.19.2 Annual trajectory

Indicator202620272028202920302031
Quality-adjusted compute price1008878706459
HBM price-per-GB1009894908682
Packaging capacity100128160195230265
API successful-task price1008270615550
Enterprise expenditure100132170215265320
Global GAI mean56.45860626466
GAI Gini0.1930.1970.2020.2070.2110.215

9.19.3 Consequences

  • Large companies obtain volume discounts and reserved capacity.
  • SMEs remain dependent on bundled subscriptions.
  • Consumer free tiers improve but remain below frontier capability.
  • Platform vendors capture a growing share of enterprise technology budgets.
  • Countries without cloud regions or bargaining scale remain price takers.

9.19.4 Early-warning indicators

  • gross margins remain elevated despite capacity expansion;
  • multi-year supply agreements dominate;
  • enterprise prices remain opaque;
  • provider-specific tools deepen switching costs;
  • priority and long-context premiums expand.

9.19.5 Invalidation events

  • successful open standards materially reduce switching;
  • new foundry or accelerator competitors capture substantial share;
  • regulators impose structural interoperability;
  • oversupply causes severe price competition.

9.20 Scenario C — Persistent Scarcity

9.20.1 Assumptions

  • inference demand exceeds forecasts;
  • HBM, packaging, substrates and power remain constrained;
  • data-centre grid connections are delayed;
  • frontier-model compute per task rises;
  • equipment prices and financing remain elevated;
  • supply expansion arrives more slowly than demand.

9.20.2 Annual trajectory

Indicator202620272028202920302031
Nominal accelerator price100110121133146161
HBM price-per-GB100112120125128130
Packaging capacity100115135155175195
API successful-task price100105110115120125
Enterprise expenditure100150215300410540
Global GAI mean56.455.554.553.552.551
GAI Gini0.1930.2050.2180.2300.2400.250

9.20.3 Consequences

  • Companies ration AI use and prioritise highest-margin tasks.
  • Heavy agentic workflows become prohibitively expensive for SMEs.
  • Students and researchers rely increasingly on inferior tiers.
  • Wealthy countries secure capacity through long-term contracts.
  • Used local hardware retains high resale values.
  • Grid and electricity access become competitive differentiators.

9.20.4 Early-warning indicators

  • HBM remains sold out more than twelve months forward;
  • packaging lead times stop improving;
  • data-centre projects are delayed by power;
  • provider rate limits tighten;
  • reserved-capacity premiums increase;
  • used accelerator prices remain above reference values.

9.20.5 Invalidation events

  • major inference-efficiency breakthrough;
  • demand growth slows materially;
  • multiple new capacity sources achieve high yields;
  • small models replace frontier systems across most workloads.

9.21 Scenario D — Fragmented AI World

9.21.1 Assumptions

  • export controls broaden;
  • countries impose data-localisation and sovereignty rules;
  • technology blocs adopt incompatible accelerators and software;
  • cross-border cloud service becomes politically contingent;
  • sanctions and foreign-currency restrictions intensify;
  • national subsidies duplicate infrastructure.

9.21.2 Annual trajectory

Indicator202620272028202920302031
Global average compute price100105112118125132
Cross-bloc price dispersion100130165205250300
Duplicated infrastructure CAPEX100140190245305370
Global API interoperability1009078675850
Global GAI mean56.455.855.254.854.354
GAI Gini0.1930.2100.2300.2500.2700.290

9.21.3 Consequences

  • Allied countries receive different hardware and cloud access.
  • Global companies operate duplicate AI stacks.
  • Compliance and switching costs rise.
  • Countries with domestic ecosystems gain sovereignty but sacrifice some scale efficiency.
  • Scientific collaboration and model reproducibility decline.
  • Smaller countries face pressure to align with a technology bloc.

9.21.4 Early-warning indicators

  • new destination-based semiconductor restrictions;
  • regional model-licensing rules;
  • sovereign-cloud mandates;
  • incompatible national AI standards;
  • cross-border model withdrawal;
  • rising duplication of national compute clusters.

9.21.5 Invalidation events

  • multilateral technology agreements;
  • effective interoperable standards;
  • relaxation of export controls;
  • broad availability of competitive hardware from multiple jurisdictions.

9.22 Scenario E — AI Utility and Financialised Access

9.22.1 Assumptions

  • compute capacity is treated as a reservable strategic service;
  • prices vary by latency, model, context, geography and scarcity;
  • reserved, priority, batch and flex markets deepen;
  • providers sell forward capacity;
  • large users hedge future compute costs;
  • spot capacity becomes dynamically priced;
  • completed-task and outcome pricing expand.

9.22.2 Annual trajectory

Indicator202620272028202920302031
Standard compute-price index1009288858382
Priority-price multiplier1.5×1.7×2.0×2.3×2.6×3.0×
Batch discount50%52%55%58%60%62%
Reserved-capacity share100135180230285350
Enterprise expenditure100140185235290350
Global GAI mean56.457585959.560
GAI Gini0.1930.2000.2080.2180.2270.235

9.22.3 Consequences

  • Large firms hedge capacity and obtain predictable service.
  • SMEs pay volatile on-demand prices.
  • Latency-sensitive sectors pay substantial premiums.
  • Batch research and back-office tasks become cheaper.
  • Households receive low-cost standard service but expensive frontier priority.
  • Compute-finance products create counterparty and concentration risks.

9.22.4 Early-warning indicators

  • expansion of reserved and priority tiers;
  • compute-capacity marketplaces;
  • forward contracts and minimum-spend commitments;
  • dynamic congestion pricing;
  • task and outcome billing;
  • financial institutions offering compute hedging.

9.22.5 Invalidation events

  • abundant capacity eliminates scarcity premiums;
  • regulation prohibits dynamic discrimination;
  • model efficiency makes reservation unnecessary;
  • compute becomes too heterogeneous for standard contracts.

9.23 Company-level consequences by scenario

Company typeABCDE
MicroenterpriseStrong access improvementSubscription dependenceExclusion riskRegional service gapsVolatile usage bills
SMELower entry costVendor lock-inMargin pressureDuplicate complianceNeed for capacity budgeting
Large corporationBroad deploymentBargaining advantageReserved-capacity raceMultiple regional stacksHedging and portfolio optimisation
Bank/insurerLower analytical costConcentrated supplier riskHigh priority premiumsData-localisation burdenFinancialised capacity contracts
HealthcareWider validated useEnterprise-provider dependenceHigh-assurance scarcityNational-system divergencePriority service becomes essential
ManufacturerCheap edge AIPlatform integrationHardware rationingSupply-chain fragmentationLong-term compute procurement
Research institutionWider experimentationTiered frontier accessPublication inequalityRestricted collaborationBatch discounts but priority disadvantage
GovernmentBroader public servicesSovereignty concernCapacity rationingNational stacksStrategic compute reserve

9.24 Household and distributional consequences

ScenarioHousehold effectStudent/researcher effectInequality effect
AFalling effective pricesBroad advanced accessConvergence
BCheap basic access, premium frontierPersistent tieringGradual divergence
CHigher prices and tighter limitsSevere exclusionStrong divergence
DAccess depends on nationality and blocInternational collaboration declinesGeographic divergence
EBasic access cheap; priority expensiveTime-sensitive users disadvantagedPrice discrimination increases

9.25 Early-warning dashboard

IndicatorAbundance signalScarcity/fragmentation signal
HBM lead timeFalling below six monthsAbove twelve months
Packaging utilisationNormalisingPersistently near effective maximum
GPU street premiumBelow 5%Above 25%
Used-market premiumNegative/normal depreciationPersistent positive premium
API list pricesBroad declineFrontier and priority increase
Batch discountExpandsContracts
Data-centre power queueShortensMulti-year delay
Memory producer marginsNormaliseRemain exceptionally high
Cloud-region expansionBroad geographic diffusionBloc-specific concentration
Export controlsStable/narrowBroaden by destination and capability
Enterprise RPOStableRapid long-term reservation
Local-model quality gapNarrowsWidens
GAI GiniDeclinesRises
Research-compute distributionBroadensConcentrates

9.26 Model validation and annual updating

The forecast should be updated every quarter for market variables and annually for country and social variables.

Forecast error is:

et = Xt − X̂t

Recommended metrics are:

MAE = (1 ÷ n)Σ|et|

RMSE = √[(1 ÷ n)Σet2]

MAPE = (100 ÷ n)Σ|et ÷ Xt

Interval coverage is:

Coverage = Number of Observations Inside Interval ÷ Total Observations

If an 80% interval contains materially fewer than 80% of subsequent observations, the model is overconfident and its variance assumptions must be revised.


9.27 Final forecast conclusions

  1. Nominal accelerator prices are likely to fall much more slowly than quality-adjusted prices.
  2. HBM and advanced packaging remain the most important semiconductor bottlenecks through at least the early forecast period.
  3. Advanced-packaging capacity could nearly triple by 2031 in the central case, yet demand can still exceed supply.
  4. US data-centre electricity demand could reach approximately 660 TWh by 2031 in the central scenario, with a very wide model-based range.
  5. Inference demand is likely to grow substantially faster than physical capacity because lower prices generate new uses.
  6. Quality-adjusted API costs can fall by approximately two-thirds while total enterprise AI expenditure triples.
  7. Enterprise expenditure will increasingly shift from experimental licences to integration, governance, security, agents and reserved capacity.
  8. Local-AI affordability improves only slowly in the central case because the threshold of adequate capability continues to rise.
  9. Global access inequality is more likely to increase modestly than collapse, even if basic AI becomes widely available.
  10. Scenario A produces broad convergence but requires simultaneous success in capacity, competition, energy and model efficiency.
  11. Scenario B is the central structural case: supply expands while bargaining power and frontier access remain concentrated.
  12. Scenario C remains plausible because demand can expand faster than memory, packaging, power and construction.
  13. Scenario D would make geography and geopolitical alignment direct determinants of AI capability.
  14. Scenario E is already partially visible in reserved, priority, batch and flexible service tiers, although a fully financialised market is not inevitable.
  15. Prediction intervals are necessarily wide because technology, demand and policy are endogenous. Narrow intervals would be statistically misleading.
  16. The five scenarios are not mutually exclusive in practice. The world can exhibit abundance in small models, oligopoly in frontier systems, scarcity in HBM, fragmentation across geopolitical blocs and financialised pricing for priority capacity at the same time.
  17. The decisive monitoring variable is not the sticker price of one GPU or one million tokens. It is the quality-adjusted cost of a successfully completed workload, available at the required latency, privacy, reliability and geographic location.
  18. The dominant five-year risk is tiered abundance: large quantities of inexpensive basic intelligence coexist with scarce, costly and concentrated frontier capability.

CHAPTER 10 — Policy Responses, Strategic Options and Final Conclusions

10.1 Final policy thesis

The evidence developed throughout this report supports neither technological fatalism nor the stronger claim that all AI hardware has become uniformly more expensive after adjusting for performance. The defensible conclusion is more specific. Raw computational capability continues to improve rapidly, and quality-adjusted prices can decline when workloads fit the memory, software and reliability constraints of newer systems. At the same time, economically useful access to AI depends on more than arithmetic throughput. Accelerator memory, memory bandwidth, advanced packaging, power availability, data-centre capacity, model quality, context length, software compatibility, privacy, concurrency and access to complementary expertise form a joint production system. A fall in price per theoretical operation therefore does not guarantee a fall in the cost of completing a real workload. Users who require large memory capacity, sustained inference, high reliability, private deployment or frontier-level model quality can face rising absolute expenditure even while price per TFLOP declines. This distinction explains why households, students and small organisations can experience genuine exclusion without contradicting semiconductor learning curves. The central policy problem is not a universal shortage of every chip. It is the concentration of strategically complementary capabilities in a limited number of firms, jurisdictions, fabrication processes, packaging facilities, memory suppliers, cloud regions and software ecosystems.

The appropriate policy objective is consequently not to make every organisation own frontier-scale infrastructure. That would waste capital, electricity and specialised labour. The objective should be to ensure contestable, resilient and affordable access to a workload-appropriate capability floor. Public intervention is justified when a combination of indivisibility, high fixed costs, knowledge spillovers, market concentration, switching costs, public-interest privacy requirements and unequal purchasing power causes socially valuable users to receive less compute than produces the highest social return. Intervention must nevertheless avoid converting scarcity into permanent subsidy dependence or protecting inefficient national suppliers from competition. The recommended architecture combines shared infrastructure, targeted demand subsidies, interoperability, transparent metering, competitive procurement, repair and refurbishment, open-weight research, grid investment and narrowly conditioned industrial support. Supply subsidies should be milestone-based, technology-neutral where possible, competitively awarded and subject to clawbacks. Access subsidies should target students, researchers, SMEs and public-interest institutions rather than reduce prices indiscriminately for the largest buyers. Competition enforcement should be evidence-driven: high concentration or margins establish a reason to investigate but do not independently prove exclusionary conduct, unlawful coordination or excessive pricing.

European policy already provides a foundation. The Commission reports an ambition to mobilise EUR 200 billion for AI, including EUR 20 billion for as many as five AI gigafactories, while an expanding network of AI Factories is intended to serve research, industry, start-ups and SMEs. AI Continent – European Commission – 2026official programme page. EuroHPC has separately described investment approaching EUR 2 billion in AI Factories; individual projects demonstrate the scale involved, including the EUR 290 million IT4LIA system, financed equally by EuroHPC and Italy. EuroHPC JU signs contract to boost AI capabilities at the IT4LIA AI Factory – EuroHPC Joint Undertaking – April 2026verified project announcement. These programmes should be treated as infrastructure platforms rather than one-time machine purchases. Their economic performance must be measured through utilisation, queue time, workload completion, geographic distribution, research output, SME survival and the share of capacity reaching users who could not otherwise purchase equivalent service.


10.2 Governing principles

A defensible access policy should satisfy seven tests:

  1. Additionality: public expenditure must add capacity, competition, resilience or access that the market would not supply on comparable terms.
  2. Capability targeting: eligibility must depend on workload requirements rather than parameter count or marketing labels.
  3. Contestability: funded infrastructure must support multiple frameworks, models and providers.
  4. Portability: users must be able to export data, prompts, embeddings, evaluation records, fine-tuning artefacts and operational metadata in documented formats.
  5. Cost transparency: metering must separate model inference, storage, retrieval, networking, tool use, reserved capacity and priority premiums.
  6. Distributional accountability: programmes must report who receives compute, not merely aggregate utilisation.
  7. Sunset and review: subsidies require expiry dates, counterfactual evaluation and recovery provisions when recipients fail investment, access or pricing obligations.

The Data Act supplies an existing European foundation for cloud switching and interoperability and has applied since 12 September 2025. Data Act Explained – European Commission – December 2025official explanation. The Digital Markets Act complements competition law by imposing obligations and prohibitions on designated gatekeepers, but its application to AI-related self-preferencing must follow the Act’s defined scope and legal tests rather than assume that every vertically integrated AI supplier is automatically covered. The Digital Markets Act – European Commission – 2026official DMA portal.


10.3 Policy-design matrix: mandate, recipients, cost and timing

All amounts below are planning estimates, not enacted appropriations, unless expressly identified as an official programme figure.

Policy instrumentProblem addressedLegal or institutional basisImplementing authorityPrincipal targetEstimated public costImplementation period
Public compute infrastructureIndivisible capital costs; private under-provision for research and public-interest workloadsNational research, industrial and digital-infrastructure statutes; in the EU, EuroHPC, Digital Europe and research competencesDigital ministry, science ministry, national HPC body, EuroHPCUniversities, public research, SMEs, hospitals, public administrationEUR 100m–1bn per national node; EUR 2bn–10bn for a multi-node system2–5 years
European AI cloud and gigafactoriesDependence on non-European frontier capacity; insufficient scaleEuroHPC Regulation, Digital Europe, Horizon Europe, InvestAI-compatible finance, applicable State-aid rulesCommission, EuroHPC JU, EIB, Member StatesEuropean research, industry, start-ups and sovereign workloadsEUR 10bn–25bn additional public and risk-sharing capital over five years; official wider ambition includes EUR 20bn for up to five gigafactories3–6 years
Shared university computeFragmented procurement and low utilisation of departmental systemsUniversity and research-funding mandatesResearch councils and university consortiaStudents, doctoral candidates and researchersEUR 25m–300m per consortium plus annual operating expenditure of approximately 10–20% of CAPEX1–4 years
Student compute vouchersIncome-based exclusion from capable AIEducation and equal-opportunity legislationEducation ministries, universities, grant agenciesLow-income students; disability and language-access usersEUR 250–1,000 per eligible student/year6–18 months
Research compute grantsEarly-career and independent researchers lack purchasing powerPublic research-grant authorityResearch councils and foundationsResearchers without institutional clustersEUR 2,000–20,000 per researcher/year; larger competitive awards for high-compute projects6–18 months
Investment tax creditsSocially useful capacity may be delayed by financing constraintsTax code; EU State-aid compatibility where applicableTreasury and tax authorityData centres, packaging, memory, energy-efficient computeCredit of 10–25% of qualifying incremental investment, with a jurisdictional fiscal cap1–3 years
Cooperative purchasingSMEs and universities face weak bargaining power and duplicative procurementProcurement and cooperative statutesCentral purchasing bodies or sector consortiaSMEs, schools, municipalities, universitiesEUR 1m–10m national administration; EUR 20m–50m for an EU-scale platform6–24 months
Open-weight model programmeDependence on closed services; weak local-language coverageResearch, cultural and digital-innovation mandatesResearch agencies, public laboratories, universitiesResearchers, SMEs, public services, minority-language communitiesEUR 25m–250m nationally or EUR 500m–2bn at EU scale over five years2–5 years
Interoperability and portability rulesCloud, data and workflow lock-inEU Data Act; national data and competition lawCommission, national competent authorities, standards bodiesAll enterprise and public-sector customersPublic administration EUR 20m–100m at EU scale; private compliance costs additional1–3 years
Standardised AI price disclosureComplex tariffs obscure effective prices and switching costsConsumer, contract, procurement and sectoral transparency lawConsumer authority, digital regulator, procurement agenciesConsumers, SMEs and public buyersEUR 5m–20m per major jurisdiction plus provider compliance cost12–24 months
Competition enforcement unitInformation asymmetry and technically complex vertical conductTFEU Articles 101 and 102; national equivalentsDG Competition and national competition authoritiesMarkets for accelerators, cloud, models and distributionEUR 50m–200m over five years across a major regional authority networkImmediate; permanent
Enhanced merger scrutinyAcquisitions can remove nascent competitors or close complementary inputsEU Merger Regulation and national merger lawCommission and national authoritiesSemiconductor, cloud, AI and data transactionsIncluded in enforcement staffing; specialist analysis EUR 5m–30m/yearImmediate
Restriction of proven self-preferencingVertically integrated platforms may disadvantage rivalsDMA where legally applicable; competition law elsewhereCommission and digital regulatorsBusiness users and competing AI servicesEnforcement cost included above1–3 years per proceeding
AI-aware public procurementPublic contracts may create lock-in or reward opaque pricingProcurement directives and national procurement lawCentral purchasing bodies and contracting authoritiesPublic administrations, schools, hospitalsAdministrative reform EUR 10m–100m; underlying purchases remain programme expenditure1–3 years
Grid and low-carbon generation investmentPower and interconnection constrain data-centre expansionEnergy, network and planning lawEnergy ministry, regulator and grid operatorsEconomy-wide users, including data centresEUR 5bn–50bn nationally or EUR 50bn–250bn regionally over five years; only a fraction attributable to AI3–10 years
Semiconductor and packaging supportAdvanced fabrication and packaging are geographically concentratedEuropean Chips Act; national industrial legislationCommission, Member States, development banksFoundries, packaging, equipment, substratesEUR 5bn–40bn per major jurisdictional programme4–10 years
HBM diversificationLimited qualified supply and long capacity cyclesIndustrial, R&D and State-aid frameworksIndustry ministry, research agencies, development banksMemory fabrication, packaging and materialsEUR 2bn–15bn per programme4–8 years
Right to repair and longer supportShort hardware life raises effective access costs and e-wasteConsumer, ecodesign and repair legislationConsumer authorities and product regulatorsHouseholds, schools, SMEs and refurbishersEUR 20m–100m administration; manufacturer transition costs additional2–5 years
Verified refurbished marketQuality uncertainty, fraud and absent warranties weaken secondary supplyConsumer safety, warranty and certification lawStandards body, consumer authority, accredited certifiersLow-income users, schools, SMEsEUR 25m–150m setup, testing network and optional guarantee reserve1–3 years
Educational and research AI entitlementAdequate AI may become a necessary educational inputEducation and research law; annual budget legislationEducation ministries and research councilsStudents, teachers and accredited researchersEUR 0.5bn–3bn/year for a large jurisdiction, depending on coverage1–3 years
International Compute Development FundCountries with weak grids, foreign-exchange constraints or absent cloud regions cannot finance accessIBRD, IDA and regional-development-bank mandates; donor agreementsWorld Bank, regional development banks, ITU-linked partners and national governmentsLow- and lower-middle-income economiesUSD 10bn–30bn over five years2–7 years

Table metadata: Unit: programme expenditure or benefit level. Currency: 2026 EUR unless USD is expressly stated. Price basis: real planning cost, excluding private co-finance unless noted. Reference year: 2026. Geographic scope: EU and representative national or international programmes. Sources: official programme records cited in this chapter plus author estimates. Method: benchmark scaling from official compute projects, programme design and beneficiary counts. Uncertainty: generally ±50–100% because facility configuration, energy works, financing structure and beneficiary coverage remain unspecified.


10.4 Policy performance, safeguards and feasibility

InstrumentExpected benefitMandatory KPIPrincipal unintended consequenceDistortion-control mechanismFeasibility
Public computeAccess to workloads too large or sensitive for ordinary usersSuccessful jobs; queue time; utilisation; cost per completed workload; share allocated to underserved usersPolitical allocation, low utilisation, rapid obsolescenceIndependent capacity auction; access quotas; external evaluation; depreciation reserveHigh, if attached to existing HPC institutions
EU AI cloudStrategic capacity and cross-border scaleDelivered accelerator-hours; workload portability; regional latency; SME/research shareDuplication, national capture, subsidy raceCompetitive site selection; cross-border access; open interfaces; clawbacksMedium–high
University sharingHigher utilisation and lower unit costActive institutions; job success; publications; cost per research outputDominance by large universitiesRing-fenced capacity and transparent queue rulesHigh
VouchersImmediate affordability improvementRedemption; user income distribution; learning or research outcomesProviders capture subsidy through higher pricesMultiple-provider eligibility; price ceilings; randomised or phased evaluationHigh
Tax creditsAccelerated investmentIncremental capacity and verified investment additionalityWindfall for already-planned projectsBaseline tests, caps, expiry and recaptureMedium
Cooperative purchasingLower prices and contract complexityDiscount against comparable standalone contracts; switching rateExcess standardisation around one vendorMulti-lot procurement and multi-provider frameworksHigh
Open-weight supportLocal control, research spillovers and language inclusionReproducible benchmarks; adoption; language coverage; documented safety testingFunding low-use models; release-related misuseMilestones, model cards, security review and usage evaluationHigh
PortabilityLower switching costs and stronger competitionMigration time, egress cost and successful export rateSuperficial formal complianceConformance testing and machine-readable export standardsHigh
Price disclosureComparable effective costsPercentage of offers reporting standard workload pricesGaming benchmark workloadsMultiple representative workloads and audit rightsHigh
Competition unitBetter ability to identify foreclosure and exclusionInvestigations completed; remedies; measured customer savingsOver-deterrence of integration and investmentEffects-based assessment and judicial reviewHigh
Merger scrutinyPreservation of future competition and input accessCases reviewed; remedies monitored; post-merger price and innovation outcomesDelayed efficient transactionsDeadlines, transparent theories of harm and targeted remediesHigh
Self-preferencing controlsNeutral access where discrimination is provenRanking/access parity; rejection and latency differentialsForced equivalence between non-equivalent servicesTechnical equivalence tests and security exceptionsMedium–high
Procurement reformLower lock-in and life-cycle costOpen-interface compliance; total switching cost; supplier diversityComplex tenders disadvantage SMEsStandard clauses, lotting and procurement assistanceHigh
Grid investmentRemoves physical constraint and limits price shocksConnection lead time; firm capacity; emissions and reliabilityCosts socialised while benefits concentrateBeneficiary connection charges and location signalsMedium, because of long lead times
Semiconductor supportGreater resilience and technological capacityQualified output, yield, private co-investment and customer diversityInternational subsidy race and stranded plantsMilestone disbursement, open access and clawbacksMedium
HBM diversificationReduced single-point dependencyQualified HBM capacity, yield and supplier countTechnology becomes obsolete before scaleStaged funding linked to customer qualificationMedium–low
Repair rightsLower lifetime cost and e-wasteRepairability score; parts availability; device lifeSafety or cybersecurity risks from unsupported componentsQualified repair rules and firmware-security obligationsHigh
Refurbished certificationTrustworthy lower-cost supplyFailure rate, warranty claims, resale price and device lifeCertification burdens small refurbishersTiered fees and public testing supportHigh
AI entitlementMinimum educational and research capabilityEligible-user access, task success and attainment gapVendor dependence; displacement of teachingProvider plurality, privacy requirements and pedagogy evaluationMedium–high
International fundInfrastructure and capability in underserved countriesAdditional reliable compute, trained users, local-language coverage and public-service outcomesDebt burden, imported systems without local capacityGrant-heavy finance, maintenance reserves and local trainingMedium

Table metadata: Unit: qualitative feasibility plus programme-specific KPIs. Currency and price basis: not applicable; costs are in the preceding table. Reference year: assessment as of August 2026. Scope: international, EU and national implementation. Source: official legal and programme sources cited herein; analytical assessment by this report. Method: institutional-feasibility and market-failure screening. Uncertainty: qualitative, approximately one feasibility category in either direction.

10.4.1 Public and shared compute

Public compute should operate as a capacity wholesaler with public-interest obligations, not as a permanently free retail cloud. Universities, accredited researchers, start-ups and privacy-sensitive public institutions should receive an initial allocation based on project merit and inability to acquire equivalent resources. Larger commercial users should normally pay a transparent cost-recovery tariff. Allocation should combine reserved public-interest capacity, competitive grants and spot access for otherwise idle resources. The accounting system must publish the fully loaded cost per accelerator-hour and per completed reference workload, including energy, maintenance, staff and depreciation. An apparently inexpensive public system can otherwise conceal low utilisation or unfunded replacement liabilities.

The EuroHPC AI Factory call provides a useful empirical scale. The Union made up to EUR 400 million available for new or upgraded AI-optimised supercomputers, supporting total acquisition costs of as much as EUR 800 million, with further Horizon Europe support for AI Factory services. EuroHPC Joint Undertaking launches AI Factories calls – EuroHPC Joint Undertaking – September 2024official call announcement. These amounts show why shared procurement is economically preferable to duplicating underutilised departmental clusters. They do not, however, prove that every funded installation generates net social benefit. Publication of queue distributions, job-failure rates, workload outcomes and capacity allocation is essential.

10.4.2 Vouchers and access entitlements

Compute vouchers should be denominated in quality-adjusted service units rather than tied to a named supplier. A student entitlement could comprise a protected baseline of inference, document analysis, coding and accessibility services, with higher allocations for technical courses or disability accommodation. Research grants should support verified workloads, including local deployment where privacy, intellectual property or research reproducibility justifies it. Voucher design should avoid a simple reimbursement percentage, which favours users who can first finance the remaining expenditure. Prepaid credits, institutional brokerage and income-sensitive co-payments are more inclusive.

A voucher programme should be suspended or redesigned when more than 25% of its value is captured by price increases relative to untreated customers, when fewer than three substitutable providers satisfy the standard, or when measured educational or research outcomes do not improve after two evaluation cycles. AI access should become an entitlement only to a minimum adequate capability, not an unlimited right to the most expensive frontier model. The entitlement must remain technologically neutral and capable of being fulfilled by shared public systems, competitive cloud services, approved open-weight deployments or hybrid delivery.

10.4.3 Competition, portability and pricing transparency

Competition enforcement must distinguish bottleneck control from competitive success. Authorities should collect transaction-level evidence on discounts, capacity reservations, supply refusals, cloud credits, egress charges, interoperability failures, ranking, bundling and discrimination between affiliated and independent downstream services. A high HHI or margin may justify scrutiny but cannot establish an infringement without a defined market, a theory of harm and evidence of effects.

Merger review should examine complementary assets rather than only horizontal overlap. A small acquisition may materially affect competition when the target controls a compiler, model-optimisation layer, dataset, interconnect technology, capacity broker or distribution channel required by competing providers. The Commission is reviewing its merger guidelines to maintain an evidence-based framework adapted to changing competitive conditions. Review of the Merger Guidelines – European Commission, DG Competition – 2026official review page.

Standardised price disclosure should require providers to report:

  • effective price per million input and output tokens;
  • cached-input treatment;
  • context-length and reasoning premiums;
  • storage, retrieval and tool-use charges;
  • reserved-capacity commitments;
  • service-level guarantees;
  • egress and migration costs;
  • prices for a set of audited reference workloads;
  • model-version retention rules;
  • unilateral price-change notice periods.

The disclosure standard should not impose a regulated price. Its purpose is to expose the effective total price and reduce information-driven lock-in.

10.4.4 Repair, refurbishment and equipment life

Repair and verified refurbishment are distribution policies as well as environmental measures. They expand the secondary supply of capable hardware, reduce the annualised ownership cost and allow schools or small organisations to obtain higher-memory systems than new-device budgets permit. Certification should test memory, thermal stability, storage health, power delivery, firmware integrity and sustained AI workload performance. Certified devices should carry a minimum warranty and disclose prior enterprise or high-duty-cycle use.

The EU Directive promoting repair was adopted in June 2024 and applies through national implementation from 31 July 2026. Directive on Repair of Goods – European Commission – June 2024official directive portal. AI-relevant computers and components should be incorporated through product-specific repairability, parts and firmware-support rules where the legal scope permits. Requirements must not oblige manufacturers to support unsafe configurations indefinitely, but firmware withdrawal should not become an artificial method of shortening useful life.

10.4.5 Semiconductor, packaging, memory and energy policy

Industrial subsidies should concentrate on bottlenecks that plausibly remain binding after demand and technology uncertainty are considered. Advanced packaging, HBM qualification, substrates, interposers, power electronics and test capacity may generate greater near-term resilience per public euro than attempting to duplicate every layer of leading-edge fabrication. The European Chips Act entered into force in September 2023 and expressly targets design, manufacturing and advanced-packaging capacity, resilience and reduced dependencies. European Chips Act – European Commission – September 2023official policy page.

Subsidies must include:

  • measurable capacity and yield milestones;
  • private co-investment;
  • non-discriminatory customer access;
  • supply-continuity obligations during shortages;
  • restrictions on relocating subsidised assets;
  • repayment for non-delivery;
  • workforce-development commitments;
  • disclosure of other public assistance;
  • environmental and grid-connection conditions.

Grid investment is a prerequisite rather than an AI subsidy. Data centres should pay cost-reflective connection charges and face locational signals where capacity is scarce. Public authorities should not reserve subsidised electricity for low-value or highly relocatable workloads while households and productive industry bear congestion costs. New capacity should therefore be evaluated by social value per constrained megawatt, reliability requirements, flexibility and recoverable heat or ancillary-grid services.

10.4.6 Lower-income countries and international development

The recommended international fund should finance more than imported accelerators. A viable package includes reliable electricity, cooling, connectivity, cybersecurity, local technical staff, foreign-exchange risk protection, educational access and long-term maintenance. Countries without sufficient demand to utilise a national cluster should receive federated access to regional systems, negotiated cloud capacity and latency-sensitive edge infrastructure. Grant finance is preferable for education, research and basic public services; revenue-generating commercial facilities can use concessional debt or guarantees.

The World Bank’s development framework treats compute, connectivity, data and skills as complementary foundations rather than interchangeable inputs. Digital Progress and Trends Report 2025: Strengthening AI Foundations – World Bank – 2025official publication. The ITU likewise documents persistent differences in connectivity and digital adoption. Facts and Figures 2025 – International Telecommunication Union – 2025official statistics portal. Funding accelerators without those complements can create expensive, underutilised installations and renewed dependence on foreign maintenance and software.


10.5 Policy portfolio by stakeholder

StakeholderFirst-priority actionSecond-priority actionAction to avoid
National governmentsShared public-interest compute plus grid planningTargeted vouchers and repair/refurbishment standardsUntargeted hardware subsidies without access obligations
European institutionsInteroperable EuroHPC/AI Factory networkCross-border capacity allocation and bottleneck investmentFragmented national clouds unable to exchange workloads
Competition authoritiesSpecialist technical and economic investigation capacityComplementary-asset merger review and switching-cost evidenceTreating concentration alone as proof of illegality
UniversitiesFederated procurement, workload scheduling and cost accountingMinimum student/research access allocationNumerous low-utilisation departmental clusters
Research organisationsReproducible benchmarks and open research artefactsPrivacy-aware hybrid infrastructureSelecting models solely by parameter count
Large companiesWorkload-level ROI gates and multi-provider architecturePortability tests and internal meteringBlanket AI deployment based on strategic fashion
SMEsCooperative purchasing and small-model/RAG specialisationContractual spending ceilings and exit plansFrontier models for tasks that bounded systems complete adequately
SchoolsTeacher-governed access entitlementPrivacy, accessibility and learning-outcome evaluationReplacing instruction with uncontrolled chatbot use
Civil societyIndependent access audits and price monitoringSupport for vulnerable groups and minority languagesAssuming free tiers provide capability equivalence
Lower-income countriesRegional infrastructure and negotiated capacity poolsLocal skills, language resources and maintenance reservesDebt-financed prestige clusters without utilisation demand
Development institutionsGrant-heavy foundational infrastructureOpen procurement and outcome-based disbursementVendor-tied finance that reproduces lock-in

Table metadata: Unit and currency: not applicable. Reference year: 2026. Scope: institutional strategy. Source: synthesis of Chapters 1–9 and official sources cited in Chapter 10. Method: prioritisation by expected access benefit, implementation speed and distortion risk. Uncertainty: qualitative.


10.6 Quantitative decision rules

Policy should be triggered by observable conditions rather than general concern. Let:

AAIg,t = Annual cost of adequate AI access for group g at time t / Median disposable resources of group g at time t

CapabilityGapg,t = SuccessRatefrontier,t − SuccessRateaccessible,g,t

QueueStresst = Median waiting time for eligible public-interest workload / Maximum operationally acceptable waiting time

PortabilityFailuret = Failed or materially incomplete migrations / Attempted audited migrations

Intervention thresholds should include:

  • student or low-income AAI above 5%: targeted voucher review;
  • AAI above 10%: presumptive access-support eligibility;
  • audited capability gap above 20 percentage points on essential educational or research tasks: capability-floor intervention;
  • public research queue stress above 1.5 for two consecutive quarters: capacity procurement or allocation reform;
  • portability failure above 10%: compliance investigation;
  • provider switching cost above 20% of annual AI expenditure: contract and interoperability review;
  • a verified effective-price increase exceeding consumer inflation by 15 percentage points without documented model or service improvement: enhanced disclosure and market inquiry;
  • critical input lead time above 12 months combined with capacity utilisation above 90%: resilience investment assessment.

These thresholds are decision rules, not claims that intervention will automatically generate net benefit. Each programme still requires a counterfactual and cost-benefit test.


10.7 Explicit answers to the report’s central questions

1. Has AI-relevant hardware genuinely become less affordable?

For important user groups and configurations, yes; for hardware as a homogeneous category, the evidence does not support a universal answer. Nominal entry prices, high-memory requirements, complementary system costs and income differences can reduce affordability even when capability per unit of arithmetic performance improves. Frontier-equivalent local deployment is substantially less affordable than merely running a small quantised model.

2. Is the increase nominal, quality-adjusted or both?

The increase is clearly nominal in selected high-demand categories and periods. It is also quality-adjusted where memory, bandwidth, reliability or completed-workload costs deteriorate relative to the requirement. It is not generally quality-adjusted when the denominator is theoretical throughput alone. The conclusion changes with the quality measure, which is why price per successfully completed workload is preferable to sticker price or TFLOP alone.

3. Is used-market inflation systematic?

The evidence supports episodic, product-specific scarcity, especially for high-memory, power-efficient or unusually capable legacy devices. It does not establish that used AI hardware generally trades at two or three times original MSRP across countries and years. Such ratios should be reported as verified observations, not as the median market condition, until transaction-level data establish representativeness.

4. Which bottlenecks are most serious?

The highest systemic significance lies in leading-edge fabrication, advanced lithography, HBM, advanced packaging, substrates and interposers, high-speed interconnects, power equipment, grid connections and specialised software ecosystems. Their criticality reflects both concentration and limited short-run substitution.

5. Does concentration produce demonstrable pricing power?

It produces the capacity and incentive for pricing power, but concentration and profitability alone do not prove its exercise or illegality. Demonstration requires evidence of persistent margins above a defensible counterfactual, customer discrimination, capacity withholding, foreclosure, non-cost-based switching barriers or price responses inconsistent with innovation and risk.

6. Are current AI-service prices economically sustainable?

Basic and highly utilised inference services may be sustainable. Frontier subscriptions, long-context reasoning and aggressively priced promotional services remain uncertain because public filings generally do not isolate complete model-level economics. Provider-wide investment and depreciation imply that current posted token prices cannot automatically be treated as long-run equilibrium prices.

7. How much could enterprise AI expenditure rise by 2031?

Using Chapter 9’s indexed scenarios, aggregate enterprise AI expenditure rises from 100 in 2026 to approximately:

  • 185 in the low case: +85%;
  • 300 in the central case: +200%;
  • 590 in the high case: +490%.

These are expenditure forecasts, not pure unit-price forecasts. They combine adoption, usage intensity, model mix, governance, infrastructure and labour costs.

8. Can local AI remain realistic?

Yes, for bounded, privacy-sensitive and repeatable workloads. It is less realistic as a universal substitute for continuously updated frontier services, especially where many concurrent users, very long contexts or near-frontier multimodality are required.

9. What constitutes genuinely usable local AI?

A system is usable only when it meets a workload-specific threshold for:

  • task success and factual reliability;
  • sufficient context and usable memory;
  • acceptable tokens per second;
  • bounded time to first token;
  • required concurrency;
  • operational uptime;
  • privacy and auditability;
  • manageable energy cost;
  • maintainable software and model updates.

Loading weights into memory is necessary but not sufficient.

10. Are students and researchers without frontier access disadvantaged?

The evidence supports a credible and testable disadvantage, particularly in coding, multilingual work, document synthesis, advanced reasoning and rapid iteration. The magnitude is not yet universally established because longitudinal causal studies remain limited. The appropriate conclusion is measurable task-level disadvantage with incomplete evidence about lifetime educational effects.

11. Which countries and groups face the greatest exposure?

Those with low disposable income, weak currencies, import duties, unreliable electricity, limited broadband, long latency, no nearby cloud region, export restrictions and weak local-language support. Within countries, exposure is highest for low-income students, independent researchers, rural users, small organisations, minority-language communities and privacy-sensitive institutions without infrastructure budgets.

12. Could AI deepen global inequality?

Yes. AI complements skilled labour, data, capital and organisational capacity, allowing initially advantaged users to compound gains. Chapter 9’s central scenario increased the illustrative AI-access Gini from 0.193 in 2026 to 0.215 in 2031; the divergence case reached approximately 0.290. Conversely, low-cost models and shared access can reduce inequalities if capability floors improve faster than frontier costs rise.

13. Is the “AI priced like Bitcoin” analogy defensible?

Only narrowly. Compute can be scarce, metered, reserved and exposed to demand-sensitive pricing. Unlike Bitcoin, however, compute is heterogeneous, depreciating, location-dependent, perishable when idle, continuously produced and inseparable from memory, energy, networking and software. It is therefore better understood as a differentiated capacity service or electricity-like industrial input than as a standard store of value.

14. Which interventions expand access without suppressing innovation?

The strongest portfolio combines:

  • shared public-interest compute;
  • targeted student and research vouchers;
  • cooperative purchasing;
  • portability and open interfaces;
  • standard price disclosure;
  • open-weight and local-language research;
  • repair and verified refurbishment;
  • evidence-based competition enforcement;
  • grid and bottleneck investment;
  • competitively awarded, milestone-based industrial support.

These measures enlarge demand and contestability without imposing general price controls.

15. Under what conditions would the report’s warning be proven wrong?

The central warning would be materially falsified if, by 2031:

  • quality-adjusted cost per completed representative workload fell by at least 70% from 2026;
  • the Global AI Access Gini fell to 0.15 or below;
  • adequate-access cost fell below 3% of median disposable resources in lower-income countries and student populations;
  • the frontier-versus-basic capability gap fell below 15 percentage points on essential tasks;
  • median lead time for critical accelerators, HBM and packaging fell below six months;
  • audited provider migration succeeded in at least 95% of cases without material functional loss;
  • no strategically essential layer maintained both HHI above 2,500 and persistent supranormal margins unsupported by innovation, risk or capital recovery;
  • at least 90% of accredited students and researchers obtained adequate, reliable AI access.

10.8 Final risk classification framework

ClassificationAI-access GiniAdequate-access affordability burden for exposed groupsEssential-task capability gapStructural supply condition
LimitedBelow 0.10Below 3%Below 10 percentage pointsCritical lead times below 6 months; broad substitution
Moderate0.10–0.153–5%10–20 pointsOccasional bottlenecks; no persistent multi-layer scarcity
HighAbove 0.15–0.22Above 5–10%Above 20–35 pointsPersistent concentration or 9–12 month lead times in critical layers
SevereAbove 0.22–0.30Above 10–20%Above 35–50 pointsCR4 above 80% or HHI above 2,500 across at least three complementary layers, with lead times above 12 months
SystemicAbove 0.30Above 20%Above 50 pointsMulti-region access failure, sustained rationing or exclusion of more than 25% of qualified users from essential workloads

Table metadata: Unit: ratios, percentage of disposable resources, percentage-point benchmark gap and market-structure indicators. Currency: not applicable. Price basis: adequate annual access cost at 2026 real purchasing power. Reference year: baseline 2026, assessment horizon 2031. Scope: global and exposed-user populations. Source: thresholds defined by this report using results from Chapters 3, 7, 8 and 9. Method: multidimensional risk-classification rule; the highest persistent dimension determines the provisional classification, subject to corroboration by at least one additional dimension. Uncertainty: classification uncertainty of approximately one category because access and benchmark datasets remain incomplete.

Final classification: HIGH, with a material upward risk toward SEVERE

The current evidence exceeds the high-risk thresholds through the estimated access Gini, affordability burdens in exposed populations, substantial capability differences between unrestricted frontier access and constrained local or free-tier access, and concentration across several complementary supply-chain layers. It does not yet justify a global severe or systemic classification because falling quality-adjusted costs, expanding capacity, open-weight innovation and public infrastructure remain credible counterforces; universal exclusion has not been demonstrated; and several critical indicators rely on estimates rather than harmonised transaction data. Under the managed-oligopoly central scenario, the risk remains high through 2031. Under persistent scarcity or geopolitical fragmentation, the access Gini and capability gap can enter the severe range. The policy implication is preventive rather than punitive: intervention should create capacity, contestability and a minimum capability floor before deprivation becomes structurally self-reinforcing.


10.9 Required-table completion and audit register

Required tablePrimary chapterPublication requirement
Global AI value-chain map1 and 3Consolidate firms, layers, dependencies and geography
Historical GPU price database2Retain product-level observations and exact dates
Historical memory-price database2Separate HBM, DRAM and NAND
Official, street and used prices2Preserve observation date and condition
Quality-adjusted hardware comparison2Publish every denominator and benchmark version
Supply-chain market shares3State market definition and share year
CR4 and HHI by segment3Preserve authority-specific interpretation
Five-year local-system TCO4Include discount rate and residual value
Model-size and memory matrix4Separate weights, KV cache, activations and overhead
Tokens per second and cost per million tokens4 and 5Report framework, quantisation and workload
Cloud-provider pricing5Timestamp every tariff
Provider cost stack5Separate observed values from inferred allocations
Enterprise cost by firm size6Publish archetype assumptions
Sectoral exposure matrix6Include regulation, privacy and failure cost
Student and researcher affordability7State income and resource denominator
Wage months for local hardware7Use net or disposable wage consistently
Country affordability table8Apply PPP and market exchange rates separately
Global AI Access Index8Publish raw indicators, weights and normalisation
Scenario assumptions and outputs9Separate scenario logic from forecasts
Annual forecasts, 2026–20319Include low, central, high and intervals
Sensitivity analysis8 and 9Report weighting and model alternatives
Early-warning dashboard9 and 10Assign trigger, frequency and responsible authority
Policy cost-benefit matrix10Included above; requires jurisdiction-specific refinement
Facts, estimates and projectionsAll chaptersFinal classification table below
Unresolved data gapsAll chaptersMaintain as a living audit register

Mandatory metadata for every final-publication table: unit; currency; nominal or real price basis; reference year; observation date; geographical scope; primary source; calculation method; sample size where relevant; uncertainty interval or confidence classification. Analytical tables appearing in individual chapters should be reissued in a consolidated statistical appendix so that metadata are not lost during publication formatting.


10.10 Evidence-status register

FindingStatusConfidencePrincipal unresolved requirement
AI capacity requires complementary hardware, energy, software and skillsConfirmed structural factHighImproved cross-layer capacity data
Selected high-memory and frontier systems show large nominal costsConfirmed for cited productsHighHarmonised worldwide street-price history
All AI hardware became quality-adjusted more expensiveNot confirmedLowConstant-workload hedonic panel
Used devices generally sell for two or three times MSRPNot confirmed as representativeLowTransaction-level marketplace data
Supply is highly concentrated in several critical layersConfirmedHighConsistent global market definitions
Concentration alone proves abusive pricingRejected inferenceHighConduct and counterfactual price evidence
Local AI is viable for bounded workloadsConfirmed conditionallyMedium–highStandardised task and reliability tests
Local AI universally substitutes for frontier cloud servicesNot confirmedHigh confidence in rejectionRapid change in model efficiency could alter result
Current frontier-service prices are fully cost-reflectiveUnresolvedLowProvider model-level cost and revenue accounts
Enterprise AI expenditure could triple by 2031Central projectionMedium–lowUsage, price and governance-cost observations
Differential access can compound educational advantageSupported mechanismMediumLongitudinal causal studies
Global access inequality rises under the central scenarioProjectionMedium–lowAnnual Global AI Access Index observations
Public/shared compute can improve accessSupported policy estimateMediumProgramme-level counterfactual evaluation
Targeted vouchers outperform universal subsidiesPolicy inferenceMediumRandomised or phased programme evidence
Compute will become Bitcoin-like moneyEconomically unsupportedHighNo plausible fungibility or store-of-value mechanism identified

Table metadata: Unit: evidence classification. Currency and price basis: not applicable. Reference date: August 2026. Geographic scope: global. Source: Chapters 1–10 and live-verified primary sources. Method: separation of directly observed facts, model estimates, forecasts and unresolved propositions. Uncertainty: categorical confidence stated in the table.

Final judgment

The emerging AI divide is not primarily a shortage of intelligence expressed through a single price. It is a layered inequality in the ability to finance, locate, power, operate and productively apply systems of different quality. Hardware learning can continue while access becomes more unequal; abundant aggregate compute can coexist with scarcity for specific users, regions, memory configurations and privacy requirements. On the evidence and thresholds established in this report, the present risk is HIGH. Without faster diffusion of shared capacity, effective portability, expanded energy and packaging supply, trustworthy secondary hardware markets and targeted educational access, the risk approaches SEVERE under persistent-scarcity or fragmented-world conditions. It becomes SYSTEMIC only if access inequality exceeds the defined thresholds and essential AI capability ceases to be realistically obtainable by a substantial share of qualified students, researchers, firms or countries.


Copyright of debugliesintel.com
Even partial reproduction of the contents is not permitted without prior authorization – Reproduction reserved

latest articles

explore more

spot_img

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Questo sito utilizza Akismet per ridurre lo spam. Scopri come vengono elaborati i dati derivati dai commenti.