Introduction — From Universal Digital Service to Rationed Computational Capability
Master Abstract
Artificial intelligence is frequently presented as an infinitely replicable software technology, yet every economically useful output depends on a finite and unusually concentrated physical system: leading-edge semiconductor fabrication, advanced packaging, high-bandwidth memory, accelerators, networking equipment, data-centre capacity, electricity, cooling infrastructure and specialised engineering labour. The resulting market cannot be evaluated through consumer subscription prices alone. Its true structure begins upstream, where technical and financial barriers restrict the number of firms capable of manufacturing the most advanced components, and continues downstream, where a small group of hyperscalers and model developers determine which users receive access, at what performance level, with which latency, context window, privacy safeguards and contractual limitations. Evidence available in August 2026 confirms an extraordinary expansion of demand and investment, but it does not yet prove that scarcity has been artificially engineered or that unlawful pricing conduct has occurred. NVIDIA reported fiscal-year 2026 revenue of $215.9 billion, an increase of 65%, with a fiscal-year GAAP gross margin of 71.1%; in the quarter ended 26 April 2026, Data Center revenue reached $75.2 billion, 92% above the corresponding prior-year period. These results demonstrate exceptional value capture in a supply-constrained technological segment, but margins and revenue growth must subsequently be decomposed into innovation rents, intellectual-property returns, scarcity rents, product-mix effects and possible market power. NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026 – NVIDIA – February 2026 NVIDIA Announces Financial Results for First Quarter Fiscal 2027 – NVIDIA – May 2026
The upstream constraint is not reducible to GPU design. AI accelerators require leading-edge logic, large quantities of HBM, complex interposers and advanced packaging technologies capable of connecting logic dies with multiple memory stacks at extreme bandwidth. TSMC reported that the annual capacity of facilities managed by the company and its subsidiaries exceeded 17 million 12-inch-equivalent wafers in 2025, while its official 2025 first-quarter investor transcript stated that the company was working to double CoWoS capacity during that year in response to customer demand. TSMC has separately announced that a 9.5-reticle-size CoWoS implementation planned for volume production in 2027 is intended to integrate twelve or more HBM stacks in one package. These disclosures show how frontier AI performance increasingly depends on co-optimisation across fabrication, memory and packaging, rather than on an isolated processor. Memory producers are consequently moving investment toward higher-value AI products: Micron projected in December 2025 that the HBM total addressable market could rise from approximately $35 billion in 2025 to around $100 billion in 2028, while Samsung announced plans to invest more than KRW 110 trillion in facilities and research and development during 2026, explicitly including HBM, foundry production and advanced packaging. These figures establish strong demand and rapid capacity formation, but they do not by themselves demonstrate permanent scarcity. The central analytical problem for this report will therefore be whether additions to fabrication, packaging and memory capacity can outrun the combined growth of training, inference, multimodal services, agentic workloads and national sovereign-compute programmes. TSMC 2025 Annual Report – Taiwan Semiconductor Manufacturing Company – April 2026 TSMC First Quarter 2025 Earnings Conference Transcript – Taiwan Semiconductor Manufacturing Company – April 2025 TSMC Unveils Next-Generation A14 Process at North America Technology Symposium – Taiwan Semiconductor Manufacturing Company – April 2025 Micron Fiscal First Quarter 2026 Financial Results – Micron Technology – December 2025 Samsung Electronics Shareholder Value Enhancement Plan – Samsung Electronics – March 2026
The downstream investment cycle indicates that today’s apparently inexpensive access to powerful models cannot automatically be treated as a stable long-term equilibrium. Alphabet reported $91.4 billion in 2025 capital expenditure, compared with $52.5 billion in 2024, and stated that it expected a significant further increase in technical-infrastructure investment during 2026; capital expenditure reached $35.7 billion in the first quarter of 2026 alone, compared with $17.2 billion in the corresponding 2025 period. This investment does not prove that consumer or API prices must rise, because higher utilisation, algorithmic efficiency, model compression, custom accelerators and competition can reduce unit costs. It does, however, establish that AI services are supported by a capital base whose depreciation, financing, electricity, network and replacement costs must eventually be recovered through advertising, subscriptions, API billing, enterprise contracts or cross-subsidisation. Energy is becoming a parallel constraint. The International Energy Agency projects global data-centre electricity consumption of approximately 945 TWh in 2030, more than twice the current level represented in its base case and just under 3% of projected global electricity consumption. The IEA further estimates average data-centre electricity-demand growth of about 15% annually between 2024 and 2030, while AI-optimised facilities account for the principal incremental driver. The report must consequently model not merely a price per token, but a multidimensional tariff in which model capability, input and output length, reasoning intensity, latency, reserved capacity, geographic location, energy availability and privacy requirements become separate price determinants. The Bitcoin analogy will be tested rather than accepted: compute can be scarce and dynamically priced, but unlike Bitcoin it is heterogeneous, depreciating, geographically constrained, continuously producible and consumed in delivering a service. Alphabet Annual Report on Form 10-K for the Fiscal Year Ended December 31, 2025 – Alphabet and U.S. Securities and Exchange Commission – February 2026 Alphabet Quarterly Report on Form 10-Q for the Quarter Ended March 31, 2026 – Alphabet and U.S. Securities and Exchange Commission – April 2026 Energy and AI: Energy Demand from AI – International Energy Agency – April 2025
For households, students, researchers and smaller organisations, the decisive divide is not between possessing and lacking nominal AI access, but between different levels of quality-adjusted computational capability. A used workstation may run a quantised model locally, yet remain inadequate for long contexts, multimodal processing, high concurrency, rapid generation or specialised scientific workloads. The RTX 3090, introduced at a starting price of $1,499, contains 24 GB of GDDR6X memory; this remains useful for defined local workloads but imposes a hard memory boundary before software overhead, context cache and concurrent inference are considered. A 2025 Mac Studio with M4 Max can be configured with 128 GB of unified memory, enabling a much larger theoretical model-memory envelope, although unified-memory capacity must not be equated with NVIDIA accelerator performance, software compatibility or effective application throughput. GeForce RTX 3090 Family Specifications – NVIDIA – September 2020 Mac Studio 2025 Technical Specifications – Apple – March 2025 The international distributional problem begins even earlier: the International Telecommunication Union estimated that 94% of people in high-income countries used the internet in 2025, compared with only 23% in low-income economies, while 2.2 billion people remained offline. Advanced AI therefore arrives on top of an unresolved connectivity, affordability, electricity, device and skills divide. Measuring Digital Development: Facts and Figures 2025 – International Telecommunication Union – November 2025 Europe’s response illustrates the strategic scale of the issue: the European Commission launched a July 2026 call intended to establish as many as seven AI Gigafactories and unlock more than €30 billion, while its architecture defines such facilities as infrastructures bringing together more than 100,000 advanced AI processors. EU Launches AI Gigafactories Call to Boost Europe’s Computing Capacity and Unlock More Than €30 Billion – European Commission – July 2026 AI Factories – European Commission – July 2026 The five-year investigation will determine whether this investment broadens access or merely relocates concentrated capacity, and whether AI becomes a productivity equaliser, a metered industrial utility or a cumulative mechanism through which capital-rich users acquire an enduring cognitive and economic advantage.
AI compute, affordability and access
Navigate the ten analytical layers, stress-test five scenarios and change policy intensity without converting estimates into observed facts.
Annual trajectory
Central, low and high paths
Critical focal points and falsification thresholds
The classification uses the highest persistent risk dimension corroborated by at least one additional dimension.
Evidence integrity ledger
| Finding | Status | Confidence | Interpretive constraint |
|---|
The Price of Intelligence: How Compute Is Becoming the New Infrastructure of Economic Power
Artificial intelligence is no longer principally a contest between algorithms. It is a contest over access to accelerators, high-bandwidth memory, advanced packaging, electricity, data centres and the capital required to assemble them into usable systems. The decisive inequality is therefore shifting from internet connectivity to computational capability: who can train, adapt and operate advanced models; who must rent that capability; and who remains confined to weaker systems. This transition has consequences for industrial productivity, scientific sovereignty, education, defence and the distribution of economic power. Europe has begun constructing a public response, but the strategic question remains unresolved: whether AI compute will develop as broadly accessible infrastructure or as a scarce, vertically controlled service whose price and availability are determined by a small group of technology providers.
The New Economic Input
Compute cannot be classified simply as another industrial input. It is simultaneously a capital good, a metered service, a component of critical infrastructure and an instrument of geopolitical leverage. Its economic value depends not only on processor throughput but also on accelerator memory, bandwidth, interconnects, software compatibility, electricity, cooling, network latency and utilisation. A nominally powerful accelerator without adequate memory or software support may be economically inferior to a slower system capable of completing the required workload reliably.
This distinction is essential when assessing claims of hardware inflation. Sticker prices alone cannot establish whether AI capability has become more expensive. A valid comparison must determine the cost of completing the same workload under equivalent requirements for accuracy, context length, latency, concurrency, privacy and reliability. Semiconductor improvements can reduce the cost of an arithmetic operation while the total cost of deploying a useful model rises because the workload requires more memory, more energy, specialised personnel or a larger cluster.
The result is an apparent paradox: aggregate computing power can become more abundant while strategically useful compute remains scarce. Scarcity becomes especially acute when organisations require frontier-level performance, sovereign data handling, guaranteed capacity or local operation. For a student, an independent researcher or an SME, the relevant barrier is not whether a small model can technically run on a computer. It is whether that system can produce competitive work within acceptable time, quality and reliability constraints.
Adoption Without Equal Access
European enterprises are already entering this transition from profoundly unequal starting positions. Eurostat reported that 19.95% of EU enterprises employing at least ten people used one or more AI technologies in 2025, while the corresponding share among large enterprises reached 55.03%. Use of Artificial Intelligence in Enterprises – Eurostat – 2025 — official statistical publication.
The disparity matters because AI adoption is cumulative. Organisations with data, engineers, cloud agreements and integration budgets can test more applications, train employees, construct proprietary workflows and learn from failure. Smaller organisations must absorb licences, cybersecurity, compliance, data preparation, integration and human-supervision costs before productivity gains are certain. AI can consequently become a de facto mandatory input: firms may have to purchase it because competitors use it, even when their own return on investment remains difficult to measure.
This creates a structural asymmetry. Large companies can negotiate capacity reservations, diversify providers and finance local infrastructure. SMEs frequently buy retail access under standard contracts and carry greater switching costs relative to revenue. If model capability, latency and access priority are increasingly differentiated by price, the market will not merely divide users into subscribers and non-subscribers. It will divide them according to the quality, speed, context and reliability of the intelligence they can afford.
Europe’s Infrastructure Response
The European Commission’s AI Continent Action Plan, presented on 09/04/2025, placed physical infrastructure at the centre of European policy. The Commission identified EUR 200 billion to be mobilised for AI investment, including EUR 20 billion intended to finance as many as five AI gigafactories. It also identified a network of 19 AI Factories supporting start-ups, industry and research. AI Continent – European Commission – 09/04/2025 — official programme page.
That network has since been complemented by 13 AI Factory Antennas, while the European High Performance Computing Joint Undertaking states that the AI Factories offer free, customised support to SMEs and start-ups. AI Factories – European High Performance Computing Joint Undertaking – 2026 — official infrastructure portal. On 30/07/2026, EuroHPC launched its AI Gigafactories call, defining the planned facilities as integrated environments combining AI-optimised supercomputers, advanced data centres, high-capacity storage, high-speed networks, secure cloud access and specialist support. The EuroHPC Joint Undertaking Launches the AI Gigafactories Call – EuroHPC Joint Undertaking – 30/07/2026 — official call announcement.
The architecture is strategically correct because it recognises that purchasing accelerators is insufficient. Europe needs an operating ecosystem able to allocate capacity across borders, support industrial users, protect sensitive data and convert computational resources into deployable applications. The danger is fragmentation. National installations that cannot exchange workloads, use compatible interfaces or provide predictable access could reproduce at public expense the same lock-in that Europe seeks to reduce.
The Energy Constraint
AI infrastructure is also energy infrastructure. On 20/12/2024, the United States Department of Energy reported that American data centres consumed 176 terawatt-hours of electricity in 2023, equivalent to approximately 4.4% of national electricity consumption. The Department’s cited range for 2028 was 325–580 terawatt-hours, or approximately 6.7–12% of US electricity consumption. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – United States Department of Energy – 20/12/2024 — official publication.
These figures are not a forecast for Europe, but they expose the physical scale of the challenge. Data-centre policy cannot be separated from generation capacity, transmission networks, substations, transformers, cooling water, permitting and local-system reliability. A country may possess capital and technical expertise yet remain unable to deploy new AI capacity within commercially useful timescales because grid connections are unavailable.
Public policy must therefore distinguish between socially valuable infrastructure and private congestion costs. Data centres should contribute to the network investments required by their load, while public authorities should assess flexibility, emissions, location and the economic value of workloads. Subsidising compute without financing the electrical system that supports it would move the bottleneck rather than remove it.
Memory as Geopolitics
High-bandwidth memory demonstrates how a specialised component can become both an industrial bottleneck and a security instrument. On 02/12/2024, the United States Department of Commerce’s Bureau of Industry and Security announced controls covering 24 types of semiconductor-manufacturing equipment, three categories of software tools, HBM and 140 additions to the Entity List. BIS explicitly described HBM as critical to AI training and inference at scale. Commerce Strengthens Export Controls to Restrict China’s Capability to Produce Advanced Semiconductors – Bureau of Industry and Security – 02/12/2024 — official announcement.
The decision illustrates why compute cannot be treated as a globally fungible commodity. Its availability depends on origin rules, export licences, technical thresholds, fabrication geography and political alignment. Unlike Bitcoin, compute is heterogeneous, location-dependent and depreciating. Unused capacity is perishable; old hardware loses relative value; and two nominally equivalent systems can produce different economic results because of memory, software or network constraints.
The European Chips Act, which entered into force on 21/09/2023, was designed to reinforce the Union’s semiconductor ecosystem, improve supply-chain resilience and reduce external dependencies. Its stated objectives include research leadership, design, manufacturing, advanced packaging, production capacity and skills. European Chips Act – European Commission – 21/09/2023 — official policy framework. The strategic lesson is that leading-edge fabrication alone is not enough. Packaging, memory, substrates, testing, equipment, software and energy must be treated as an interdependent industrial system.
The Cost of Dependence
Cloud AI solves one access problem while creating another. It allows organisations to use advanced systems without owning expensive infrastructure, but it transfers control over pricing, capacity, model continuity and data-processing conditions to the service provider. Providers may differentiate prices according to input, output, context, latency, reasoning intensity, tool use, reserved capacity or service priority. This resembles utility pricing more than the sale of conventional software.
The danger is not metering itself. Metering can improve efficiency and transparency. The danger arises when customers cannot compare the full cost of equivalent workloads or move their data, embeddings, evaluation records and applications without material loss. The European Data Act, Regulation (EU) 2023/2854, has applied since 12/09/2025 and establishes requirements intended to facilitate switching between cloud and edge services and promote interoperability. Data Act Explained – European Commission – 15/12/2025 — official explanation.
The next regulatory step should be standardised effective-price disclosure. Providers should separate inference, storage, retrieval, networking, tool use, reserved capacity and migration charges. Public procurement should require exportable operational data, documented interfaces, model-version notice periods and tested exit procedures. Competition authorities, meanwhile, should examine self-preferencing, data access and cloud dependencies without presuming that vertical integration is inherently unlawful. The Commission’s DMA review, published on 28/04/2026, identified interoperability, potential self-preferencing, access to data and cloud dependencies among the AI-related issues raised during consultation. Commission Staff Working Document SWD(2026) 123 final – European Commission – 28/04/2026 — official document.
The Global Capability Divide
The AI divide begins before access to an accelerator. According to the International Telecommunication Union’s figures released on 17/11/2025, approximately six billion people were online in 2025, while 2.2 billion remained offline. The ITU estimated 5G coverage at 84% of the population in high-income countries, compared with 4% in low-income countries. Facts and Figures 2025 – International Telecommunication Union – 17/11/2025 — official release.
The internet-use divide is equally stark. ITU data published on 15/10/2025 recorded internet use at 94% in high-income countries but 23% in low-income economies; the corresponding figure for Africa was 36%. Facts and Figures 2025: Internet Use – International Telecommunication Union – 15/10/2025 — official statistical analysis.
AI can lower barriers by translating information, supporting education and distributing sophisticated analytical tools. Yet those benefits presuppose connectivity, electricity, compatible languages, payment capacity and digital skills. Where these foundations are absent, the arrival of more powerful models can widen the productive distance between countries rather than close it. The relevant development objective is therefore not the installation of prestige clusters. It is reliable, maintainable and affordable access linked to local skills, regional infrastructure and public-service demand.
A Capability Floor
The most effective response is not universal ownership of frontier hardware. It is a guaranteed capability floor combined with competitive provision. Students and researchers should receive workload-denominated compute entitlements usable across approved public, cloud and open-weight systems. Universities should pool procurement and publish utilisation, queue times and completed-workload costs. SMEs should gain access through cooperative purchasing and the AI Factory network. Privacy-sensitive institutions should be able to choose local, sovereign or hybrid deployment without accepting an automatic penalty in capability.
Hardware life also matters. Directive (EU) 2024/1799 requires Member States to apply the European right-to-repair framework from 31/07/2026. The Commission states that manufacturers of covered products must offer repair within a reasonable time and at a reasonable price, cannot use hardware or software techniques to obstruct repair, and must provide access to spare parts. Transposition, Entry into Force or Application of EU Legislation: Key Milestones – European Commission – 01/07/2026 — official implementation calendar. Extending useful hardware life and creating trustworthy refurbished markets can lower access costs without suppressing semiconductor innovation.
The Strategic Choice
Europe’s AI strategy will succeed only if it measures access rather than announcing capacity. The decisive indicators are not the number of installed accelerators or the nominal value of investment. They are the share of capacity reaching SMEs, universities and public-interest users; waiting times; cost per completed workload; migration success; energy availability; local-language performance; and the proportion of students and researchers able to use systems adequate for their tasks.
The emerging market will not price intelligence in the manner of a uniform commodity. It will price differentiated access to capability. Model quality, memory, context, latency, privacy and guaranteed capacity will determine what users can accomplish and what they must pay. If that architecture remains concentrated and difficult to contest, AI will magnify existing advantages in capital, education and geography. If Europe combines infrastructure, portability, competition, repair, open models and targeted access, compute can instead become the foundation of a broader productive system. The central political decision is therefore not whether to regulate or subsidise AI. It is whether advanced computational capability will remain purchasable privilege or become accessible economic infrastructure.
Chapter 1 — The Political Economy of AI Compute
1.1 Purpose, Scope and Analytical Position
AI compute is neither a single product nor a homogeneous quantity. It is a layered production capability created by combining semiconductor-design intellectual property, fabrication capacity, advanced packaging, high-bandwidth memory, servers, interconnects, storage, data centres, electricity, cooling, software frameworks, trained models and specialised labour. The economic object purchased by the final user is therefore not simply a GPU or a token. It is access to a time-bounded, workload-specific configuration of capital, energy, memory bandwidth, software compatibility and model capability. This distinction is essential because different forms of apparent scarcity can arise at different layers. A shortage of accelerator boards may reflect insufficient leading-edge fabrication, constrained advanced packaging, unavailable HBM, export restrictions, hyperscaler procurement, distributor behaviour or a temporary demand shock. A high API price may instead reflect training-cost recovery, inference expense, reserved capacity, latency guarantees, model scarcity, commercial segmentation or pricing power. The report will not infer unlawful market manipulation merely from increasing prices, exceptional margins or a concentrated supplier structure. It will determine whether prices exceed competitive economic benchmarks after controlling for product quality, capacity cycles, depreciation, input costs, demand growth, risk and innovation. The European Commission has independently identified data, AI accelerators, computing infrastructure, cloud capacity and technical expertise as possible barriers to entry in generative-AI markets, while cautioning that a bottleneck becomes a competition concern only in the relevant economic and legal context. Competition Policy Brief: Competition in Generative AI and Virtual Worlds – European Commission – September 2024
The principal proposition of this report is therefore conditional: if the supply of quality-adjusted AI compute expands more slowly than demand, if effective competition remains restricted at several complementary layers, and if access to superior systems produces cumulative educational or economic gains, then compute scarcity can become a mechanism of social and industrial stratification. The contrary outcome remains possible. Fabrication capacity, alternative accelerators, open-weight models, quantisation, distillation, inference optimisation, edge deployment and public compute infrastructure could lower the cost of a fixed level of capability even while the nominal price of premium hardware increases. Chapter 1 establishes the framework needed to discriminate between these trajectories. It defines six economically distinct interpretations of compute, maps the value chain, locates potential sources of scarcity and bargaining power, formalises the concepts used throughout the report and specifies eight falsifiable hypotheses. No hypothesis will be treated as confirmed in this chapter. In particular, claims concerning used-market prices, Apple or Samsung product inflation, the real cost of local AI and the sustainability of current token pricing require longitudinal datasets that will be constructed in Chapters 2, 4 and 5. The evidentiary status at the beginning of the investigation is consequently “open”: upstream capacity pressure, extraordinary investment and unequal adoption are already observable, while persistent real hardware inflation, systematic second-hand premiums, above-normal rents and future token-price escalation remain matters for formal testing.
1.2 The Six Economic Identities of AI Compute
AI compute can assume six economic identities simultaneously, but the relative importance of each identity changes according to the workload, purchaser and time horizon. For a software developer purchasing several hours of inference, compute resembles a variable industrial input. For a hyperscaler committing billions of dollars to accelerators, data centres and power procurement, it is a durable but rapidly depreciating capital good. When delivered continuously through standardised APIs with metering, capacity tiers and service-level commitments, it acquires utility-like characteristics. When governments rely on it for defence, intelligence, healthcare, energy-system optimisation or public administration, it becomes strategic infrastructure. When access is conditioned by export controls, domestic fabrication, allied supply chains or national cloud capacity, it becomes a geopolitical resource. Finally, when future capacity is reserved through long-term contracts, priority tiers, capacity options or transferable claims, it can acquire commodity-like and partially financialised attributes. None of these classifications is complete in isolation. Unlike electricity, compute is heterogeneous: one accelerator-hour cannot be substituted perfectly for another because memory, precision, architecture, software, networking, geography and permitted workloads differ. Unlike oil, compute is not stored and later burned; unused capacity is a perishable service flow generated by depreciating equipment. Unlike Bitcoin, it is produced through industrial investment and has no fixed algorithmic supply. Nevertheless, compute can still be rationed, reserved, traded contractually and priced dynamically, making the commodity analogy economically relevant within narrow limits.
| Economic identity | Defining characteristic | Typical purchaser | Principal price unit | Main scarcity mechanism | Most appropriate regulatory lens |
|---|---|---|---|---|---|
| Normal industrial input | Consumed in producing another good or service | Enterprise, developer, laboratory | Accelerator-hour, token, task or API call | Short-run workload demand | Input-cost pass-through and productivity |
| Scarce capital good | Requires large upfront investment and depreciates | Hyperscaler, state, large enterprise | Server, cluster, installed MW or five-year TCO | Fabrication, financing and installation lead time | Investment, depreciation and capacity formation |
| Utility-like service | Continuously delivered and metered | Household, SME, institution | Subscription, token or reserved capacity | Congestion and service availability | Transparency, reliability and nondiscrimination |
| Strategic infrastructure | Supports essential economic or state functions | Government, critical infrastructure, hospital | Sovereign capacity and service availability | Resilience, security and domestic control | Critical-infrastructure governance |
| Geopolitical resource | Access depends on jurisdiction and international alignment | State, defence firm, sanctioned or controlled entity | Licensable performance and permitted end use | Export controls and territorial concentration | Trade, security and industrial policy |
| Potentially financialised commodity | Future capacity is reserved, tiered or contracted | Hyperscaler, model provider, financial investor | Forward capacity, priority or availability commitment | Expectations of future scarcity | Market design, disclosure and systemic leverage |
The classification produces an immediate methodological consequence: there can be no universal “price of AI compute.” The price vector must instead be expressed as P = f(H, M, B, L, E, G, S, Q, R, T), where H denotes hardware architecture, M available memory, B memory and network bandwidth, L latency, E energy burden, G geography, S software ecosystem, Q model or task quality, R reliability and T contract duration. Two services charging the same amount per million tokens may have radically different effective prices if one produces more useful output, requires fewer retries, supports longer context, preserves privacy or supplies guaranteed capacity. Similarly, two computers with identical purchase prices may have different economic costs because one consumes more electricity, has lower utilisation, lacks software compatibility or becomes obsolete sooner. The report will therefore use both nominal prices and quality-adjusted prices, and it will measure access as a bundle rather than a binary condition.
1.3 Compute as an Industrial Input and Capital Good
As an industrial input, AI compute enters a production function alongside labour, data, conventional capital and organisational knowledge. Its marginal value depends not simply on how much compute is consumed, but on whether the enterprise possesses complementary assets capable of converting model output into reliable decisions or revenue. A firm without clean data, process redesign, evaluation systems, cybersecurity, legal governance and trained personnel may purchase large volumes of tokens without achieving measurable productivity gains. Conversely, a specialised small model integrated into a well-designed process may produce a greater return than a more expensive frontier model used without organisational adaptation. A general production specification for enterprise i in period t is Yi,t = Ai,t × F(Ki,t, Li,t, Ci,t, Di,t, Oi,t), where Y is output, K is conventional capital, L is labour, C is quality-adjusted compute, D is usable data, O is organisational capability and A captures other productivity determinants. Compute may complement highly skilled labour, substitute for routine tasks or generate no material return if organisational capability is absent. This complementarity helps explain why a nominally equal subscription does not create equal economic access: large companies can spread integration, governance and evaluation costs across more users and transactions. Eurostat’s 2025 enterprise survey supplies an early observable indicator of this scale effect. AI technologies were used by 55.03% of large EU enterprises, compared with 30.36% of medium-sized enterprises and 17.00% of small enterprises. The survey does not establish that compute prices caused the difference, but it supports testing affordability, expertise and scale as contributing mechanisms. Use of Artificial Intelligence in Enterprises – Eurostat – December 2025
As a capital good, compute is characterised by high initial expenditure, rapid technological obsolescence, uncertain residual value and strong utilisation effects. A cluster that operates near capacity can distribute depreciation and facility costs across far more billable output than an underutilised local installation. The relevant average cost is ACt = FCt ÷ Qt + VCt, where FC includes depreciation, financing, buildings, networking and fixed labour; Q is quality-adjusted output; and VC includes marginal energy, cooling, maintenance and workload-specific costs. Scale reduces FC ÷ Q until congestion, grid limitations, networking complexity or management costs offset the gain. Capital intensity also changes bargaining power. Buyers able to commit to long-term volumes can secure capacity and preferential commercial terms, while students, researchers, small firms and lower-income governments generally purchase at retail or shared-service prices. Alphabet’s capital expenditure increased from $52.5 billion in 2024 to $91.4 billion in 2025, with the company stating that the expenditure primarily reflected technical infrastructure and that it expected a significant further increase in 2026. This does not identify the precise AI share or prove future price increases, but it establishes the scale of capital mobilisation available to a leading platform operator. Alphabet Annual Report on Form 10-K for the Fiscal Year Ended December 31, 2025 – Alphabet and U.S. Securities and Exchange Commission – February 2026
1.4 Compute as a Utility-Like Service
Cloud AI resembles a utility because it transforms a complex capital system into continuous, remotely delivered and metered access. Users do not need to own a fabrication plant, an accelerator cluster or a data centre; they pay for a service expressed through subscriptions, input tokens, output tokens, images, audio duration, storage, provisioned throughput or completed operations. Utility-like delivery can democratise access because it converts capital expenditure into variable expenditure and permits small buyers to consume fractions of expensive infrastructure. At the same time, it can create dependency if essential functions become inseparable from the provider’s identity, model interface, data formats, security architecture or proprietary tools. The utility analogy is therefore strongest in billing, continuity and dependency, but weaker in technical fungibility. Electricity of a specified voltage is substantially standardised; AI outputs are not. A cheaper model may be an inadequate substitute if it has lower reliability, weaker reasoning, insufficient language coverage, shorter context or no access to required tools. The effective unit is consequently not one token but one successfully completed, quality-controlled task. A useful measurement is ECPm,w = Cm,w ÷ Um,w, where ECP is the effective completion price of model m on workload w, C is total expenditure including retries and supervision, and U is the number of outputs that meet a predefined quality threshold.
The utility analogy also raises questions of tariff design, capacity reservation, universal access, switching rights and service continuity. If providers face congestion, they can respond through waiting times, usage caps, degraded model access, higher priority-tier prices or geographic rationing. Dynamic pricing is not automatically abusive: it can allocate scarce capacity and encourage off-peak demand. It becomes socially consequential when access to education, research, healthcare or competitive business functions depends on the ability to pay peak or frontier-model premiums. Cloud dependence is already broad enough to make this question material. Eurostat reported that 52.74% of EU enterprises purchased cloud-computing services in 2025, while 40.89% of all surveyed enterprises purchased at least one service classified as sophisticated. Among enterprises using paid cloud services, 77.53% were classified as highly dependent on sophisticated cloud functions. These figures concern cloud services generally rather than frontier AI, and they must not be misrepresented as an AI-dependency rate. They nevertheless establish the institutional substrate through which metered AI can diffuse rapidly. Cloud Computing: Statistics on the Use by Enterprises – Eurostat – January 2026
| Utility characteristic | Electricity or telecommunications analogue | AI-compute equivalent | Critical difference |
|---|---|---|---|
| Metering | kWh, minute, GB | Token, image, accelerator-second, completed task | Units are not quality-equivalent |
| Capacity reservation | Contracted power or bandwidth | Provisioned throughput and reserved accelerators | Hardware and model availability vary |
| Congestion management | Peak tariffs or throttling | Rate limits, queues, latency and model tiers | Quality may be rationed as well as quantity |
| Reliability obligation | Uptime and continuity | Service-level agreement and model availability | Model behaviour can change after updates |
| Switching | Change supplier using standard connection | Port data, prompts, tools and workflows | Proprietary APIs and model behaviour impede portability |
| Universal access | Regulated baseline service | Possible educational or research compute entitlement | No accepted minimum AI-capability standard exists |
| Cross-subsidy | One user class subsidises another | Advertising, enterprise contracts or investor capital subsidise consumer access | Cross-subsidy is often opaque |
1.5 Strategic Infrastructure and Compute Sovereignty
Compute becomes strategic infrastructure when its unavailability can impair economic continuity, state capacity, security or essential services. Sovereignty in this context does not require that every component be domestically manufactured. No major economy is fully autonomous across lithography, design software, advanced logic, memory, substrates, networking, servers and energy. Operational compute sovereignty is better defined as the ability of a jurisdiction or institution to obtain, govern, relocate, audit and sustain the minimum quality-adjusted capacity necessary for its critical functions under plausible disruption. This definition contains five dimensions: physical availability, jurisdictional control, operational competence, software and data portability, and supply-chain resilience. A country that owns a data centre but depends entirely on foreign accelerators, proprietary frameworks, remote updates and imported specialist labour possesses infrastructure without complete operational sovereignty. Conversely, a country using foreign-origin hardware may retain meaningful control if it holds installed capacity, diversified spares, local expertise, transferable models and legally enforceable continuity arrangements.
A sovereignty index used later in the report will take the form CSIj = Σwkzj,k, where CSI is the Compute Sovereignty Index for jurisdiction j, z represents normalised indicators and the weights will be tested under equal-weight, principal-component and expert-weight specifications. Indicators will cover installed accelerator capacity, domestic data-centre power, redundancy, cloud-region availability, network connectivity, access to replacement components, public research compute, workforce capability, legal control, model portability and exposure to foreign export restrictions. Europe’s policy response demonstrates that governments no longer treat frontier compute as an ordinary import. The European Commission describes AI Gigafactories as facilities bringing together more than 100,000 advanced AI processors, reliable power, advanced networking and secure supply chains. In July 2026, it launched a call intended to support up to seven such facilities and unlock more than €30 billion in investment. These are policy targets rather than completed capacity and must remain classified as announced mobilisation. AI Factories – European Commission – July 2026 EU Launches AI Gigafactories Call to Boost Europe’s Computing Capacity and Unlock More Than €30 Billion – European Commission – July 2026
| Sovereignty dimension | Operational question | Proposed indicator | Failure condition |
|---|---|---|---|
| Physical capacity | Is sufficient compute installed or contractually guaranteed? | Quality-adjusted accelerator capacity per million inhabitants and per unit of GDP | Critical workloads cannot obtain capacity |
| Jurisdictional control | Can foreign law or supplier action interrupt access? | Share of critical capacity under domestic or allied legal control | Provider or exporting state can terminate service |
| Energy resilience | Can the grid support sustained workloads? | Firm MW, redundancy and expected outage-adjusted availability | Capacity exists but cannot be powered reliably |
| Supply-chain resilience | Can failed systems be repaired or expanded? | Supplier diversity, replacement lead time and spare inventory | Single upstream disruption immobilises capacity |
| Software portability | Can workloads move between providers or architectures? | Migration time, compatible frameworks and open formats | Switching destroys functionality or requires prohibitive redevelopment |
| Data governance | Can sensitive data remain controlled and auditable? | Local processing, encryption, audit and contractual enforceability | Compute access requires unacceptable data exposure |
| Human capability | Can local institutions operate and optimise the system? | Engineers, administrators and AI researchers per installed unit | Hardware remains underused or foreign-operated |
| Research access | Can universities and independent researchers obtain meaningful capacity? | Public compute hours per researcher | Frontier research becomes institutionally exclusive |
1.6 Compute as a Geopolitical Resource
Compute is geopolitical because advanced AI capability is territorially concentrated, supply chains cross multiple jurisdictions and access can be altered by export licensing, sanctions, investment controls and alliance relationships. The relevant unit of geopolitical power is not the semiconductor alone but the controlled combination of design, fabrication, packaging, memory, equipment, software and cloud deployment. A state may possess mineral inputs yet lack fabrication technology; it may host fabrication while depending on foreign tools; it may operate data centres while lacking advanced accelerators; or it may develop models that cannot be deployed economically without foreign cloud infrastructure. This interdependence produces bargaining power at chokepoints and vulnerability where substitution is slow. It also creates a distinction between global aggregate capacity and politically accessible capacity. Additional accelerators produced worldwide do not reduce scarcity for every buyer if export controls, allocation agreements or security rules exclude particular jurisdictions.
The semiconductor chain demonstrates this concentration of technical functions. TSMC reported that advanced technologies defined as 7 nanometres and below accounted for 77% of wafer revenue in the second quarter of 2026, while its high-performance-computing platform accounted for 66% of quarterly revenue. These are revenue shares within TSMC, not global market shares, but they show the company’s increasing exposure to leading-edge and HPC demand. TSMC Second Quarter 2026 Earnings Conference Transcript – Taiwan Semiconductor Manufacturing Company – July 2026 NVIDIA’s fiscal-year 2026 filing states that the company depends on foundries to manufacture its semiconductor wafers and presents itself as a full-stack computing-infrastructure company rather than a component supplier. This combination of external manufacturing dependence and downstream platform breadth illustrates why power in the AI chain is distributed but asymmetric: a design firm may control the architecture and software ecosystem without owning the fabrication plant, while the foundry controls a scarce production process without independently controlling final demand. NVIDIA Annual Report on Form 10-K for the Fiscal Year Ended January 25, 2026 – NVIDIA and U.S. Securities and Exchange Commission – February 2026
1.7 Potential Financialisation: What the Bitcoin Analogy Explains—and What It Distorts
AI compute may become more financialised without becoming a cryptocurrency or a conventionally traded commodity. Financialisation occurs when rights to future capacity, priority, revenue or infrastructure are separated from immediate physical use and become objects of contracting, financing or speculation. Long-term accelerator orders, take-or-pay cloud contracts, reserved data-centre capacity, power-purchase agreements, equipment-backed debt and capacity options can transfer scarcity expectations into present asset values. A model provider may secure future compute through a strategic investment from a hyperscaler; the hyperscaler may receive preferred access, distribution rights or revenue participation; and smaller buyers may then face a residual spot market. The Federal Trade Commission’s January 2025 study of major AI partnerships identified agreements involving substantial cloud commitments, exclusivity-related provisions, access to sensitive information and potential switching implications. The report does not establish that every partnership is anticompetitive, but it confirms that investment, cloud purchasing and competitive alignment can be contractually interconnected. FTC Staff Report on AI Partnerships and Investments 6(b) Study – Federal Trade Commission – January 2025
The Bitcoin analogy is valid only at the level of scarcity narratives, divisible consumption units and expectation-driven pricing. It fails in five fundamental respects. First, compute supply is endogenous: companies can manufacture more accelerators and build more data centres, although with delays. Second, compute is heterogeneous: an hour on one architecture is not necessarily substitutable for an hour on another. Third, compute depreciates as newer architectures improve price-performance. Fourth, compute is geographically and electrically constrained. Fifth, it is productive capacity rather than a bearer asset; value arises from the workloads it executes. A more accurate analogy is a hybrid of electricity, cloud bandwidth, industrial machinery and reserved transport capacity. The report will test financialisation through the ratio FCIt = Vreserved,t ÷ Vtotal,t, where Vreserved is the estimated value of capacity committed under forward or preferential arrangements and Vtotal is total available capacity. A high ratio would indicate that quoted retail prices describe only a residual market. Because contractual disclosure is incomplete, the measure will require intervals and must never be presented with false precision.
| Attribute | Bitcoin | Electricity | Industrial machinery | AI compute |
|---|---|---|---|---|
| Supply rule | Protocol-limited | Capacity and fuel constrained | Investment-driven | Fabrication, packaging, memory, power and investment constrained |
| Homogeneity | Relatively fungible units | Standardised within technical limits | Highly heterogeneous | Highly heterogeneous |
| Storage | Digital bearer asset | Limited and costly | Asset can remain idle | Capacity flow is perishable; hardware is durable but depreciating |
| Depreciation | No physical depreciation | Not applicable to consumed energy | Physical and technological depreciation | Rapid technological and economic depreciation |
| Geographic dependence | Network access | Grid-specific | Installation-specific | Data-centre, grid, jurisdiction and network-specific |
| Productive use | Indirect | Universal energy input | Direct production asset | Direct cognitive and computational input |
| Dynamic pricing potential | Exchange price | Spot and peak tariffs | Rental or leasing rates | Token, latency, model, workload and capacity-tier pricing |
| Appropriate analogy score | Low | Moderate | High | Hybrid category |
1.8 The AI Compute Value Chain
The AI value chain is best analysed as a set of complementary layers rather than a simple linear sequence. A shortage at any indispensable layer can reduce the output of the entire system, while control of a technically replaceable layer may generate little bargaining power if switching is inexpensive. The relevant measure is therefore bottleneck-adjusted capacity. If effective capacity is constrained by the minimum available complementary input, then Ceffective = min(Clogic, Cmemory, Cpackaging, Cnetwork, Cpower, Csoftware). This Leontief-style representation is deliberately stringent: in practice, some substitution is possible through lower precision, slower networks, different models or reduced service quality. The model will later be expanded into a constant-elasticity-of-substitution structure. For Chapter 1, it captures the central point that surplus accelerator-design demand is economically irrelevant if HBM or packaging prevents finished systems from being delivered.
Table 1.1 — AI Compute Value Chain, Scarcity Nodes and Bargaining Power
| Layer | Core economic function | Capital intensity | Short-run substitutability | Potential scarcity indicator | Source of bargaining power | Primary later chapter |
|---|---|---|---|---|---|---|
| Semiconductor equipment | Enables leading-edge fabrication | Extreme | Very low | Tool backlog and delivery lead time | Proprietary process technology and installed base | Chapter 3 |
| Design software and IP | Converts architecture into manufacturable design | High knowledge intensity | Low | Licence concentration and switching time | Standards, libraries and engineering integration | Chapter 3 |
| Accelerator architecture | Executes AI workloads | High R&D | Medium over longer periods | Quality-adjusted shipments and order backlog | Performance, software ecosystem and developer base | Chapters 2–3 |
| Leading-edge foundry | Manufactures logic dies | Extreme | Very low in the short run | Advanced-node utilisation and capacity | Yield, process leadership and scale | Chapter 3 |
| HBM and advanced memory | Supplies model weights and data at high bandwidth | Extreme | Low | Contracted supply, bit output and price | Manufacturing know-how and qualification | Chapters 2–3 |
| Advanced packaging | Integrates logic and memory | Extreme | Very low | CoWoS-equivalent capacity and lead time | Process qualification and limited capacity | Chapter 3 |
| Substrates and components | Connect and support packaged systems | High | Low to medium | Lead times and supplier concentration | Qualification and production constraints | Chapter 3 |
| Server integration | Converts components into deployable systems | High | Medium | Rack delivery and liquid-cooling availability | Integration, warranty and customer qualification | Chapters 2–4 |
| Networking | Connects accelerators at cluster scale | High | Low for frontier training | Bandwidth, topology and port availability | Scale efficiency and architecture compatibility | Chapters 3–5 |
| Data-centre facilities | Houses and cools systems | Extreme | Low locally | MW under construction and connection queue | Land, permits, cooling and grid access | Chapter 5 |
| Electricity | Operates the infrastructure | Extreme system capital | Low during local grid constraint | Firm MW, prices and connection delay | Location-specific availability | Chapters 5 and 9 |
| Cloud orchestration | Allocates and meters resources | Extreme | Medium, but migration is costly | Utilisation, reserved capacity and availability | Scale, customer base and integrated services | Chapters 3 and 5 |
| Foundation models | Converts infrastructure into capability | Extreme training and R&D | Task-dependent | Benchmark-adjusted access and price | Model quality, brand, data and distribution | Chapters 5 and 7 |
| Application layer | Embeds models in workflows | Variable | Medium to high | Customer retention and integration depth | Data, workflow and distribution lock-in | Chapters 6–7 |
| User organisation | Converts output into economic value | Variable | Not substitutable | Skills, governance and complementary capital | Proprietary data and domain competence | Chapters 6–8 |
Three conclusions follow from this map. First, market concentration must be measured separately at every layer; a single “AI market share” is analytically meaningless. Second, high concentration does not necessarily identify the binding bottleneck. A highly concentrated layer with ample capacity and credible substitution may exercise less power than a somewhat less concentrated layer with long qualification periods and no inventory. Third, bargaining power can move over time. When accelerators are scarce, the designer or cloud allocator may dominate; when accelerators become abundant but electricity connections are delayed, power availability and data-centre permits become the constraint. The U.S. Department of Energy reported that American data centres consumed approximately 176 TWh in 2023, or 4.4% of total U.S. electricity, and projected a range of 325–580 TWh by 2028, equivalent to approximately 6.7–12% of national consumption. These are scenarios rather than fixed outcomes, but they demonstrate why power must be treated as part of the compute chain rather than an external operating expense. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – U.S. Department of Energy – December 2024
1.9 Where Value, Scarcity and Bargaining Power Accumulate
Economic value is captured where an actor controls a scarce complement, creates a difficult-to-replicate performance advantage, owns distribution or raises the cost of switching. These conditions do not always coincide. A foundry may possess strong technical scarcity but remain exposed to a few large customers; a hyperscaler may have abundant infrastructure but possess bargaining power through customer relationships and bundled services; a model provider may have superior capability but depend on a cloud partner for capital and distribution. The balance can be represented through a bargaining-power score BPi = w₁Si + w₂Li + w₃Di + w₄Ei + w₅Vi − w₆Bi, where S is scarcity, L lock-in, D differentiation, E ecosystem control, V vertical reach and B buyer concentration. This is a research index, not an established law. Each component will be normalised and tested with alternative weights.
Table 1.2 — Preliminary Bargaining-Power Assessment
| Node | Physical scarcity | Ecosystem lock-in | Buyer concentration | Replication time | Preliminary power assessment | Evidentiary status |
|---|---|---|---|---|---|---|
| Leading-edge fabrication | High | Medium | High | Very long | High but constrained by major buyers and capital cycles | Framework inference |
| Advanced packaging | High | Medium | High | Long | High during AI-capacity expansion | Supported bottleneck hypothesis |
| HBM | High | Medium | High | Long | High if contracted output leaves a small residual market | To be quantified |
| Accelerator design | High | Very high for established software ecosystems | High | Long | Very high where performance and software reinforce one another | To be quantified |
| Data-centre power connection | Locally high | Low | Variable | Long | Potentially decisive in constrained regions | Supported by DOE scenarios |
| Hyperscale cloud | High capacity barrier | Very high | Low buyer power for retail users | Very long | High through bundling, scale and allocation | To be quantified |
| Frontier model | High R&D barrier | High | Mixed | Uncertain | High if task quality is not reproducible by alternatives | Benchmark-dependent |
| Open-weight model | Lower legal access barrier | Lower | Low | Shorter | Limited direct pricing power, but infrastructure remains necessary | Workload-dependent |
| Enterprise application | Usually lower physical barrier | Potentially high workflow lock-in | Fragmented | Medium | Can capture value through integration and proprietary data | Sector-dependent |
| End user | Low individually | Often locked in | Very low individually | Not applicable | Price taker unless organised purchasing exists | Strong prior expectation |
NVIDIA’s fiscal 2026 results illustrate value capture without resolving its cause. Fiscal-year revenue was $215.9 billion, 65% above the preceding year, and fiscal-year GAAP gross margin was 71.1%. In the first quarter of fiscal 2027, Data Center revenue reached $75.2 billion, 92% above the corresponding prior-year quarter. These data demonstrate exceptional commercial performance, not unlawful conduct. Later chapters will compare margins with historical baselines, invested capital, R&D, supply commitments, competitive alternatives and changes in product mix before estimating economic rent. NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026 – NVIDIA – February 2026 NVIDIA Announces Financial Results for First Quarter Fiscal 2027 – NVIDIA – May 2026
1.10 Core Concepts and Operational Definitions
Table 1.3 — Conceptual Dictionary
| Concept | Operational definition for this report | Proposed measurement | What must not be inferred automatically |
|---|---|---|---|
| Compute scarcity | Quality-adjusted demand exceeds capacity available at the prevailing price and service conditions | Utilisation, queues, lead times, stock-outs, premiums and unmet demand | A high price alone does not prove scarcity |
| Physical scarcity | Insufficient deployable hardware, power or facilities | Units, MW, delivery times and availability | It does not necessarily imply market abuse |
| Economic scarcity | Capacity exists but is unaffordable to a defined user group | Cost-to-income or cost-to-revenue ratio | It is not equivalent to physical shortage |
| Artificial scarcity | Supply is strategically withheld below the profit-maximising competitive benchmark | Capacity, output, inventories, contracts and counterfactual supply | It cannot be inferred from concentration alone |
| Compute sovereignty | Ability to sustain governed access to essential compute under disruption | Composite resilience and control index | Domestic ownership alone is insufficient |
| Access inequality | Unequal distribution of quality-adjusted AI capability | Gini, Theil, percentile ratios and affordability indices | Subscription ownership is not capability equality |
| Digital exclusion | Inability to obtain meaningful AI use because of connectivity, devices, skills, language or affordability | Multidimensional deprivation rate | Being online does not imply meaningful AI access |
| Economic rent | Return above the minimum required to keep capital and capability in their current use | Excess ROIC, residual income and margin decomposition | Accounting profit is not identical to rent |
| Network effects | Product value increases as users, developers, tools or complementary data increase | Developer base, integrations and retention | Scale alone is not proof of a network effect |
| Economies of scale | Average cost falls as output or utilisation rises | Cost elasticity relative to output | Large size is not proof of continuing scale economies |
| Economies of scope | Joint production costs less than separate production | Multi-product cost function | Bundling is not always efficient or harmful |
| Vertical integration | One firm controls multiple adjacent layers | Revenue, ownership and contractual control map | Integration can create efficiency or foreclosure |
| Switching cost | Economic loss caused by moving provider or architecture | Migration expenditure, time, performance loss and retraining | Contract termination fees capture only one component |
| Data lock-in | User data, embeddings, histories or workflows cannot move without loss | Portability completeness and migration cost | Data possession alone does not create lock-in |
| Ecosystem lock-in | Complementary software and skills make alternatives costly | Compatibility, developer tools and retraining | Popularity alone is insufficient |
| Technological dependence | Critical capability requires a supplier or jurisdiction without timely substitute | Replacement time and criticality-adjusted exposure | Import dependence is not always vulnerability |
Compute scarcity will be decomposed into five observable forms: quantity scarcity, quality scarcity, temporal scarcity, geographic scarcity and institutional scarcity. Quantity scarcity exists when total units are inadequate. Quality scarcity exists when lower-performance substitutes are available but cannot complete the required workload. Temporal scarcity appears as waiting time, delivery delay or queueing. Geographic scarcity arises when capacity is available globally but not within the required jurisdiction or latency boundary. Institutional scarcity exists when capacity is accessible to hyperscalers or elite universities but not to small companies, independent researchers or public institutions. This decomposition prevents a common analytical error: concluding that scarcity has disappeared because some form of compute remains purchasable. A user able to run a small, highly quantised local model may possess quantity access but remain quality-excluded from workloads requiring long context, multimodality, high reliability or rapid concurrent inference.
1.11 Network Effects, Scale, Scope and Vertical Integration
AI markets can exhibit several mutually reinforcing scale mechanisms. Infrastructure scale lowers average fixed cost by raising utilisation and spreading data-centre, networking and engineering expenditure across more workloads. Model-development scale allows training, safety, evaluation and post-training costs to be amortised across more users. Data scale can improve products when legally and technically usable interactions generate feedback. Distribution scale lowers customer-acquisition costs. Developer-network effects increase the value of an architecture as more libraries, optimisation tools and trained personnel become available. These mechanisms can generate genuine efficiency while simultaneously raising entry barriers. The analytical task is not to label scale as harmful, but to determine whether efficiency gains are passed to users through lower quality-adjusted prices or retained as durable rent.
Economies of scale will be tested through ln(ACi,t) = α + βln(Qi,t) + γXi,t + εi,t. A statistically negative β would be consistent with falling average cost as output expands, subject to data quality and identification limitations. Economies of scope will be examined through SC = C(Y₁,0) + C(0,Y₂) − C(Y₁,Y₂). A positive value indicates a cost advantage from joint production. Vertical integration will be evaluated through both efficiency and foreclosure channels. Integration between accelerators, networking, cloud, models and applications can reduce coordination costs and optimise the stack. It can also make competing models or hardware less attractive through preferential access, bundling, interoperability restrictions or discriminatory commercial terms. The FTC’s AI-partnership report provides a relevant empirical starting point because it examined how investments and partnerships could affect access to cloud resources, switching, exclusivity and sensitive information. It does not support a blanket conclusion that integration is anticompetitive; it establishes the contractual mechanisms that must be tested.. FTC Staff Report on AI Partnerships and Investments 6(b) Study – Federal Trade Commission – January 2025
Table 1.4 — Efficiency and Exclusion Tests
| Mechanism | Efficiency hypothesis | Exclusion hypothesis | Required evidence |
|---|---|---|---|
| Hardware–software integration | Improves performance and lowers workload cost | Locks developers into one architecture | Cross-platform benchmark, migration cost and price comparison |
| Cloud–model integration | Reduces deployment cost and latency | Preferentially allocates capacity to affiliated model | Capacity terms, pricing and availability by model |
| Model–application integration | Improves user experience and workflow automation | Makes rival models technically or commercially inaccessible | Interoperability, default settings and contractual terms |
| Data–model integration | Improves relevance and personalisation | Entrenches incumbent through inaccessible feedback data | Data portability, learning effects and incremental performance |
| Bundled enterprise services | Reduces procurement and security cost | Uses strength in one market to foreclose another | Stand-alone versus bundled prices and customer switching |
| Long-term capacity contract | Finances new infrastructure and stabilises demand | Removes capacity from competitors or residual buyers | Contract duration, reserved share and alternative capacity |
| Developer ecosystem | Reduces implementation cost | Raises retraining and redevelopment barriers | Labour-market skills, library compatibility and porting time |
1.12 Market Concentration, Dominance and Economic Rent
Concentration is an indicator, not a verdict. The report will calculate concentration at narrowly defined functional layers using CR₄ = s₁ + s₂ + s₃ + s₄ and HHI = Σsi², with shares expressed as percentages. Market definition will precede calculation: data-centre GPUs, consumer GPUs, HBM, leading-edge foundry services, general cloud infrastructure and frontier-model subscriptions cannot be placed in a single denominator. The U.S. 2023 Merger Guidelines state that a merger producing or increasing a highly concentrated market may raise a presumption of illegality under specified conditions, but the guidelines are an enforcement framework, not a declaration that every concentrated market is unlawful. 2023 Merger Guidelines – U.S. Department of Justice and Federal Trade Commission – December 2023 The report will therefore supplement HHI with entry conditions, capacity constraints, price-cost margins, buyer power, innovation, multi-homing, switching and credible substitution.
Economic rent will be estimated through multiple methods because no single accounting ratio isolates market power. The first measure will be excess return: ERi,t = ROICi,t − WACCi,t. The second will be residual income: RIi,t = NOPATi,t − WACCi,t × ICi,t. The third will decompose gross-margin changes into selling price, product mix, unit cost, utilisation and accounting effects. A positive excess return may represent innovation rent, temporary scarcity rent, intangible capital omitted from the balance sheet, superior management or durable market power. The classification will depend on persistence, entry, price response to capacity expansion and whether returns remain above benchmarks after risk and intangible investment are accounted for.
Table 1.5 — Rent Classification Framework
| Rent category | Economic origin | Expected duration | Competition implication | Empirical signature |
|---|---|---|---|---|
| Innovation rent | Superior technology or product | Temporary unless continuously renewed | Often pro-competitive | High R&D, rapid performance improvement and eventual entry |
| Scarcity rent | Demand temporarily exceeds capacity | Cyclical or capacity-dependent | Not inherently anticompetitive | Prices and margins fall after capacity expansion |
| Quasi-rent | Recovery on specialised sunk investment | Asset-life dependent | Usually normal | Return falls as capital depreciates or contracts expire |
| Risk premium | Compensation for uncertain investment | Linked to risk | Normal if proportionate | Return correlates with investment risk |
| Ecosystem rent | Complementary tools and users reinforce incumbent | Potentially persistent | Ambiguous | High switching costs and developer dependence |
| Monopoly or dominance rent | Durable market power restricts competitive pressure | Persistent | Potential concern | High margins, weak entry, limited substitution and strategic exclusion |
| Regulatory rent | Rules or controls restrict entry | Policy-dependent | May be intended or distortive | Profit tied to licences, quotas or protected access |
| Geopolitical rent | Jurisdictional access or controls create privileged supply | Event-dependent | Security and trade issue | Regional price and availability divergence |
1.13 Access Inequality and Digital Exclusion
AI access inequality must be measured as a distribution of capabilities rather than a distribution of accounts. A free model, a paid frontier subscription, a dedicated enterprise instance and a locally controlled high-memory system are not equivalent resources. Quality-adjusted access will therefore incorporate model performance, usage limits, context capacity, latency, privacy, availability, language coverage and tool access. The proposed individual measure is QAAi,t = ΣwkAi,k,t, where A represents normalised access dimensions. Affordability will be measured as AAIi,t = Cadequate,i,t ÷ Ydisposable,i,t. For companies, the denominator becomes revenue, operating expenditure or labour cost. For universities, it becomes research expenditure or funding per researcher. A system will be deemed affordable only relative to a defined workload; a low-cost model that cannot complete the task is not an affordable substitute.
The baseline digital divide remains severe. The ITU estimated that 6.0 billion people were online in 2025, while 2.2 billion remained offline, mainly in low- and middle-income economies. It reported internet use of 94% in high-income countries and 23% in low-income countries. Measuring Digital Development: Facts and Figures 2025 – International Telecommunication Union – November 2025 These figures do not measure AI access, but they define its lower boundary: users without meaningful connectivity cannot participate fully in cloud AI, while users with weak connectivity may be unable to sustain multimodal or interactive workloads. The ITU’s 2025 connectivity analysis frames meaningful connectivity through quality, availability, affordability, devices, skills and security, a multidimensional structure that this report will extend to AI compute. Global Connectivity Report 2025 – International Telecommunication Union – 2025
Table 1.6 — AI Access Ladder
| Level | Accessible capability | Typical constraint | Economic consequence |
|---|---|---|---|
| 0 — Excluded | No reliable internet or compatible device | Connectivity, electricity, income | No regular AI participation |
| 1 — Nominal access | Free, severely limited or intermittent service | Caps, queues and older models | Basic assistance but weak continuity |
| 2 — Standard consumer access | Paid general-purpose model | Subscription affordability and privacy | Individual productivity improvement |
| 3 — Advanced professional access | Higher limits, tools, larger context and multimodality | Higher recurring cost | Material professional advantage |
| 4 — Enterprise controlled access | Security, integration, governance and support | Organisational scale and implementation cost | Workflow-level productivity and data integration |
| 5 — Local high-performance access | Private inference with substantial memory and throughput | Hardware, energy and technical expertise | Sovereignty and predictable availability |
| 6 — Frontier research access | Large clusters, training and large-scale experimentation | Extreme capital and institutional concentration | Ability to create rather than merely consume frontier systems |
1.14 Falsifiable Research Hypotheses
H1 — AI-Relevant Hardware Prices Have Risen Faster Than General Consumer-Electronics Prices
H1 will be supported only if a representative, transaction-weighted AI-hardware price index rises faster than the appropriate consumer-electronics comparator over the defined period. The test cannot rely on selected flagship products or isolated resale listings. Chapter 2 will construct matched and hedonic indices for GPUs, memory, workstations, laptops and relevant Apple and Samsung configurations. The principal regression will be ln(Pi,t) = α + βXi,t + γt + δb + εi,t, where X captures memory, bandwidth, compute performance, energy, condition and other measurable characteristics; γ is the time effect and δ the brand or product-family effect. H1 will be rejected if nominal prices rise but quality-adjusted prices decline at a rate consistent with or faster than the comparator index.
H2 — The Inflation-Adjusted Cost of a Fixed Level of AI Capability Has Increased
H2 is stricter than H1. The unit of analysis will be a reproducible workload, not hardware. Cost per successful task will include capital cost, electricity, software, time, retries and human correction. The fixed-capability basket will cover reasoning, coding, long-document analysis, multilingual work and multimodal processing. H2 will be supported when the real minimum cost of reaching a constant benchmark threshold rises significantly across multiple workloads and deployment modes. It will be rejected if improved algorithms, quantisation or hardware efficiency reduce real workload cost despite higher purchase prices.
H3 — Used-Market Prices Indicate Persistent Scarcity Rather Than Temporary Speculation
H3 will compare used prices with launch MSRP, contemporaneous new prices, depreciation benchmarks and product availability. The scarcity premium will be USPi,t = (Pused,i,t − Pfundamental,i,t) ÷ Pfundamental,i,t. Persistence will require premiums across multiple months, regions and product classes, accompanied by low availability or delivery delays. H3 will be rejected if premiums are confined to brief launches, collectible products, unreliable listings, tax distortions or speculative episodes without transaction evidence.
H4 — Concentration Enables Suppliers to Retain Above-Normal Economic Rents
H4 will require a relationship between concentration and persistent excess returns after controlling for R&D, risk, intangible capital, scale and product mix. The panel specification will be ERi,t = α + βHHIm,t + γZi,t + μi + τt + εi,t. A positive β alone will not prove causality. Event studies around capacity expansion, entry or supply shocks will test whether margins respond competitively. H4 will be rejected if high returns dissipate with entry or are explained by innovation and risk-adjusted investment.
H5 — Frontier AI Access Produces Measurable Economic and Educational Advantages
H5 will be tested through matched or randomised comparisons where credible data exist. Outcomes will include task accuracy, completion time, grades, research output, coding productivity, revenue and error rates. Treatment must represent quality-adjusted access rather than any AI use. The preferred estimand is ATE = E[Y(1) − Y(0)], with controls for prior ability, institution, income and task. H5 will be rejected for workloads where smaller or cheaper systems achieve statistically equivalent results.
H6 — Lower-Income Users and Countries Are Increasingly Restricted to Lower-Capability Systems
H6 will compare quality-adjusted access, not internet penetration alone, across income groups. The tests will use affordability, available model tiers, cloud regions, hardware-wage ratios, electricity reliability, language support and public research compute. Increasing restriction requires a widening gap through time. H6 will be rejected if low-cost access, open models and shared infrastructure cause capability convergence after purchasing-power adjustment.
H7 — Current AI-Service Pricing Does Not Fully Reflect Long-Run Infrastructure Cost and May Rise Materially
H7 will estimate provider break-even prices using depreciation, financing, energy, facilities, networking, labour, training and utilisation. The cost-recovery ratio will be CRRp,t = RevenueAI,p,t ÷ FullyAllocatedCostAI,p,t. Because public segment data are incomplete, estimates will be ranges. H7 will be supported if recurring prices remain below long-run cost under plausible utilisation and cross-subsidy assumptions. It will be rejected if current pricing covers cost and required capital return or if unit costs decline faster than tariff pressure.
H8 — AI Access Is Evolving into a Dynamically Priced, Metered Resource
H8 concerns market design rather than simply rising prices. It will be supported by documented movement toward multiple metering dimensions: model class, reasoning intensity, latency, context length, time of demand, geography, capacity reservation and service priority. A Dynamic Pricing Complexity Index will count and weight independently priced dimensions. H8 will be rejected if markets converge toward simple, stable, interoperable and progressively cheaper tariffs.
1.15 Hypothesis Test Matrix
Table 1.7 — Falsification Criteria
| Hypothesis | Primary dependent variable | Core comparator | Support threshold | Rejection condition | Principal risk |
|---|---|---|---|---|---|
| H1 | AI-hardware price index | Consumer-electronics index | Statistically higher cumulative growth | Quality-adjusted growth is equal or lower | Product-selection bias |
| H2 | Real cost per fixed-capability task | Constant benchmark threshold | Significant real increase across robust specifications | Real task cost declines | Benchmark drift |
| H3 | Used-market premium and duration | Depreciation-adjusted fair value | Persistent multi-market premium with confirmed transactions | Short-lived or listing-only premium | Selection and fake listings |
| H4 | Risk-adjusted excess return | Competitive sector benchmark | Persistent excess return associated with concentration and weak entry | Returns explained by innovation, risk or cycle | Market-definition error |
| H5 | Productivity or educational outcome | Comparable non-frontier access | Positive statistically and economically significant treatment effect | Equivalent outcomes from lower-cost systems | User self-selection |
| H6 | Quality-adjusted access gap | Income and regional groups | Gap widens through 2026–2031 | Convergence after PPP and quality adjustment | Missing lower-income data |
| H7 | Cost-recovery ratio | Fully allocated long-run cost | Ratio below 1 under central assumptions | Current prices cover cost and capital return | Opaque segment accounts |
| H8 | Pricing-complexity index | Historical tariff structure | More independently priced dimensions and capacity tiers | Simplification and commoditisation | Contract confidentiality |
1.16 Research Map: Hypotheses, Data and Later Chapters
Table 1.8 — Full Research Architecture
| Hypothesis | Dataset required | Minimum frequency | Method | Robustness test | Developed in |
|---|---|---|---|---|---|
| H1 | Launch MSRP, official current prices, verified transaction prices, specifications and electronics indices | Monthly or quarterly, 2020–2031 | Hedonic price regression and matched-model index | Alternative quality measures and geography fixed effects | Chapter 2 |
| H2 | Hardware TCO, API prices, benchmarks, electricity and human-review cost | Quarterly or annual | Cost-per-successful-task frontier | Alternative benchmarks, quantisation and utilisation | Chapters 2, 4 and 5 |
| H3 | Verified second-hand transactions, age, condition, warranty and new availability | Weekly or monthly | Premium-duration and event-study models | Listings versus completed transactions | Chapter 2 |
| H4 | Segment shares, CR₄, HHI, margins, ROIC, WACC, entry and capacity | Quarterly or annual | Panel regression, residual income and event study | Intangible-capital adjustment and market definitions | Chapter 3 |
| H5 | User-level access, task outcomes, time, errors and background variables | Experimental or panel | RCT, matching or difference-in-differences | Placebo tasks and heterogeneous effects | Chapter 7 |
| H6 | Income, hardware affordability, connectivity, cloud access, language and public compute | Annual | Composite index, Gini, Theil and decomposition | PPP versus market FX and alternative weights | Chapter 8 |
| H7 | Provider capex, depreciation, energy, utilisation, revenue and tariffs | Quarterly | Fully allocated cost and break-even model | Monte Carlo utilisation and asset-life assumptions | Chapter 5 |
| H8 | Public API tariffs, enterprise contracts, priority products and reserved capacity | Monthly or event-based | Tariff taxonomy and complexity index | Provider and regional comparison | Chapters 5 and 9 |
Table 1.9 — Cross-Chapter Dependency Map
| Later chapter | Contribution to Chapter 1 hypotheses | Principal output |
|---|---|---|
| Chapter 2 — Hardware Inflation | Tests H1 and H3; supplies hardware inputs for H2 | Nominal, real and quality-adjusted price indices |
| Chapter 3 — Industrial Concentration | Tests H4 and maps upstream dependence | CR₄, HHI, entry barriers and rent decomposition |
| Chapter 4 — Real Cost of Local AI | Tests the local component of H2 and H6 | Workload-specific five-year TCO |
| Chapter 5 — Cloud AI Economics | Tests H7 and H8 and cloud component of H2 | Cost per successful task and break-even tariffs |
| Chapter 6 — Enterprise Exposure | Translates H2, H7 and H8 into firm-level effects | AI cost burden, ROI and sector vulnerability |
| Chapter 7 — Capability Divide | Tests H5 and individual component of H6 | Educational and productivity treatment effects |
| Chapter 8 — Global Inequality | Tests country-level H6 | Global AI Access and Compute Sovereignty indices |
| Chapter 9 — Forecasts | Projects all supported relationships to 2031 | Scenario ranges, prediction intervals and warning indicators |
| Chapter 10 — Policy | Responds only to empirically supported failures | Costed interventions and measurable KPIs |
1.17 Preliminary Evidence Register
Table 1.10 — What Chapter 1 Establishes and What Remains Unproven
| Proposition | Current status | Evidence presently available | Required next test |
|---|---|---|---|
| AI compute depends on multiple complementary physical and software layers | Established conceptual fact | European Commission supply-chain and bottleneck analysis | Quantify each layer |
| AI infrastructure investment is exceptionally large | Established | Alphabet 2025 capex and European Gigafactory mobilisation | Compare firms, regions and historical baseline |
| Advanced fabrication and HPC demand are strategically important | Established | TSMC revenue composition | Global shares and alternative capacity |
| Energy can become a binding compute constraint | Strongly supported | DOE data-centre consumption scenarios | Regional power-price and grid-connection model |
| AI adoption differs substantially by enterprise size | Established correlation | Eurostat 2025 survey | Identify causal contributions of cost, skills and scale |
| Global connectivity remains highly unequal | Established | ITU 2025 indicators | Extend from connectivity to quality-adjusted AI access |
| AI hardware has become universally more expensive in real terms | Unproven | No complete quality-adjusted price index yet | Chapter 2 |
| Used AI hardware systematically trades at two or three times launch price | Unproven | Requires transaction-level evidence | Chapter 2 |
| Supplier concentration causes unlawful pricing | Unproven | Concentration and high margins are insufficient | Chapter 3 |
| Current token prices are unsustainable | Unproven | Provider-specific cost allocation is incomplete | Chapter 5 |
| Frontier access improves all educational and professional outcomes | Unproven and likely task-dependent | Requires causal outcome evidence | Chapter 7 |
| Lower-income populations are trapped permanently in inferior AI | Unproven | Connectivity and enterprise-adoption gaps support investigation | Chapter 8 |
| Compute will be traded like Bitcoin | Economically overstated | Some capacity can be reserved and dynamically priced | Chapters 5 and 9 |
1.18 Chapter Conclusion
AI compute should be understood as a hybrid economic institution: an industrial input when consumed, a capital good when owned, a utility-like service when delivered through the cloud, strategic infrastructure when essential functions depend upon it, a geopolitical resource when access is jurisdictionally controlled and a potentially financialised capacity when future availability is reserved or monetised. This hybrid status explains why conventional consumer-electronics analysis is insufficient. A laptop price cannot reveal the full cost of local AI; a token tariff cannot reveal long-run infrastructure economics; a global shipment figure cannot reveal politically accessible supply; and an internet-penetration rate cannot reveal access to frontier capability.
The initial evidence establishes four conditions that justify the report. First, capital mobilisation is occurring at a scale that can advantage a small number of firms and states able to finance infrastructure before demand is fully realised. Second, the chain contains complementary nodes—advanced logic, HBM, packaging, networking and power—where substitution can be slow. Third, enterprise adoption already varies sharply by company size, although causality has not yet been assigned. Fourth, global connectivity and income inequalities form a pre-existing structure through which AI inequality may be transmitted. None of these conditions proves artificial scarcity, monopolistic abuse or inevitable price escalation. The central scientific obligation of the remaining chapters is to determine whether the observed system reflects a temporary investment cycle followed by falling quality-adjusted costs, or a persistent architecture in which concentrated control and unequal purchasing power transform intelligence into a rationed economic resource.
The report’s eight hypotheses are deliberately falsifiable. H1 and H2 separate sticker-price inflation from the cost of fixed capability. H3 distinguishes durable scarcity from second-hand speculation. H4 separates concentration from demonstrable economic rent. H5 measures advantage rather than assuming it. H6 converts the idea of an AI divide into a comparative access distribution. H7 tests whether current tariffs recover long-run costs. H8 identifies whether metering is becoming multidimensional and dynamic. Together they create a research programme capable of confirming, rejecting or qualifying the report’s central warning without converting a legitimate concern into a predetermined conclusion.
You were right about the formatting error in Chapter 1. All mathematical subscripts below use HTML <sub> tags. Hypothesis labels such as H1, H2 and H3 remain ordinary text.
Chapter 2 — Hardware Inflation: GPUs, Memory, Computers and the Used Market
2.1 Scope, Evidentiary Standard and Central Finding
This chapter tests whether AI-relevant hardware became structurally less affordable between 2020 and 2026 and whether the observed increase represents nominal inflation, quality-adjusted inflation, temporary scarcity, regional price distortion or speculative resale behaviour. The first conclusion is methodologically decisive: no single “AI hardware price” exists. A consumer GPU, professional accelerator, data-centre server, high-memory Apple workstation and AI-capable laptop provide different combinations of memory capacity, bandwidth, compute throughput, power efficiency, software compatibility, portability, warranty and useful life. A valid price comparison must therefore preserve either the product configuration or the workload capability. Comparing the launch price of a 2020 GPU with the price of a technically different 2025 GPU without controlling for capability can exaggerate or conceal inflation. The chapter consequently separates four objects: the nominal purchase price; the general-inflation-adjusted price; the specification-adjusted price; and the cost of completing a constant AI workload.
The verified flagship NVIDIA sequence illustrates the difference. The GeForce RTX 3090 launched in September 2020 at $1,499, the RTX 4090 in October 2022 at $1,599, and the RTX 5090 in January 2025 at $1,999. The nominal flagship entry price therefore rose by 33.4% between the RTX 3090 and RTX 5090. However, memory capacity increased from 24 GB to 32 GB, while official memory bandwidth rose from approximately 936 GB/s to 1,792 GB/s. Once the launch prices are converted into July 2026 dollars using the U.S. Consumer Price Index, the estimated real launch price rises by approximately 9%, while the real price per GB of accelerator memory declines by approximately 18% and the real price per GB/s of bandwidth declines by more than 40%. These calculations do not prove that AI became cheaper in practical terms because memory capacity and bandwidth alone do not measure model quality, software overhead, context length, energy cost or tokens per second. They do demonstrate that the claim “high-end GPUs became two or three times more expensive” is not representative of the verified NVIDIA Founders Edition flagship MSRP sequence. Introducing GeForce RTX 30 Series Graphics Cards – NVIDIA – September 2020 NVIDIA Delivers Quantum Leap in Performance, Introduces New Era of Neural Rendering with GeForce RTX 40 Series – NVIDIA – September 2022 NVIDIA Blackwell GeForce RTX 50 Series Opens New World of AI Computer Graphics – NVIDIA – January 2025
That result applies only to launch MSRP and selected technical denominators. It does not describe actual street prices, availability during launch windows, board-partner premiums, taxes, import costs or second-hand transactions. Under the evidence protocol governing this report, an asking price cannot be represented as a transaction, a single listing cannot establish a median, and a marketplace advertisement cannot prove that a product was sold. No official manufacturer, government statistical agency or audited corporate filing publishes a complete 2020–2026 international series of transaction-level used GPU prices. Consequently, the used-market medians requested for the United States, EU, United Kingdom, China, Japan, India, Brazil and lower-income markets are marked ND — not disclosed or not verifiable under the permitted source hierarchy. This is not a missing detail that can be responsibly filled with estimates: it is a material data limitation. The claim that ordinary used AI hardware systematically sells at two or three times original MSRP remains unproven until completed-transaction microdata are obtained, cleaned and audited.
2.2 Price Concepts and Measurement Architecture
Hardware inflation must be decomposed into at least seven components:
- general monetary inflation;
- change in technical quality;
- change in product position within the manufacturer’s range;
- changes in taxes, tariffs and exchange rates;
- temporary scarcity;
- distributor or reseller margin;
- speculative or exceptional resale premiums.
The observed local price can be represented as:
Pobserved,i,r,t = Pbase,i,t × FXr,t × (1 + Taxr,t) × (1 + Tariffr,t) × (1 + Distributioni,r,t) × (1 + Scarcityi,r,t) × (1 + Speculationi,r,t)
Where:
- Pobserved,i,r,t is the observed price of product i in region r and period t;
- Pbase,i,t is the manufacturer reference price;
- FXr,t is the applicable exchange-rate conversion factor;
- Taxr,t includes VAT, sales tax and other mandatory charges;
- Tariffr,t includes customs duties and import charges;
- Distributioni,r,t is the wholesale and retail margin;
- Scarcityi,r,t measures the premium associated with supply-demand imbalance;
- Speculationi,r,t is the residual premium not explained by ordinary distribution and scarcity.
This multiplicative formulation prevents taxes and currency changes from being incorrectly classified as manufacturer inflation. It also avoids treating the U.S. pre-sales-tax MSRP as directly comparable with an EU consumer price that normally includes VAT. A €2,000 EU price including 22% VAT is not economically equivalent to a $2,000 U.S. price excluding sales tax. Geographic comparisons must remove recoverable taxes for enterprise purchasers, retain taxes for household affordability analysis and use both market exchange rates and purchasing-power-parity conversions where appropriate.
Table 2.1 — Mandatory Price Fields and Evidentiary Status
| Field | Definition | Acceptable evidence | Current status |
|---|---|---|---|
| Launch date | First official commercial availability | Manufacturer announcement | Verifiable |
| Launch MSRP or SEP | Manufacturer’s recommended launch price | Official manufacturer announcement | Verifiable for many consumer products |
| Current official price | Manufacturer’s currently displayed direct price | Live manufacturer store or official product page | Verifiable only while published |
| Street price | Actual authorised-retailer transaction price | Audited retailer transactions or official statistical scanner data | Generally unavailable |
| Used asking price | Seller’s requested amount | Marketplace listing | Observable but not a completed transaction |
| Used transaction price | Amount paid in a completed sale | Completed-sale microdata | Not available within current source constraints |
| Used-market median | Median of cleaned completed transactions | Representative transaction dataset | ND |
| Inflation-adjusted price | Price converted into constant-period currency | Official national CPI or HICP | Calculable |
| Quality-adjusted price | Price divided by a stable capability measure | Verified price and specification data | Partly calculable |
| Workload-adjusted price | Cost per successful constant workload | Reproducible benchmark and complete TCO | Requires Chapter 4 testing |
| Warranty status | Remaining transferable manufacturer coverage | Serial or invoice-level evidence | Cannot be inferred for aggregate used listings |
| Regional premium | Local net-of-tax price relative to common baseline | Official regional prices, FX and tax data | Requires configuration matching |
2.3 Verified Consumer GPU Price History, 2020–2026
2.3.1 NVIDIA Flagship Sequence
The following table uses only manufacturer-published launch prices and technical specifications. The conversion into July 2026 dollars uses the U.S. CPI-U all-items index. September 2020, October 2022 and January 2025 CPI values are aligned with the respective launch months, while July 2026 is the common valuation month. The Bureau of Labor Statistics reports a July 2026 CPI-U level of 333.918, based on 1982–1984 = 100. Consumer Price Index Summary, July 2026 – U.S. Bureau of Labor Statistics – August 2026 The inflation-adjustment formula is:
P2026,i = Plaunch,i × CPIJuly 2026 ÷ CPIlaunch month,i
Table 2.2 — Verified NVIDIA Flagship Launch History
| Product | Official launch | U.S. launch MSRP | Memory | Official/derived bandwidth | Board power | July 2026-dollar MSRP | Real price per GB | Real price per GB/s |
|---|---|---|---|---|---|---|---|---|
| GeForce RTX 3090 | 24 Sep 2020 | $1,499 | 24 GB GDDR6X | ≈936 GB/s | 350 W | ≈$1,923 | ≈$80.13 | ≈$2.05 |
| GeForce RTX 4090 | 12 Oct 2022 | $1,599 | 24 GB GDDR6X | 1,008 GB/s | 450 W | ≈$1,792 | ≈$74.67 | ≈$1.78 |
| GeForce RTX 5090 | 30 Jan 2025 | $1,999 | 32 GB GDDR7 | 1,792 GB/s | 575 W | ≈$2,102 | ≈$65.69 | ≈$1.17 |
Interpretation: the nominal launch price increased by 33.4% from the RTX 3090 to RTX 5090. In common July 2026 dollars, the estimated increase is approximately 9%. The real price per installed GB declined from approximately $80 to $66, and the real price per GB/s of memory bandwidth declined from approximately $2.05 to $1.17. The absolute power envelope increased substantially, however, meaning purchase-price efficiency cannot be treated as equivalent to operating-cost efficiency. The RTX 5090 provides more memory and bandwidth but requires a more demanding power supply, creates greater heat density and may raise total system cost. Its 32 GB capacity also remains far below the memory required for many high-parameter models at higher precision or with large KV caches.
NVIDIA officially specifies the RTX 5090 with 32 GB GDDR7 and 1,792 GB/s memory bandwidth. The RTX 4090 carries 24 GB GDDR6X, 1,008 GB/s bandwidth and a 450 W total graphics power rating. GeForce RTX 5090 Specifications – NVIDIA – January 2025 GeForce RTX 4090 Specifications – NVIDIA – September 2022 The RTX 3090 launch announcement records 24 GB GDDR6X and a starting price of $1,499. Introducing GeForce RTX 30 Series Graphics Cards – NVIDIA – September 2020
Table 2.3 — Change Across NVIDIA Flagship Generations
| Metric | RTX 3090 → RTX 4090 | RTX 4090 → RTX 5090 | RTX 3090 → RTX 5090 |
|---|---|---|---|
| Nominal launch MSRP | +6.7% | +25.0% | +33.4% |
| Installed memory | 0.0% | +33.3% | +33.3% |
| Memory bandwidth | ≈+7.7% | +77.8% | ≈+91.5% |
| Board power | +28.6% | +27.8% | +64.3% |
| Real launch price in July 2026 dollars | ≈−6.8% | ≈+17.3% | ≈+9.3% |
| Real price per GB | ≈−6.8% | ≈−12.0% | ≈−18.0% |
| Real price per GB/s | ≈−13.2% | ≈−34.3% | ≈−43.0% |
The RTX sequence therefore supports a nuanced rather than categorical conclusion. The cost of entering the current flagship category rose sharply in nominal dollars, but specification-adjusted affordability improved on memory-capacity and bandwidth measures. At the same time, minimum absolute expenditure rose: a buyer still needs approximately $2,000 before tax for the reference flagship GPU alone, excluding the computer, high-capacity system memory, storage, cooling and power supply. Quality-adjusted deflation can coexist with worsening access inequality because a lower price per unit of capability does not help a household that cannot finance the larger minimum purchase.
2.4 AMD Consumer GPU Counterfactual
AMD provides an essential competitive comparator because changes in the NVIDIA flagship series cannot establish market-wide inflation. The Radeon RX 7900 XTX was announced with a suggested e-tail price of $999 and 24 GB GDDR6, while the RX 9070 XT launched with a suggested e-tail price of $599 and 16 GB GDDR6. AMD lists the RX 7900 XTX at 355 W total board power and the RX 9070 XT at 304 W. The RX 9070 XT has an official memory bandwidth of up to 640 GB/s. AMD Unveils World’s Most Advanced Gaming Graphics Cards, Built on Groundbreaking AMD RDNA 3 Architecture – AMD – November 2022 AMD Unveils Next-Generation RDNA 4 Architecture with the Launch of AMD Radeon RX 9000 Series Graphics Cards – AMD Investor Relations – February 2025 Radeon RX 9070 XT Specifications – AMD – 2025
Table 2.4 — Verified AMD Reference Products
| Product | Launch availability | U.S. SEP | Memory | Bandwidth | Board power | Nominal price per GB | Nominal price per GB/s |
|---|---|---|---|---|---|---|---|
| Radeon RX 7900 XTX | 13 Dec 2022 | $999 | 24 GB GDDR6 | 960 GB/s | 355 W | $41.63 | $1.04 |
| Radeon RX 7900 XT | 13 Dec 2022 | $899 | 20 GB GDDR6 | 800 GB/s | 315 W | $44.95 | $1.12 |
| Radeon RX 9070 XT | 6 Mar 2025 | $599 | 16 GB GDDR6 | 640 GB/s | 304 W | $37.44 | $0.94 |
| Radeon RX 9070 | 6 Mar 2025 | $549 | 16 GB GDDR6 | 640 GB/s | 220 W | $34.31 | $0.86 |
These specification ratios must not be interpreted as equivalent LLM performance. NVIDIA and AMD differ in software maturity, supported data types, kernel optimisation, framework integration and application compatibility. A nominally lower price per GB is not an economic saving if the required application cannot use the hardware efficiently. Conversely, excluding non-NVIDIA hardware because the dominant software ecosystem favours NVIDIA would obscure the extent to which ecosystem lock-in, rather than physical semiconductor cost, determines effective price. Chapter 4 will therefore measure both hardware-neutral physical capability and workload-specific realised capability.
2.5 Professional Workstation GPUs
Professional GPUs differ from consumer products through certified drivers, memory capacity, error protection, virtualisation support, enterprise servicing, warranties and application certification. Their prices cannot be compared directly with GeForce or Radeon gaming cards. A professional card may provide lower nominal compute per dollar yet be economically preferable where downtime, certification or memory integrity has a high expected cost. The relevant formula is:
EACi = PurchasePricei + ExpectedDowntimeLossi + FailureRiskCosti + SoftwareIncompatibilityCosti − WarrantyValuei
Where EAC is the expected adjusted cost. A consumer GPU can appear substantially cheaper until the cost of unsupported deployment, business interruption or limited memory is included.
Table 2.5 — Professional GPU Data Requirements
| Variable | Consumer card | Professional card | Required adjustment |
|---|---|---|---|
| Memory capacity | Usually lower | Often materially higher | Price per usable GB |
| ECC or error protection | Limited or product-dependent | More commonly supported | Expected error-cost adjustment |
| Driver certification | Consumer applications | Professional application certification | Workflow-compatibility value |
| Virtualisation | Restricted or limited | Enterprise-oriented | Multi-user utilisation value |
| Warranty | Consumer terms | Professional support terms | Expected service value |
| Availability horizon | Shorter retail generation | Often longer enterprise lifecycle | Replacement and fleet standardisation |
| Price disclosure | Usually public MSRP | Often reseller or quotation based | Official price gaps must remain ND |
| LLM suitability | Strong if software-supported | Stronger for memory-intensive workloads | Tokens per second and concurrency test |
A complete professional-GPU price history cannot be created from audited corporate filings because manufacturers do not consistently publish configuration-specific transaction prices. Quotation-only prices must not be converted into invented MSRPs. Chapter 2 therefore treats professional price measurement as a controlled data-acquisition task rather than filling the table with unverified retailer values.
2.6 Data-Centre Accelerators and AI Servers
Data-centre accelerators pose the greatest pricing-transparency problem. Public discussion frequently assigns unit prices to products such as NVIDIA H100, H200, B200 or AMD Instinct accelerators, but manufacturer transaction prices depend on volume, system configuration, networking, support, delivery terms and customer agreements. The price of a single accelerator is economically incomplete because large training and inference systems are sold as integrated platforms containing processors, HBM, networking, switches, cooling, storage, racks and software. Under the report’s source rules, unofficial estimates cannot be presented as verified prices.
Table 2.6 — Data-Centre Accelerator Price Transparency
| Product category | Official technical specifications | Public manufacturer unit MSRP | Observable audited average transaction price | Analytical treatment |
|---|---|---|---|---|
| Individual AI accelerator | Usually available | Generally ND | ND | Specifications only |
| Multi-GPU server | Configuration available | Sometimes quotation-based | ND | Full-system quotation required |
| Integrated rack | Architecture available | Usually ND | ND | Price per delivered capacity |
| Cloud accelerator instance | Public tariff often available | Service price, not hardware price | Partly observable | Cost per accelerator-hour |
| Reserved cloud capacity | Contract-specific | Rarely fully public | ND | Scenario or disclosed-contract analysis |
| Sovereign AI cluster | Procurement-specific | Sometimes public through tender | Potentially observable | Tender-by-tender comparison |
For data-centre hardware, the preferred price unit is not dollars per accelerator. It is the levelised cost per successful workload:
LCWw = [CAPEX × CRF + OPEXannual] ÷ SuccessfulWorkloadsannual,w
The capital-recovery factor is:
CRF = r(1 + r)n ÷ [(1 + r)n − 1]
Where r is the annual discount rate and n the economic life in years. CAPEX includes accelerator systems, networking, facility modifications and installation. OPEX includes power, cooling, maintenance, software, specialised labour, insurance and downtime. A lower accelerator purchase price may not reduce LCW if utilisation, software support or networking performance is inferior.
NVIDIA’s fiscal 2026 annual filing identifies the company as a full-stack computing-infrastructure provider and states that it depends on foundries and other suppliers for manufacturing. It does not provide a standard realised unit price for data-centre GPUs. NVIDIA Annual Report on Form 10-K for the Fiscal Year Ended January 25, 2026 – NVIDIA and U.S. Securities and Exchange Commission – February 2026 This absence is itself economically relevant: enterprise buyers operate in a negotiated market while households purchase transparent retail products.
2.7 Memory Inflation: HBM, DRAM and NAND
Memory must be separated into three markets. Conventional DRAM supplies system memory; NAND supplies persistent storage; HBM supplies the extreme bandwidth required by high-performance accelerators. Their prices do not move identically because their production complexity, demand, qualification and end markets differ. HBM consumes more wafer capacity per delivered bit than conventional DRAM and requires advanced packaging integration. A shift in producer capacity toward HBM can therefore tighten conventional-memory supply even when total memory investment increases.
The minimum measurement structure is:
MPIm,t = Σwj,0 × Pj,t ÷ Pj,0 × 100
Where MPI is the memory price index for memory class m, P is the verified contract or transaction price and w is the base-period quantity weight. A chained index is preferable when product generations change rapidly. An unweighted average of advertised module prices would be invalid because it would overrepresent products with many listings and ignore transaction volumes.
Table 2.7 — Memory-Market Measurement Framework
| Memory segment | Principal AI role | Appropriate unit | Required quality controls | Primary scarcity indicator |
|---|---|---|---|---|
| HBM2e/HBM3/HBM3E/HBM4 | Accelerator bandwidth and capacity | Price per stack, GB and GB/s | Generation, stack height, bandwidth and qualification | Contracted capacity and delivery lead time |
| Server DRAM | CPU-side model hosting, preprocessing and databases | Price per GB | DDR generation, speed, ECC and module type | Contract price and inventory |
| Consumer DRAM | Local workstation system memory | Price per GB | DDR generation, speed and latency | Retail transaction price |
| Enterprise NAND | Model storage, datasets and checkpointing | Price per TB and endurance unit | Interface, endurance, latency and warranty | Contract price and utilisation |
| Consumer NAND | Local model storage | Price per TB | Interface, endurance and performance | Retail transaction price |
Micron forecast in December 2025 that the HBM total addressable market could grow from approximately $35 billion in 2025 to around $100 billion in 2028, implying an approximate 40% compound annual growth rate. This is a company forecast, not an observed market outcome. Micron Fiscal First Quarter 2026 Financial Results – Micron Technology – December 2025 Samsung announced plans to invest more than KRW 110 trillion in facilities and R&D during 2026, including HBM, foundry operations and advanced packaging. Samsung Electronics Shareholder Value Enhancement Plan – Samsung Electronics – March 2026 These disclosures support rapid demand and investment but do not supply a public, auditable HBM price-per-GB history. Accordingly, the chapter does not invent an HBM price series.
Table 2.8 — What Can Be Verified for Memory
| Requested variable | HBM | DRAM | NAND |
|---|---|---|---|
| Official product generation | Yes | Yes | Yes |
| Manufacturer capacity commentary | Yes | Yes | Yes |
| Audited segment revenue | Often | Often | Often |
| Public transaction-weighted unit price | Generally no | Generally no | Generally no |
| Public international street-price median | No | No | No |
| Completed used-market median | Not meaningful for chips; no | No | No |
| Quality-adjusted public index | Requires external transaction dataset | Requires dataset | Requires dataset |
| Five-year projection | Scenario-based only | Scenario-based only | Scenario-based only |
2.8 CPUs, NPUs and the Problem of Incomparable AI Throughput
The emergence of integrated NPUs changes the definition of an AI-capable computer. A laptop may perform certain local inference tasks without a discrete GPU, but NPU TOPS cannot be compared directly with GPU TFLOPS or accelerator AI TOPS. The numerical value depends on precision, sparsity, supported operators, memory architecture and workload. A 50-TOPS NPU does not necessarily deliver one-fiftieth of a 2,500-TOPS accelerator’s useful performance, because the larger device may use a different precision, assume structured sparsity or perform workloads unsupported by the NPU.
The general quality-adjusted CPU/NPU price is:
PCPU-NPU,i,tQA = Psystem,i,t ÷ G(QCPU, QNPU, M, B, E, W)
Where G is a workload-specific capability aggregator, M is accessible memory, B bandwidth, E energy efficiency and W software compatibility. A geometric mean can prevent one very high component from dominating:
Gi = ∏Qk,iwₖ
With Σwk = 1.
Table 2.9 — CPU/NPU Comparison Controls
| Control | Reason |
|---|---|
| Precision | INT4, INT8, FP8, FP16 and FP32 values are not interchangeable |
| Sparsity | Some advertised throughput assumes structured sparsity |
| Sustained power | Peak throughput may not be maintained in thin laptops |
| Accessible memory | Shared memory may be large but bandwidth-limited |
| Operator support | Unsupported operators may fall back to CPU or GPU |
| Framework maturity | Theoretical hardware may lack practical software paths |
| Batch size | Throughput and latency respond differently |
| Model type | Vision, audio and language workloads use hardware differently |
| Thermal condition | Battery and plugged-in performance may diverge |
| System price | Integrated NPU cost cannot be isolated reliably from the device |
2.9 Apple Computers: Price, Memory and Configuration Drift
Apple systems require configuration-level rather than brand-level analysis. Statements that “Apple raised all prices” are not scientifically acceptable unless identical or hedonic configurations are compared by geography and date. Apple has sometimes raised prices, held starting prices stable, reduced them or increased included memory. For example, the 2023 M2 Mac mini started at $599, while the 2024 M4 iMac started at $1,299 with 16 GB unified memory. These products are not longitudinal substitutes, but they demonstrate why company-wide price claims require product-family datasets. Apple Introduces New Mac mini with M2 and M2 Pro – Apple – January 2023 Apple Introduces New iMac Supercharged by M4 and Apple Intelligence – Apple – October 2024
The 2023 14-inch MacBook Pro with M2 Pro started at $1,999, while the 16-inch version started at $2,499. The M2 Max supported up to 96 GB unified memory with 400 GB/s bandwidth. Apple Unveils MacBook Pro Featuring M2 Pro and M2 Max – Apple – January 2023 The 2025 Mac Studio with M4 Max could be configured with 128 GB unified memory, while the M3 Ultra version provided a different higher-memory architecture. Mac Studio 2025 Technical Specifications – Apple – March 2025
Table 2.10 — Apple AI-Workstation Data Schema
| Field | Required configuration detail |
|---|---|
| Product family | Mac mini, Mac Studio, MacBook Pro, iMac |
| Chip | Exact generation and tier |
| CPU/GPU cores | Exact selected configuration |
| Unified memory | Installed capacity |
| Memory bandwidth | Chip-specific bandwidth |
| Storage | Installed SSD capacity |
| Launch price | Price for identical configuration |
| Current price | Same geography, tax treatment and configuration |
| AI workload | Model, quantisation, context and inference engine |
| Tokens per second | Median sustained result after warm-up |
| Energy | Wall-power measurement |
| Warranty | Standard and extended coverage |
| Upgradeability | Memory and storage replacement constraints |
| Residual value | Verified completed resale transaction only |
The economic attraction of high-memory Apple systems is that unified memory can permit models too large for a 24 GB or 32 GB consumer GPU. The limitation is that capacity does not establish speed, framework compatibility or enterprise manageability. Price per GB therefore favours high-memory unified systems only when the workload can exploit the architecture at an acceptable rate.
2.10 Samsung Devices and Relevant Competitors
Samsung spans memory production, mobile devices, laptops and semiconductor manufacturing, making it economically distinct from a pure device vendor. However, the company does not publish a single globally comparable AI-computer price series in audited financial reports. Model names, processors, memory configurations, launch timing, promotional bundles and taxes differ by country. A valid Samsung comparison must therefore use exact SKU-level official launch prices and exclude temporary trade-in credits unless trade-in value is separately modelled.
Table 2.11 — Required Samsung and Competitor Sample
| Category | Samsung product family | Required competitors | Comparison unit |
|---|---|---|---|
| Premium AI laptop | Galaxy Book Ultra/Pro class | Apple MacBook Pro, Dell Precision/XPS, Lenovo ThinkPad/Legion, HP ZBook | Cost per successful local workload |
| Mobile AI device | Galaxy S Ultra class | Apple iPhone Pro, Google Pixel Pro | On-device AI capability per real price |
| Memory | Samsung DRAM/NAND/HBM | Micron, SK hynix | Contract-price index by memory generation |
| Enterprise storage | Samsung enterprise SSD | Micron and other qualified vendors | Cost per TB adjusted for endurance |
| AI workstation component | Samsung memory-equipped systems | Competing memory configurations | Cost per usable GB and bandwidth |
No Samsung price increase is treated as established unless an identical configuration and geography are documented at two dates. Product-mix changes—such as more memory, different processors or bundled AI features—must be removed through hedonic adjustment.
2.11 High-Performance Laptops and Desktop Workstations
Laptop prices present additional measurement problems because the GPU name does not fully specify performance. Mobile GPU power limits vary by manufacturer and chassis; a laptop GPU with the same family designation as another may run at a different power envelope. Cooling, system memory, screen, battery, storage and warranty also affect price. The correct unit is a complete SKU, not a nominal GPU class.
Table 2.12 — Mandatory Laptop Controls
| Variable | Measurement requirement |
|---|---|
| GPU model | Exact mobile GPU designation |
| GPU power | Sustained configured power, not only maximum permitted range |
| VRAM | Capacity and type |
| System RAM | Capacity, bandwidth and upgradeability |
| CPU/NPU | Exact model and enabled power profile |
| Cooling | Sustained benchmark after thermal equilibrium |
| Battery mode | Separate from plugged-in result |
| Display | Resolution and refresh rate as price controls |
| Storage | Capacity and performance |
| Weight | Portability control |
| Warranty | Standardised expected service value |
| AI benchmark | Same model, precision, context and inference engine |
| Geography | Net-of-tax and tax-inclusive price |
| Promotions | Recorded separately from standard price |
For desktop workstations, the chassis, motherboard, power supply, cooling, memory, storage and professional support must be added to GPU cost. A $1,999 GPU does not create a $1,999 local AI system. The workstation price is:
Pworkstation = PGPU + PCPU + PRAM + Pstorage + Pboard + PPSU + Pcooling + Pcase + POS + Passembly + Pwarranty
2.12 Quality-Adjusted Price Measures
No single denominator captures AI capability. This chapter therefore adopts six complementary measures.
2.12.1 Price per Usable Accelerator-Memory GB
Pi,tVRAM = Pi,t ÷ Musable,i,t
Usable memory must exclude operating overhead, display reservation and unusable fragmentation. Installed memory is acceptable only for the preliminary hardware table.
2.12.2 Price per GB/s of Memory Bandwidth
Pi,tBW = Pi,t ÷ Bi,t
This is relevant for memory-bound inference but does not capture compute, caching or software.
2.12.3 Price per Comparable TFLOP
Pi,tFLOP = Pi,t ÷ Fi,t,p
The precision p must be identical. FP32 cannot be divided by an FP16 or FP8 throughput value.
2.12.4 Price per Token per Second
Pi,t,m,cTPS = Psystem,i,t ÷ TPSi,m,c
Where m is the model and c the complete benchmark configuration.
2.12.5 Price per Benchmark Unit
Pi,tBENCH = Pi,t ÷ Scorei,t
Only independently reproducible benchmarks with identical settings qualify.
2.12.6 Price per Successfully Completed Workload
Pi,wSUCCESS = TCOi,w ÷ Nsuccessful,i,w
This is the preferred economic metric because it includes reliability, retries, supervision and total cost.
Table 2.13 — Strengths and Limitations of Quality Metrics
| Metric | Principal strength | Principal weakness |
|---|---|---|
| Price per installed GB | Simple and relevant to model fit | Ignores speed and usable capacity |
| Price per usable GB | Better model-capacity measure | Requires runtime-level measurement |
| Price per GB/s | Captures memory movement | Ignores compute and software |
| Price per TFLOP | Captures arithmetic throughput | Precision and sparsity can mislead |
| Price per token/s | Direct inference measure | Model and settings specific |
| Price per benchmark point | Enables composite comparison | Benchmark design can bias results |
| Price per successful task | Closest to economic utility | Expensive to measure and workload-specific |
2.13 Hedonic Price Model
The principal hedonic specification is:
ln(Pi,r,t) = α + Σk=1KβkXk,i,t + γt + δb + θr + λc + εi,r,t
Where:
- Pi,r,t is the net-of-tax transaction price;
- Xk,i,t contains technical characteristics;
- γt is the time fixed effect;
- δb is the brand fixed effect;
- θr is the region fixed effect;
- λc is the condition class: new, refurbished or used;
- εi,r,t is the residual.
The feature vector will include:
- usable accelerator memory;
- memory bandwidth;
- comparable throughput by precision;
- power consumption;
- age;
- product class;
- warranty;
- system RAM;
- storage;
- portability;
- professional certification;
- software ecosystem;
- condition;
- seller type;
- regional tax;
- exchange rate;
- verified availability.
The time dummy exponent produces a quality-adjusted price index:
HPIt = 100 × exp(γt − γbase)
If the model is estimated in log form, percentage interpretation must use exp(β) − 1 rather than treating large coefficients as exact percentages.
Table 2.14 — Model Diagnostics
| Diagnostic | Test | Required response |
|---|---|---|
| Heteroskedasticity | Breusch–Pagan or White | Robust standard errors |
| Multicollinearity | VIF and condition number | Combine or remove redundant hardware variables |
| Serial correlation | Time-clustered residual test | Cluster by product family and period |
| Brand endogeneity | Fixed effects and matched products | Report sensitivity |
| Sample selection | Availability and seller controls | Weight or model selection |
| Functional form | RESET and spline comparison | Use nonlinear terms where justified |
| Outliers | Leverage and Cook’s distance | Verify, winsorise only under disclosed rule |
| Missing data | Missingness map | Multiple imputation only where defensible |
| Product turnover | Chained index | Avoid comparing unmatched generations directly |
| Benchmark drift | Fixed workload suite | Freeze model and settings within test window |
2.14 Inflation Adjustment and Geographic Comparability
The real price in reference-period currency is:
Pi,t→Treal = Pi,t × CPIT ÷ CPIt
The real price change is:
πi,t→Treal = [Pi,Tnominal ÷ Pi,t→Treal − 1] × 100
For cross-border comparison:
Pi,r,tcommon = [Pgross,i,r,t ÷ (1 + VATr,t)] ÷ FXr,t
Household affordability retains VAT:
HAIi,r,t = Pgross,i,r,t ÷ MedianMonthlyDisposableIncomer,t
Enterprise affordability removes recoverable VAT where applicable:
EAIi,r,t = Pnet,i,r,t ÷ MedianMonthlyEnterpriseValueAddedr,t
Table 2.15 — Geographic Normalisation Protocol
| Market | Consumer inflation source | Price treatment | Currency treatment | Principal distortion |
|---|---|---|---|---|
| United States | BLS CPI-U | Usually exclude sales tax from MSRP, add local tax for affordability | USD base | State and local sales-tax variation |
| European Union | Eurostat HICP | Record VAT-inclusive and net-of-VAT | EUR and local currency where applicable | Different VAT and regional official pricing |
| United Kingdom | ONS CPI | VAT-inclusive and net-of-VAT | GBP, then common currency | GBP movements and VAT |
| China | Official national CPI required | Include applicable taxes | CNY and common currency | Local SKU and channel differences |
| Japan | Official national CPI required | Include consumption tax | JPY and common currency | Yen volatility |
| India | Official national CPI required | Include GST and import effects | INR and PPP comparison | Import duties and income differential |
| Brazil | Official national CPI required | Include indirect taxes and import costs | BRL and PPP comparison | Tax complexity and currency volatility |
| Lower-income sample | National statistics or official international series | Consumer gross price | Local currency, USD and PPP | Thin formal distribution and income constraint |
Eurostat defines the HICP as a comparable measure of changes in household consumer prices across European countries. Harmonised Indices of Consumer Prices – Eurostat – 2026 The U.S. Bureau of Labor Statistics maintains a dedicated index for computers, peripherals and smart-home-assistant devices and applies quality adjustment to changing computer products. How BLS Measures Price Change for Computers, Peripherals and Smart Home Assistant Devices – U.S. Bureau of Labor Statistics – February 2026
2.15 Used and Refurbished Market
The used-market premium must be calculated only from completed, comparable transactions:
UMPi,r,t = [Pused,i,r,t − Preference,i,r,t] ÷ Preference,i,r,t × 100
The reference price depends on the research question:
- launch MSRP tests whether used price exceeds original nominal price;
- inflation-adjusted MSRP tests whether it exceeds original real price;
- contemporaneous new price tests whether used hardware is irrationally priced relative to new supply;
- depreciation-adjusted fair value tests whether scarcity offsets ordinary ageing;
- replacement-capability price tests whether the old product is cheaper than an equivalent current product.
A two-times-MSRP claim is:
MultiplierMSRP = Pused ÷ Plaunch MSRP
But this is not sufficient evidence of market inflation. A discontinued limited product may have collector value; a listing may never sell; the product may include a complete system; tax may be included in one price and excluded in the other; or the listing may be fraudulent.
Table 2.16 — Used-Market Cleaning Rules
| Issue | Required treatment |
|---|---|
| Asking versus sold price | Retain only completed transactions |
| Duplicate listing | Deduplicate by identifier, seller, image and timestamp |
| Bundle | Separate GPU-only from complete system |
| Condition | New-sealed, open-box, refurbished, used-working and parts-only |
| Warranty | Record transferable remaining coverage |
| Fraud/anomaly | Exclude under a disclosed, reproducible rule |
| Shipping | Separate product price from delivery |
| Tax | Record separately |
| Currency | Convert using transaction-date official FX |
| Board partner | Record exact manufacturer and model |
| Factory overclock | Treat as product attribute |
| Location | Use actual shipping origin and destination |
| Transaction volume | Report count and confidence interval |
| Thin markets | Suppress median below minimum observation threshold |
Table 2.17 — Statistical Test for Persistent Used Scarcity
| Test | Null hypothesis | Evidence supporting persistent scarcity |
|---|---|---|
| Median premium test | Median UMP ≤ 0 | Median UMP significantly above zero |
| Duration test | Premium lasts no longer than launch period | Premium persists beyond defined threshold |
| Cross-region test | Premium is geographically isolated | Premium observed across multiple tax-adjusted markets |
| Availability test | Premium unrelated to new-product availability | Premium rises when verified availability falls |
| Transaction-volume test | Premium driven by thin trades | Premium remains with adequate volume |
| Event study | Premium unrelated to supply event | Premium responds to launch, export or supply shock |
| Mean-reversion test | Premium is persistent | Slow or incomplete reversion |
| Replacement test | Used product is not scarce in capability terms | Equivalent new alternatives remain more expensive or unavailable |
Determination
Under the current source restrictions, the statement that AI-capable used products generally sell for two or three times original MSRP cannot be verified. The available primary sources establish MSRP and specifications, but not representative used transaction medians. The claim must therefore remain unconfirmed, not accepted and not rejected. Isolated cases may exist, especially during launch shortages or in geographically constrained markets, but isolated cases cannot be generalised to the entire used market.
2.16 Warranty and Risk-Adjusted Used Price
A used product’s nominal price understates its economic cost if warranty, failure probability and remaining useful life differ from a new product. The risk-adjusted price is:
Pused,irisk = Ptransaction,i + pfailure,i × Lfailure,i + Cdowntime,i − Vwarranty,i
Where:
- pfailure,i is expected failure probability;
- Lfailure,i is replacement or repair loss;
- Cdowntime,i is the economic cost of unavailability;
- Vwarranty,i is remaining warranty value.
For a student, downtime may mean lost study time. For a hospital, defence contractor or financial institution, it may be unacceptable. The same used GPU can therefore have different risk-adjusted prices for different users.
2.17 Preliminary 2020–2026 Findings
Table 2.18 — Evidence-Based Findings
| Question | Preliminary result | Confidence |
|---|---|---|
| Did flagship NVIDIA MSRP rise nominally? | Yes: $1,499 to $1,999, approximately +33.4% | High |
| Did flagship NVIDIA MSRP rise after general inflation? | Moderately, approximately +9% using common July 2026 dollars | Medium-high |
| Did real price per GB rise? | No in the selected flagship sequence; it declined | High for specification ratio |
| Did real price per GB/s rise? | No in the selected flagship sequence; it declined materially | High for specification ratio |
| Did absolute minimum flagship expenditure rise? | Yes | High |
| Did power requirements rise? | Yes, substantially | High |
| Did practical LLM cost necessarily decline? | Not established | Open |
| Did AMD maintain lower reference prices in selected products? | Yes, but software and workloads are not equivalent | High |
| Did HBM demand and investment rise? | Strong corporate evidence supports rapid expansion | High |
| Is there a verified public HBM transaction-price history? | No under permitted sources | High |
| Are data-centre accelerator unit prices transparent? | Generally no | High |
| Are international used-market medians verified? | No | High |
| Are two-times or three-times used prices representative? | Not demonstrated | Open |
The strongest finding is therefore not universal hardware inflation but a widening divergence between absolute access cost and quality-adjusted unit cost. A flagship GPU can deliver more memory bandwidth per real dollar while remaining inaccessible to more households because its minimum purchase price has increased. This is the threshold effect:
Accessi = 1 if Wealthi or CreditCapacityi ≥ MinimumSystemCost; otherwise 0.
A falling price per unit of capability does not reduce exclusion when the indivisible minimum system remains above the buyer’s budget.
2.18 Projection Framework, 2027–2031
A scientifically valid five-year forecast cannot be produced by extending MSRP alone. The model must forecast:
- nominal hardware price;
- general inflation;
- quality growth;
- memory capacity;
- bandwidth;
- power;
- availability;
- software compatibility;
- workload performance;
- depreciation;
- used-market liquidity;
- regional taxation and FX.
The quality-adjusted price evolves as:
Pt+1QA = PtQA × (1 + gnominal,t) ÷ [(1 + πt)(1 + gquality,t)]
A supply-demand specification is:
Δln(Pt) = α + β₁Δln(Demandt) − β₂Δln(Capacityt) + β₃EnergyShockt + β₄TradeShockt + β₅FXShockt + εt
The 2027–2031 values below are not claimed as observed or externally certified forecasts. They are transparent scenario indices with 2026 = 100, intended to show how different annual quality-adjusted price paths compound.
Scenario assumptions
- Abundant-supply scenario: quality-adjusted price declines 10% annually.
- Central-efficiency scenario: quality-adjusted price declines 3% annually.
- Persistent-scarcity scenario: quality-adjusted price rises 8% annually.
- Geopolitical-fragmentation scenario: quality-adjusted price rises 12% annually in import-dependent markets.
- Dual-market scenario: frontier hardware rises 8% annually while mainstream capability declines 8% annually.
Table 2.19 — Quality-Adjusted Hardware Price Index, 2026 = 100
| Scenario | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Abundant supply | 100.0 | 90.0 | 81.0 | 72.9 | 65.6 | 59.0 |
| Central efficiency | 100.0 | 97.0 | 94.1 | 91.3 | 88.5 | 85.9 |
| Persistent scarcity | 100.0 | 108.0 | 116.6 | 126.0 | 136.0 | 146.9 |
| Geopolitical fragmentation | 100.0 | 112.0 | 125.4 | 140.5 | 157.4 | 176.2 |
| Mainstream side of dual market | 100.0 | 92.0 | 84.6 | 77.9 | 71.6 | 65.9 |
| Frontier side of dual market | 100.0 | 108.0 | 116.6 | 126.0 | 136.0 | 146.9 |
Table 2.20 — Scenario Interpretation
| Scenario | Semiconductor condition | Memory condition | Used market | Access consequence |
|---|---|---|---|---|
| Abundant supply | Capacity expansion exceeds demand | HBM and DRAM supply normalises | Rapid depreciation | Broadening access |
| Central efficiency | Supply broadly tracks demand | Periodic but manageable constraint | Normal depreciation with launch spikes | Gradual affordability improvement |
| Persistent scarcity | Demand repeatedly outruns capacity | HBM/packaging remain binding | Premiums persist for high-memory products | Rising absolute and real burden |
| Geopolitical fragmentation | Capacity divided by controls and blocs | Regional allocation differs sharply | Import-dependent used markets tighten | Large geographic inequality |
| Dual market | Mainstream improves; frontier remains scarce | Premium memory concentrated in frontier systems | High-memory older systems retain value | Nominal access broadens, capability gap widens |
The dual-market scenario is the most important. Technological progress may reduce the cost of small and medium models while frontier capability becomes more expensive because it requires larger memory pools, more complex packaging, liquid cooling, advanced networking and higher power density. Under that scenario, statements that “AI hardware prices are falling” and “frontier AI is becoming unaffordable” can both be true.
2.19 Geographic Exposure, 2027–2031
Table 2.21 — Geographic Price-Transmission Mechanisms
| Market | Main strengths | Main exposure | Expected source of premium |
|---|---|---|---|
| United States | Large cloud market, direct vendor channels, capital access | Local scarcity, energy and trade-policy effects | Launch availability and state tax |
| European Union | Large market and public compute investment | VAT, energy price, limited frontier-hardware production | VAT, FX, regional allocation and energy |
| United Kingdom | Mature technology market | Sterling volatility and import dependence | FX, VAT and distribution |
| China | Large domestic technology base | Access restrictions to highest-end foreign accelerators | Product substitution and controlled performance |
| Japan | Strong electronics ecosystem | Yen volatility and imported accelerator dependence | FX and local official pricing |
| India | Large skills base and growing digital market | Income constraint, taxes and imported hardware | Import cost and affordability |
| Brazil | Large market | Currency volatility, taxes and import costs | Tax and distribution |
| Lower-income markets | Potential open-model adoption | Income, electricity, finance, thin distribution and warranty | Small volume, FX, credit and logistics |
For lower-income markets, a product can become cheaper in U.S. dollars while becoming less affordable locally. The affordability change is:
ΔAffordabilityr = ΔLocalHardwarePricer − ΔMedianDisposableIncomer
If local hardware prices rise 5% but disposable income rises only 2%, affordability deteriorates by approximately 3%, even if quality improves. The later country analysis will therefore report three outcomes simultaneously: dollars per capability unit, months of median income required and access to financing.
2.20 Final Assessment of H1, H2 and H3
H1: Prices of AI-relevant hardware have risen faster than general consumer-electronics prices
Status: not yet confirmed across the full market. The selected NVIDIA flagship MSRP sequence shows a nominal increase and a smaller real increase, but it does not establish a market-wide index. AMD reference products and Apple configuration changes demonstrate that product-level paths differ. A complete H1 determination requires the hedonic dataset.
H2: The inflation-adjusted cost of a fixed level of AI capability has increased
Status: unresolved and unlikely to have a universal answer. The selected NVIDIA sequence shows lower real prices per GB and per GB/s, but these are incomplete capability measures. Fixed-workload testing may show declining cost for smaller models and rising cost for privacy-sensitive, high-memory or frontier workloads.
H3: Used-market prices indicate persistent scarcity rather than temporary speculation
Status: unproven under available primary data. No representative, verified international completed-transaction dataset has been identified within the permitted evidence hierarchy. Claims of two-times or three-times MSRP cannot be generalised.
2.21 Chapter Conclusion
The verified evidence does not support a simplistic conclusion that every form of AI hardware became two or three times more expensive between 2020 and 2026. It supports a more consequential finding: the market is splitting between improving technical price-performance and increasing minimum access costs. The NVIDIA flagship launch price rose from $1,499 to $1,999, but memory capacity and bandwidth increased sufficiently to lower simple real price-per-capability ratios. The improvement does not eliminate exclusion because a user must still finance the entire device and supporting system. The result resembles a high-speed rail ticket becoming cheaper per kilometre while the minimum fare rises beyond the budget of some passengers.
Memory and data-centre hardware present a different problem: price transparency is substantially weaker. HBM demand, capacity commitments and investment are documented, but public transaction prices are not. Data-centre accelerators are commonly sold through negotiated systems and cloud arrangements, preventing a reliable public unit-price history. Used-market analysis is weaker still because asking prices are frequently mistaken for completed sales. Under an academic evidentiary standard, those gaps must remain visible.
The principal five-year risk is therefore a dual market. Mainstream AI capability may become cheaper through efficiency, open models and better consumer hardware, while frontier, private and high-concurrency capability becomes more expensive because of HBM, packaging, networking, electricity and institutional demand. Such a market would not exclude poorer users from AI entirely. It would confine them to a lower capability tier. That distinction—between access to some AI and access to economically competitive AI—is the central bridge from hardware inflation to Chapters 4, 7 and 8.
Chapter 3 — Industrial Concentration and the Architecture of Scarcity
3.1 Analytical Objective
The AI industrial system is not a single market governed by one concentration ratio. It is a chain of technically complementary and economically interdependent markets spanning semiconductor-design software, intellectual property, lithography, fabrication, memory, packaging, substrates, accelerator design, networking, server assembly, cloud infrastructure, foundation models and enterprise distribution. Concentration at one layer can be offset by substitution or intensified by dependence on another. A company may possess strong accelerator-design capabilities while relying on an external foundry, HBM suppliers, packaging capacity and cloud customers. A foundry may control leading-edge production while remaining exposed to a small number of customers and equipment suppliers. A cloud provider can operate enormous infrastructure but depend on accelerators designed and fabricated elsewhere. The appropriate analytical object is therefore not “the AI monopoly,” but a network of differentiated bottlenecks whose control, scarcity, substitution time and contractual allocation jointly determine effective AI capacity.
This chapter makes five distinctions. First, market concentration measures how economic activity is distributed among firms. Second, technical concentration measures how many suppliers can satisfy a particular performance requirement. Third, dependency concentration measures whether downstream firms rely on the same upstream node. Fourth, allocation concentration measures whether future output is reserved by a small number of purchasers. Fifth, economic rent measures returns exceeding the amount required to retain capital and capability in their current use. None of these concepts independently establishes unlawful conduct. A firm can earn high margins because it created a superior product, assumed substantial technological risk, invested before demand materialised or temporarily controls scarce capacity. Conversely, a market can exhibit exclusionary effects even when accounting margins are not extraordinary, particularly where interoperability restrictions, licensing terms or long-term capacity commitments raise rivals’ costs.
The European Commission identifies data, AI accelerator chips, computing infrastructure, cloud capacity and technical expertise as potentially critical inputs and possible barriers to entry in generative-AI markets. Its analysis explicitly treats bottlenecks as context-dependent: a constrained input can reduce competition without every constraint constituting an anticompetitive practice. Competition Policy Brief: Competition in Generative AI and Virtual Worlds – European Commission – September 2024
3.2 Concentration Metrics and Interpretation
The n-firm concentration ratio is:
CRn = Σi=1nsi
Where si is firm i’s market share. CR1 measures the largest firm; CR3 and CR4 measure the combined share of the three or four largest firms. The Herfindahl–Hirschman Index is:
HHI = Σi=1N(100si)2
When market shares are expressed as decimal fractions. If shares are already expressed as percentage points, the equivalent formula is:
HHI = Σi=1Nsi2
The HHI ranges from close to zero in an extremely fragmented market to 10,000 in a pure single-firm market. Five equally sized firms generate an HHI of 2,000; four equal firms generate 2,500; three equal firms generate approximately 3,333; two equal firms generate 5,000.
Under the current U.S. Merger Guidelines:
- HHI above 1,800 indicates a highly concentrated market;
- an increase exceeding 100 points is considered significant;
- a transaction producing both conditions can generate a structural presumption that competition may be substantially lessened.
These are merger-screening thresholds, not a legal declaration that every market above 1,800 is unlawful. Market definition, entry, innovation, buyer power, substitution and efficiencies remain essential. 2023 Merger Guidelines – U.S. Department of Justice and Federal Trade Commission – December 2023
Table 3.1 — Mathematical Meaning of HHI
| Market structure | Equal-share assumption | HHI |
|---|---|---|
| One supplier | 100% | 10,000 |
| Two suppliers | 50% each | 5,000 |
| Three suppliers | 33.33% each | 3,333 |
| Four suppliers | 25% each | 2,500 |
| Five suppliers | 20% each | 2,000 |
| Six suppliers | 16.67% each | 1,667 |
| Ten suppliers | 10% each | 1,000 |
| Twenty suppliers | 5% each | 500 |
Equal shares produce the minimum HHI for a fixed number of firms. If one supplier is larger, HHI rises. A nominal three-firm market therefore cannot have an HHI below approximately 3,333 unless additional suppliers are included in the correctly defined market.
Market shares can understate AI concentration for four reasons. First, revenue-based shares do not necessarily measure available frontier capacity. A supplier can have a modest share of total semiconductor revenue but control a much larger share of a technically indispensable advanced segment. Second, nominal competitors may not be substitutes. An accelerator that cannot run the required framework, fit the model or satisfy export restrictions does not constrain price effectively. Third, the same upstream node may support several nominal downstream competitors. Multiple cloud providers can offer different services while relying on the same foundry, lithography equipment or memory suppliers. Fourth, long-term contracts can allocate most future output before it enters any observable spot market.
The report therefore supplements CRn and HHI with four indices:
Supplier Dependency Index:
SDIj = Σi=1Nwidi,j
Where di,j is firm i’s dependence on supplier j and wi is firm i’s downstream importance.
Effective Substitutability Index:
ESIj = 1 ÷ [1 + Tswitch,j + Cswitch,j + Lperformance,j]
Where T is switching time, C is normalised switching cost and L is performance loss.
Reserved Capacity Ratio:
RCRj,t = Capacitycontracted,j,t ÷ Capacitytotal,j,t
Bottleneck-Adjusted Concentration:
BACj = HHIj × Criticalityj × [1 − ESIj]
A large HHI in a non-critical, easily substituted component may have limited systemic effect. A somewhat lower HHI in an indispensable input with multi-year qualification requirements can be more consequential.
3.4 Full AI Industrial-Control Map
Table 3.2 — Principal Firms by AI Supply-Chain Layer
| Layer | Principal firms or platforms | Function | Central dependency |
|---|---|---|---|
| Semiconductor EDA | Synopsys, Cadence, Siemens EDA; specialised vendors | Design, verification, physical implementation and simulation | Proprietary formats, process-design kits and engineering workflows |
| Semiconductor IP | Arm, Synopsys, Cadence, Imagination and specialised IP providers | CPU, interface, memory and connectivity IP | Licensing, standards and integration |
| EUV lithography | ASML | Patterning of leading-edge semiconductor layers | Extremely complex equipment and supplier network |
| DUV lithography | ASML, Nikon, Canon | Mature and selected advanced process steps | Tool capability, throughput and process qualification |
| Deposition and etch | Applied Materials, Lam Research, Tokyo Electron and specialised firms | Material deposition and pattern transfer | Process recipes, service network and qualification |
| Metrology and inspection | KLA, ASML and specialised firms | Defect detection, yield control and process measurement | Yield learning and process control |
| Leading-edge foundry | TSMC, Samsung Foundry, Intel Foundry; selected Chinese domestic development | Advanced logic fabrication | Process maturity, yield, tool access and customer qualification |
| Mature-node foundry | TSMC, GlobalFoundries, UMC, SMIC, Hua Hong, Tower and others | Supporting logic, control and interface chips | Geographic capacity and application qualification |
| HBM | SK hynix, Samsung Electronics, Micron | High-bandwidth accelerator memory | Yield, stack technology, packaging integration and qualification |
| Conventional DRAM | Samsung, SK hynix, Micron; additional regional suppliers | System and server memory | Cyclical capacity allocation |
| NAND | Samsung, SK hynix/Solidigm, Kioxia, Western Digital/SanDisk, Micron and others | Persistent model and dataset storage | Layer transitions, controller integration and cyclicality |
| Advanced packaging | TSMC, ASE, Amkor, Samsung, Intel and other OSAT providers | Logic-memory integration and chiplet packaging | CoWoS-equivalent capacity, substrates and thermal design |
| Substrates | Ibiden, Shinko, Unimicron, Nan Ya PCB, AT&S and others | Physical package interconnection | Qualification, yield and expansion lead time |
| Silicon interposers | Foundries and packaging specialists | High-density connection between logic and HBM | Large-die processing and packaging integration |
| GPU design | NVIDIA, AMD, Intel and regional challengers | General accelerated computing | Architecture, software and foundry access |
| Custom AI accelerators | Google TPU, AWS Trainium/Inferentia, Microsoft Maia, Meta silicon programmes and others | Workload-specific cloud acceleration | Internal cloud scale and software integration |
| Networking | NVIDIA, Broadcom, Marvell, Cisco, Arista and others | Scale-up and scale-out AI interconnect | Bandwidth, latency, optics and switching |
| AI servers | Supermicro, Dell, HPE, Lenovo, Inspur and ODMs including Quanta, Wiwynn and Foxconn | Integration of accelerators, CPUs, memory and cooling | GPU allocation, rack power and liquid cooling |
| Public cloud | AWS, Microsoft Azure, Google Cloud, Oracle and others | Metered compute and AI infrastructure | Capital expenditure, installed capacity and customer ecosystem |
| Frontier models | OpenAI, Google DeepMind, Anthropic, Meta, xAI and other regional developers | Foundation-model creation and inference | Compute, data, talent, capital and distribution |
| Enterprise AI distribution | Microsoft, Google, AWS, Salesforce, Oracle, SAP, ServiceNow, Adobe and specialist firms | Embedding AI into enterprise workflows | Installed software base, customer data and switching cost |
The table identifies principal actors, not formal market shares. In most layers, audited corporate reports do not publish a complete common-denominator market dataset. Consequently, the presence of named firms must not be mistaken for a calculated CR4.
3.5 Lithography: The Highest-Order Equipment Bottleneck
Lithography is the most structurally concentrated upstream layer because leading-edge fabrication depends on exceptionally complex equipment with long development cycles, extensive intellectual property and thousands of specialised suppliers. ASML reported €32.7 billion in 2025 net sales, a 52.8% gross margin, €11.3 billion operating income, €9.6 billion net income, €4.7 billion R&D expenditure and a supplier network of approximately 5,100 companies. It sold 48 EUV systems, 279 DUV systems and 208 metrology and inspection systems during the year. ASML 2025 Annual Report – ASML – January 2026
The verified primary dataset identifies ASML as the supplier reporting commercial EUV-system sales, but it does not provide an audited global market-share denominator constructed under a competition-law market definition. A formal CR1 and HHI are therefore not stated as observed statistics. The conditional mathematical result is nevertheless clear: if the relevant market is defined as commercial EUV lithography systems and ASML supplies 100%, then:
CR1 = 100%
HHI = 1002 = 10,000
That would be a technical single-source market. It would not mean ASML can act without constraint. Its power is limited by semiconductor-cycle demand, customer capital budgets, export licensing, supplier capacity, technological execution and the risk that customers delay node transitions. Its exceptionally high R&D expenditure and supplier dependence also show why accounting margin cannot be classified entirely as monopoly rent.
Table 3.3 — Lithography Entry Barriers
| Barrier | Severity | Explanation |
|---|---|---|
| Optical and light-source complexity | Extreme | Leading-edge lithography requires capabilities accumulated over decades |
| Precision manufacturing | Extreme | System performance depends on sub-nanometre control and exceptionally complex integration |
| Supplier network | Extreme | Thousands of specialised suppliers contribute irreplaceable subsystems |
| Customer qualification | Extreme | Foundries must integrate tools into complete process flows |
| Intellectual property | Extreme | Patents, trade secrets and cumulative engineering knowledge |
| Capital requirement | Extreme | Development and production require sustained multi-billion investment |
| Service infrastructure | Very high | Installed tools require continuous support and calibration |
| Export licensing | High | Government controls can restrict addressable markets |
| Time to credible entry | Very long | Entry would require technological, manufacturing and customer-validation cycles |
3.6 Electronic-Design Automation and Semiconductor IP
EDA tools sit upstream of every advanced chip. They support architecture design, verification, timing, physical implementation, power analysis and sign-off. Their economic importance exceeds their direct share of semiconductor expenditure because an advanced design cannot be fabricated reliably without validated tools and foundry-compatible process-design kits. Switching costs include retraining engineers, rebuilding design flows, validating scripts, converting libraries, repeating verification and accepting schedule risk. These costs create ecosystem dependence even where alternative tools exist.
The European Commission’s review of Synopsys’ acquisition of Ansys identified competition concerns in three global software markets and required divestitures covering the relevant overlaps. The Commission described Synopsys as an EDA and semiconductor-IP supplier and Ansys as a simulation and analysis provider whose semiconductor tools overlapped with parts of Synopsys’ portfolio. Buying Ansys: A Synopsys of a Chip Design Story – European Commission – 2025
No complete primary-source revenue shares for the entire EDA market were identified under the permitted source hierarchy. Therefore:
- CR3: ND;
- CR4: ND;
- HHI: ND.
A three-firm illustration must not be misreported as an estimate. If a correctly defined EDA submarket contained exactly three equal firms, its minimum HHI would be approximately 3,333. Actual concentration depends on tool category, specialised competitors, internal tools and the relevant geographic market.
Table 3.4 — EDA Lock-In Mechanisms
| Mechanism | Economic effect | Observable test |
|---|---|---|
| Foundry-certified design flow | Limits tools that can reach production sign-off | Number of certified alternatives by process node |
| Proprietary file and scripting formats | Raises migration cost | Conversion time and defect rate |
| Engineer skill concentration | Makes alternative adoption expensive | Labour-market skill distribution |
| Integrated tool suite | Creates scope efficiencies | Separate versus integrated workflow cost |
| Semiconductor IP bundle | Reinforces tool adoption | Share of designs using bundled IP |
| Historical design libraries | Embeds prior investment | Cost of library requalification |
| Long design cycles | Makes mid-project switching impractical | Switching frequency by project stage |
| Sign-off liability | Favours incumbent validated tools | Customer acceptance and insurance requirements |
3.7 Leading-Edge Foundry Capacity
Leading-edge foundry concentration results from capital intensity, process complexity, yield learning, tool access and customer qualification. TSMC, Samsung and Intel represent the principal global organisations pursuing the most advanced logic processes, although their commercially available capacity, yields and customer models differ. Mature-node capacity is more diversified and must not be included in a leading-edge denominator merely because it fabricates semiconductors.
TSMC reported 2025 revenue of approximately $122 billion, up 35.9% in U.S.-dollar terms. Its 2025 gross margin was 59.9%, operating margin 50.8% and net margin 45.1%. The company attributed the margin increase partly to higher capacity utilisation and cost improvement, while overseas fabrication expansion diluted profitability. TSMC 2025 Annual Report – Taiwan Semiconductor Manufacturing Company – April 2026 In the second quarter of 2026, process technologies at 7 nanometres and below accounted for 77% of TSMC wafer revenue, while HPC represented 66% of quarterly revenue. These are TSMC internal revenue shares, not global foundry shares. TSMC Second Quarter 2026 Earnings Conference Transcript – Taiwan Semiconductor Manufacturing Company – July 2026
Primary audited materials available in this research session do not provide a harmonised 2025 global leading-edge foundry share series. Consequently:
- CR1: ND;
- CR3: ND;
- HHI: ND.
Table 3.5 — Foundry Barriers
| Barrier | Leading-edge effect | Substitution period |
|---|---|---|
| Fabrication-plant capital | Restricts entry and simultaneous expansion | Multi-year |
| Yield learning | Makes nominal process availability different from economic capacity | Multi-quarter to multi-year |
| Customer qualification | Prevents instant design migration | Often multiple quarters |
| Process-design kits | Ties design tools and libraries to a node | Project-specific |
| Advanced equipment | Depends on lithography, deposition, etch and metrology suppliers | Multi-year |
| Skilled workforce | Limits rapid geographic replication | Multi-year |
| Utility requirement | Large, stable electricity and ultrapure-water demand | Location-specific |
| Export controls | Restricts equipment and customer access | Policy-dependent |
| Packaging integration | Logic capacity is insufficient without downstream packaging | Multi-quarter |
| Geographic risk | Concentrated sites create correlated disruption | Potentially immediate impact |
3.8 Advanced Packaging: The Hidden Capacity Multiplier
Advanced packaging transforms fabricated logic and HBM into a usable accelerator. For AI systems, packaging is no longer a low-value final step; it determines how many compute dies and HBM stacks can communicate at sufficient bandwidth and power density. TSMC’s CoWoS platform, Intel’s EMIB and Foveros technologies, Samsung’s advanced packaging, and services supplied by ASE, Amkor and other OSAT firms form a heterogeneous market. These offerings cannot automatically be included in one market because technical substitutability varies by design.
TSMC stated in April 2025 that it was working to double CoWoS capacity during that year. The company also announced plans for a 9.5-reticle-size CoWoS implementation in volume production during 2027, intended to support twelve or more HBM stacks. TSMC First Quarter 2025 Earnings Conference Transcript – Taiwan Semiconductor Manufacturing Company – April 2025 TSMC Unveils Next-Generation A14 Process at North America Technology Symposium – Taiwan Semiconductor Manufacturing Company – April 2025
No primary, audited capacity-share denominator permits a defensible packaging HHI. The appropriate concentration measure must use qualified AI-package output rather than total packaging revenue.
Table 3.6 — Packaging Capacity Measurement
| Variable | Correct unit | Why revenue share is insufficient |
|---|---|---|
| Interposer capacity | Qualified interposer area per month | Package sizes differ |
| AI packages completed | Qualified packages per quarter | Complexity and HBM count vary |
| HBM stacks integrated | Stacks per package and total stacks | Stack intensity changes capacity use |
| Yield | Good packages divided by starts | Nominal capacity overstates output |
| Lead time | Order-to-delivery months | Reveals effective scarcity |
| Customer allocation | Share committed under contract | Shows residual open capacity |
| Thermal capability | Supported power density | Determines substitutability |
| Reticle-equivalent area | Reticle units per package | Normalises package scale |
3.9 HBM and Conventional Memory
HBM is a stronger potential bottleneck than conventional DRAM because AI accelerators require high bandwidth, dense stacking and close integration with advanced packaging. Principal HBM suppliers include SK hynix, Samsung Electronics and Micron. Conventional DRAM is also concentrated around these large producers, although exact shares differ by product generation and period. NAND includes a broader set of firms.
Micron disclosed that its HBM business was supported by multi-year strategic customer agreements and forecast the HBM total addressable market to increase from approximately $35 billion in 2025 to around $100 billion in 2028. Micron Fiscal First Quarter 2026 Financial Results – Micron Technology – December 2025 Samsung announced plans to invest more than KRW 110 trillion in facilities and R&D during 2026, including HBM, foundry operations and advanced packaging. Samsung Electronics Shareholder Value Enhancement Plan – Samsung Electronics – March 2026
Official corporate disclosures do not provide a common audited global share denominator for HBM. Therefore, CR3 and HHI remain ND. If exactly three suppliers constituted the complete relevant market, the mathematical minimum HHI would be approximately 3,333; unequal shares would produce a higher figure. This is a structural illustration, not a reported market estimate.
Table 3.7 — HBM Scarcity Mechanisms
| Mechanism | Effect |
|---|---|
| Wafer-intensity per delivered bit | HBM capacity competes with conventional memory capacity |
| Complex stacking | Yield losses can compound across layers |
| Customer qualification | Accelerator suppliers cannot instantly change memory |
| Packaging co-design | Memory must be integrated with logic and interposer design |
| Long-term agreements | Future supply can be allocated before spot availability |
| Generation transition | HBM3E and HBM4 cannot be treated as identical output |
| Capital discipline | Memory producers may avoid uncontrolled commodity overcapacity |
| Export controls | Eligible demand differs by jurisdiction |
| Thermal and power constraints | Nominal capacity may not be deployable in all systems |
The U.S. Bureau of Industry and Security added controls on HBM in December 2024 alongside new controls on semiconductor-manufacturing equipment and software. This policy confirms that HBM is treated as a strategic input, but it does not quantify commercial concentration. Commerce Strengthens Export Controls to Restrict China’s Capability to Produce Advanced Semiconductors for Military Applications – U.S. Bureau of Industry and Security – December 2024
3.10 Substrates, Interposers and Secondary Bottlenecks
Package substrates and interposers can constrain AI capacity even though their revenue is small relative to accelerators. They require high layer counts, fine interconnect geometry, low defect rates and qualification for large, high-power packages. A substrate shortage can immobilise completed logic and memory, creating a disproportionate output loss. This is the “low-value/high-criticality paradox”: a component can represent a small share of system cost while controlling the availability of the entire system.
Criticality-adjusted value is:
CAVj = RevenueSharej × OutputLossElasticityj
A low-revenue input with a very large output-loss elasticity can rank above a more expensive but substitutable component.
No audited global market shares for AI-qualified substrates or silicon interposers were identified. CR4 and HHI are therefore ND. Chapter monitoring should instead use:
- substrate lead time;
- qualified supplier count;
- capacity-expansion announcements;
- yield;
- customer concentration;
- package-area requirements;
- geographic concentration;
- inventory coverage.
3.11 Accelerator Design and CUDA Ecosystem Dependence
Accelerator competition occurs across discrete GPUs, integrated accelerators, custom ASICs and internally developed cloud chips. Principal suppliers include NVIDIA, AMD and Intel, while Google, Amazon and Microsoft have developed custom accelerators for internal cloud deployment. Chinese and other regional firms pursue domestic alternatives under different software and export-control conditions.
NVIDIA’s competitive position cannot be reduced to silicon. Its system includes GPU architecture, CUDA, libraries, compilers, communications software, NVLink, InfiniBand and Ethernet networking, reference systems and developer support. The economic switching cost includes code porting, kernel optimisation, model validation, workforce retraining, operational risk and performance uncertainty. A customer can theoretically purchase another accelerator but still face a large total migration cost.
The CUDA Lock-In Index proposed for later estimation is:
CLIi = w₁CodePortingi + w₂PerformanceLossi + w₃Retrainingi + w₄LibraryGapi + w₅OperationalRiski
Table 3.8 — Accelerator Competition Layers
| Layer | NVIDIA position | Alternative mechanism | Lock-in test |
|---|---|---|---|
| Hardware | GPU and integrated systems | AMD, Intel, custom ASICs and regional accelerators | Constant-workload cost |
| Programming model | CUDA | ROCm, SYCL, Open standards and proprietary cloud stacks | Porting time |
| Libraries | CUDA-X ecosystem | Alternative libraries and framework abstractions | Feature completeness |
| Scale-up interconnect | NVLink/NVSwitch | Alternative proprietary or open interconnects | Cluster efficiency |
| Scale-out networking | InfiniBand and Ethernet portfolio | Broadcom, Marvell, Cisco, Arista and others | End-to-end performance |
| Systems | DGX, HGX and rack-scale platforms | OEM and custom systems | Deployment time |
| Cloud distribution | Instances across hyperscalers | Custom cloud chips and rival GPUs | Availability and tariff |
| Developer base | Large installed skill base | Cross-platform software and compiler abstraction | Labour-market switching cost |
No official authority or audited corporate filing supplies a complete 2025 accelerator market denominator that simultaneously covers merchant GPUs and captive custom chips. A formal global accelerator HHI is therefore ND. Treating custom TPU or Trainium capacity as zero because it is not sold as a discrete card would overstate merchant-market concentration while understating user dependence on individual clouds.
3.12 AI-Server Manufacturing
AI-server manufacturing includes branded OEMs, original-design manufacturers and cloud operators that design systems internally. Supermicro, Dell, HPE, Lenovo and Inspur coexist with manufacturing partners such as Quanta, Wiwynn and Foxconn. Their bargaining power depends less on processor intellectual property than on allocation, integration, liquid cooling, rack power, delivery capability and customer certification.
AI-server concentration cannot be calculated reliably from total server revenue because conventional enterprise servers are not substitutes for dense accelerator systems. A valid market denominator must include only systems meeting defined accelerator, networking and cooling criteria.
Table 3.9 — AI-Server Entry Barriers
| Barrier | Severity | Explanation |
|---|---|---|
| Accelerator allocation | Very high | Integrators cannot ship systems without scarce GPUs |
| Thermal engineering | High | Rack-scale AI systems require advanced cooling |
| Power density | High | Facility compatibility limits deployment |
| Networking integration | High | Performance depends on topology and configuration |
| Firmware and validation | High | Enterprise customers require tested configurations |
| Working capital | High | Expensive components create financing exposure |
| Customer support | High | Downtime costs require global servicing |
| Manufacturing scale | Medium-high | Volume and supplier relationships affect delivery |
| Brand trust | Medium-high | Critical buyers favour established support providers |
3.13 Cloud Infrastructure: A Verifiable Concentration Case
Cloud infrastructure is one segment where an official authority has published usable share ranges. The UK Competition and Markets Authority’s 2024 revenue analysis placed AWS and Microsoft individually in the 30–40% range for combined IaaS and PaaS, Google in the 5–10% range, IBM and Oracle each in the 0–5% range, and other providers collectively in the 20–30% range. The CMA stated that AWS and Microsoft’s combined share remained within 60–80% in 2024. Appendix D: Market Structure, Concentration Methodology and UK Share of Supply by Revenue – UK Competition and Markets Authority – 2025
Because exact shares are redacted into ranges and “other” aggregates many firms, a single exact HHI would be false precision. A strict concentration floor can nevertheless be calculated.
If AWS and Microsoft have a combined 60% and are divided equally, and Google has 5%, then:
HHI floor = 302 + 302 + 52 = 1,825
The remaining 35% necessarily adds a positive amount. Therefore actual HHI is greater than 1,825.
If AWS and Microsoft have a combined 80%, divided equally, and Google has 10%, then:
Partial HHI = 402 + 402 + 102 = 3,300
Again, remaining firms add to the value. The result establishes that the UK IaaS/PaaS market is highly concentrated under the U.S. 1,800 threshold even at the lowest mathematically favourable endpoint of the CMA ranges.
Table 3.10 — UK Cloud Concentration
| Measure | Official range or derived value | Interpretation |
|---|---|---|
| AWS share, 2024 | 30–40% | One of two market leaders |
| Microsoft share, 2024 | 30–40% | One of two market leaders |
| Google share, 2024 | 5–10% | Distant third |
| IBM share, 2024 | 0–5% | Small |
| Oracle share, 2024 | 0–5% | Small |
| Other combined | 20–30% | Aggregate, not one firm |
| CR2 | 60–80% | Very high two-firm concentration |
| CR3 | 65–90% before range reconciliation | High; exact value redacted |
| HHI lower bound | Greater than 1,825 | Highly concentrated |
| Plausible partial HHI at upper endpoints | At least 3,300 | Very highly concentrated |
The CMA subsequently stated that its investigation found AWS and Microsoft to possess positions of significant market power and identified data-egress charges, interoperability barriers and Microsoft software-licensing practices as limitations on customer choice. CMA Announces Package of Actions on Business Software and Cloud Services – UK Competition and Markets Authority – March 2026
3.14 Frontier-Model Development
Frontier models require capital, compute, data, specialised talent, evaluation systems, distribution and sustained inference capacity. Principal developers include OpenAI, Google DeepMind, Anthropic, Meta and xAI, alongside major Chinese and other national developers. Market definition remains unresolved. Possible markets include model training, API access, consumer assistants, enterprise models, open-weight models, multimodal systems and task-specific inference. User counts, API revenue, token volume and benchmark capability would produce different shares.
No primary regulator or audited filing provides a complete frontier-model revenue or quality-adjusted usage denominator. Therefore:
- CR3: ND;
- CR4: ND;
- HHI: ND.
The absence of a calculable HHI does not imply low concentration. It means the relevant market and denominator remain empirically incomplete.
Table 3.11 — Frontier-Model Entry Barriers
| Barrier | Mechanism |
|---|---|
| Training compute | Large upfront cluster access |
| Inference capacity | Continuing expenditure after model launch |
| Data | Quality, legality, cleaning and domain coverage |
| Talent | Small global pool of experienced researchers and infrastructure engineers |
| Evaluation | Safety, reliability and benchmark infrastructure |
| Distribution | Consumer platform, cloud or enterprise installed base |
| Brand and trust | Buyers prefer recognised models for critical deployment |
| Regulation | Compliance fixed costs favour scale |
| Feedback loops | User interactions can improve products and distribution |
| Capital | Long investment horizon and uncertain monetisation |
3.15 Vertical Integration Between Clouds and Model Developers
Vertical and quasi-vertical relationships join infrastructure providers with model developers through equity investment, cloud commitments, revenue sharing, distribution, technical cooperation and preferential access. The FTC’s January 2025 study examined partnerships involving Microsoft and OpenAI, Amazon and Anthropic, and Alphabet and Anthropic. It found that cloud partners received various equity, revenue-sharing, consultation, control or exclusivity-related rights and that model developers obtained capital, cloud resources and distribution. Partnerships Between Cloud Service Providers and AI Developers – Federal Trade Commission – January 2025
These partnerships can create efficiencies:
- finance otherwise unaffordable model development;
- secure infrastructure supply;
- integrate models into enterprise products;
- reduce deployment latency;
- distribute model services to a large customer base;
- align technical road maps.
They can also create competitive risks:
- make model developers dependent on one cloud;
- reserve scarce compute;
- provide cloud partners access to sensitive technical information;
- limit multi-cloud deployment;
- favour affiliated models in distribution;
- increase switching cost;
- create circular revenue and procurement relationships.
Table 3.12 — Vertical Integration Test
| Question | Efficiency interpretation | Competition-risk interpretation |
|---|---|---|
| Is cloud capacity guaranteed? | Enables investment and stable deployment | Removes scarce capacity from rivals |
| Is model distribution bundled? | Reduces customer acquisition cost | Favours affiliated model |
| Are commitments exclusive? | Supports relationship-specific investment | Prevents multi-cloud competition |
| Is sensitive information shared? | Improves engineering coordination | Weakens future rivalry |
| Are revenues shared? | Aligns incentives | Reinforces circular dependence |
| Can the developer switch cloud? | Relationship remains contestable | Lock-in becomes structural |
| Can customers export workloads? | Interoperability preserved | Data and workflow lock-in |
| Are rival models offered equally? | Platform remains open | Self-preferencing |
3.16 Long-Term Supply Agreements and Foundry Allocation
Long-term agreements can stabilise investment in fabrication, HBM, packaging and cloud infrastructure. Capital-intensive suppliers may not expand without credible future demand; purchasers may not invest in model programmes without capacity assurance. Agreements are therefore economically productive when they share risk and finance expansion.
The competitive concern arises when a small number of purchasers reserve a dominant share of output, leaving independent firms to buy from a thin residual market. The appropriate metrics are:
Customer Allocation Ratio:
CARk,j = Capacityallocated to customer k,j ÷ Capacitytotal,j
Residual Market Ratio:
RMRj = 1 − Σk=1KCARk,j
Allocation HHI:
AHHIj = Σk=1K(100CARk,j)2
NVIDIA disclosed that in fiscal 2026 one direct customer represented 22% of total revenue and another represented 14%, primarily in Compute & Networking. This is customer concentration, not supplier market share. It shows substantial buyer concentration within NVIDIA’s revenue base and complicates a simple seller-power narrative. NVIDIA Annual Report on Form 10-K for the Fiscal Year Ended January 25, 2026 – NVIDIA and U.S. Securities and Exchange Commission – February 2026
3.17 Export Controls as an Architecture of Geographic Scarcity
Export controls divide global physical capacity from legally accessible capacity. The U.S. Bureau of Industry and Security controls designated advanced computing semiconductors, semiconductor-manufacturing equipment, HBM and related software according to performance, end user and destination. In December 2024, BIS added controls on 24 types of semiconductor-manufacturing equipment, three types of software and specified HBM. Commerce Strengthens Export Controls to Restrict China’s Capability to Produce Advanced Semiconductors for Military Applications – U.S. Bureau of Industry and Security – December 2024
In January 2026, BIS revised its licensing approach for NVIDIA H200, AMD MI325X and similar products destined for approved Chinese customers, moving to case-by-case review subject to stated conditions. Department of Commerce Revises License Review Policy for Semiconductors Exported to China – U.S. Bureau of Industry and Security – January 2026
The geopolitical availability ratio is:
GARr,t = LegallyAccessibleFrontierCapacityr,t ÷ GlobalFrontierCapacityt
A country can face severe effective scarcity even while global production rises if GAR declines.
Table 3.13 — Export-Control Effects
| Effect | Intended function | Economic consequence |
|---|---|---|
| Restrict advanced accelerator exports | National-security control | Regional performance ceiling |
| Restrict manufacturing equipment | Limit indigenous advanced production | Prolonged fabrication dependence |
| Control HBM | Restrict complete AI-system performance | Memory-based capacity constraint |
| Control design/manufacturing software | Limit advanced-node productivity | EDA dependence |
| Entity-list restrictions | Target specific organisations | Due-diligence and transaction costs |
| Anti-diversion rules | Prevent circumvention | Wider compliance burden |
| Licensing uncertainty | Preserve government discretion | Investment and allocation uncertainty |
| Domestic substitution | Encourage alternative ecosystems | Duplication cost and fragmentation |
3.18 Subsidies and Industrial Policy
Governments increasingly treat semiconductor and AI capacity as strategic infrastructure. Subsidies can reduce entry barriers, accelerate domestic fabrication, diversify geography and internalise resilience benefits not rewarded by short-term commercial returns. They can also create overcapacity, political allocation, duplicated infrastructure or subsidy competition among jurisdictions.
The economic test is:
NetSocialValuep = ResilienceBenefitp + InnovationSpilloverp + SecurityBenefitp + TaxEmploymentBenefitp − FiscalCostp − DistortionCostp − DuplicationCostp
Table 3.14 — Industrial-Policy Evaluation
| Policy | Potential benefit | Principal risk | Required KPI |
|---|---|---|---|
| Fab subsidy | Geographic diversification | Subsidy competition | Qualified output per public dollar |
| R&D credit | Technology development | Subsidising activity that would occur anyway | Incremental patent and process output |
| Public compute | Research and SME access | Low utilisation or political allocation | Successful workloads and user diversity |
| AI Gigafactory | Frontier-scale capacity | Reinforcing large incumbents | Open-access share and cost per task |
| Workforce programme | Reduces skill bottleneck | Training mismatch | Qualified workers retained |
| Energy investment | Enables data-centre expansion | Cost socialisation | Firm capacity and grid reliability |
| Procurement preference | Creates demand for domestic suppliers | Higher cost and lock-in | Performance-adjusted procurement cost |
| Open standards | Reduces switching costs | Slow consensus | Migration cost and interoperability |
3.19 Firm-Level Profitability and Economic Rent
High profitability is evidence of value capture, not proof of unlawful market power. This chapter calculates reported margin indicators for selected critical firms using official filings.
Operating margin is:
OMi,t = OperatingIncomei,t ÷ Revenuei,t × 100
Net margin is:
NMi,t = NetIncomei,t ÷ Revenuei,t × 100
Return on invested capital is:
ROICi,t = NOPATi,t ÷ AverageInvestedCapitali,t
Economic profit is:
EPi,t = [ROICi,t − WACCi,t] × InvestedCapitali,t
Free cash flow is:
FCFi,t = OperatingCashFlowi,t − CapitalExpenditurei,t
Market value added is:
MVAi,t = EnterpriseMarketValuei,t − InvestedCapitali,t
Tobin’s q is approximated as:
qi,t = MarketValueAssetsi,t ÷ ReplacementCostAssetsi,t
A credible Tobin’s q cannot be calculated from market capitalisation alone because the replacement cost of intangible and physical assets is not directly observable.
Table 3.15 — Verified Profitability Indicators
| Firm and period | Revenue | Gross margin | Operating income or margin | Net income or margin | Evidentiary interpretation |
|---|---|---|---|---|---|
| NVIDIA FY2026 | $215.938bn | 71.1% | $130.387bn; ≈60.4% | $120.067bn; ≈55.6% | Exceptional value capture across accelerated-computing platform |
| TSMC 2025 | ≈$122bn | 59.9% | 50.8% margin | 45.1% margin | Strong utilisation, node leadership and scale |
| ASML 2025 | €32.7bn | 52.8% | €11.3bn; ≈34.6% | €9.6bn; ≈29.4% | High-return equipment and service ecosystem |
| TSMC 2024 | Reported annual revenue | 56.1% | 45.7% | 40.5% | Lower than 2025, supporting utilisation and mix effects |
| TSMC 2023 | $69.30bn | 54.4% | 42.6% | 38.8% | Semiconductor downturn demonstrates cyclicality |
NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026 – NVIDIA – February 2026 TSMC 2025 Annual Report – Taiwan Semiconductor Manufacturing Company – April 2026 ASML 2025 Annual Report – ASML – January 2026 TSMC 2023 Annual Report – Taiwan Semiconductor Manufacturing Company – February 2024
Table 3.16 — Profit Explanation Matrix
| Observed profit source | NVIDIA | TSMC | ASML |
|---|---|---|---|
| Innovation rent | Very high | Very high | Very high |
| Intellectual-property rent | Very high | High through process know-how | Very high |
| Scarcity rent | High during constrained supply | High at leading nodes | High for critical tools |
| Scale economy | Very high | Very high | High |
| Ecosystem rent | Very high | Medium-high | High through installed base |
| Risk compensation | High | Very high due capital cycle | Very high due R&D cycle |
| Market-power component | Plausible and requiring formal test | Plausible in leading-edge segments | Structurally plausible |
| Evidence of unlawful conduct | Not established by margins | Not established by margins | Not established by margins |
NVIDIA’s fiscal 2026 operating margin of approximately 60.4% is extraordinarily high, but the gross margin declined from 75.0% in FY2025 to 71.1% in FY2026, partly reflecting the transition to full-scale Blackwell data-centre solutions and a charge associated with H20 inventory and purchase obligations. This movement is inconsistent with a simplistic assumption that market power mechanically forces margins upward every year. It supports a mixed explanation involving innovation, system mix, scarcity and platform power.
3.20 Formal Bottleneck Analysis
A bottleneck must be evaluated by more than supplier count. This chapter defines the Global AI Bottleneck Score:
GABSj = 0.25Cj + 0.20Sj + 0.20Tj + 0.15Gj + 0.10Ij + 0.10Aj
Where each variable is scored from 0 to 100:
- Cj: concentration;
- Sj: lack of substitutability;
- Tj: replacement or expansion time;
- Gj: geographic concentration;
- Ij: impact on downstream output;
- Aj: degree of pre-allocation.
The scores below are structured analytical assessments, not audited statistics.
Table 3.17 — Global AI Bottleneck Ranking
| Rank | Component | Concentration | Substitutability | Expansion time | Disruption impact | Preliminary risk |
|---|---|---|---|---|---|---|
| 1 | EUV lithography | Extreme | Extremely low | Extremely long | Systemic | Critical |
| 2 | Leading-edge foundry capacity | Very high | Very low | Very long | Systemic | Critical |
| 3 | HBM | Very high | Low | Long | Severe | Critical |
| 4 | Advanced packaging | High | Low | Long | Severe | Critical |
| 5 | Accelerator software ecosystem | Very high | Low-medium | Long organisationally | Severe | Critical |
| 6 | EDA and process-design integration | Very high | Low | Long | Systemic for new designs | Critical |
| 7 | AI data-centre electricity | Regionally high | Low locally | Long | Severe by region | High |
| 8 | High-speed networking and optics | High in specialised products | Medium-low | Medium-long | Severe at cluster scale | High |
| 9 | Advanced substrates/interposers | High in qualified supply | Low short term | Medium-long | High | High |
| 10 | AI-server liquid cooling | Moderate | Medium | Medium | High locally | Medium-high |
| 11 | Frontier-model talent | Highly concentrated | Low short term | Long | High for innovation | High |
| 12 | Conventional DRAM/NAND | Concentrated but cyclical | Medium | Medium | Moderate | Medium |
3.21 Disruption Transmission
The output impact of a component disruption is:
ΔCapacityAI = Elasticityj × ΔSupplyj
For strict complements, elasticity can approach one over the binding range: a 10% reduction in the bottleneck may reduce deliverable systems by close to 10%. For inputs with inventories or substitutes, short-run elasticity will be lower.
Table 3.18 — Bottleneck Failure Scenarios
| Disrupted node | Immediate effect | Second-order effect | Recovery condition |
|---|---|---|---|
| EUV equipment supply | Delayed new-node capacity and tool replacement | Reduced future accelerator output | Tool delivery, service and process recovery |
| Leading-edge foundry | Loss of logic-die production | GPU, CPU and custom-accelerator shortages | Alternative qualification or restored fab |
| HBM | Completed logic cannot be configured at target memory | Lower accelerator shipments or reduced specifications | Qualified alternative memory |
| Advanced packaging | Logic and memory inventories accumulate unfinished | Server delays and cloud-capacity shortfall | Packaging expansion and yield restoration |
| EDA/sign-off tools | New designs delayed | Slower architecture transition | Tool restoration or revalidation |
| Power connection | Installed systems cannot operate at intended scale | Lower utilisation and regional queueing | Grid and generation capacity |
| Networking | Accelerators operate below cluster potential | Higher cost per training or inference task | Alternative fabric and retuning |
| Cloud platform | Service interruption and workload migration | Enterprise operational disruption | Multi-cloud portability and restored service |
| Model API | Application failure despite available hardware | Business-process interruption | Model substitution and validation |
| Export licence | Geographic access terminated or delayed | Regional price divergence and substitution | Licence, redesign or domestic alternative |
3.22 Systemic Concentration: The Common-Dependency Problem
Downstream diversity can conceal upstream commonality. Multiple AI-model providers may use different models while relying on the same cloud, accelerator architecture, foundry, lithography equipment and HBM suppliers. The probability of correlated failure is therefore higher than a simple count of model providers suggests.
The Common Dependency Ratio is:
CDRj = DownstreamCapacityDependentOnNodej ÷ TotalDownstreamCapacity
The systemic loss expectation is:
E(Lossj) = ProbabilityDisruptionj × CDRj × EconomicValueAtRisk
Table 3.19 — Common-Dependency Layers
| Downstream appearance | Hidden common dependency |
|---|---|
| Several cloud providers | Same leading-edge foundry and equipment chain |
| Several model providers | Same cloud and accelerator infrastructure |
| Several accelerator brands | Same packaging, substrates or memory suppliers |
| Several server brands | Same accelerator allocation and networking |
| Several enterprise AI products | Same frontier-model API |
| Several national AI programmes | Same imported accelerators and software stack |
| Several local models | Same consumer GPU ecosystem and memory constraints |
3.23 Final Concentration Register
Table 3.20 — CRn and HHI Status by Segment
| Segment | CRn | HHI | Determination |
|---|---|---|---|
| Commercial EUV lithography | Conditional CR1 = 100% | Conditional HHI = 10,000 | Technical single-source condition; formal denominator not independently published |
| EDA | ND | ND | High structural concentration; exact primary shares unavailable |
| Leading-edge foundry | ND | ND | Very high concentration expected; no harmonised primary denominator |
| HBM | ND | ND | Three principal suppliers; exact shares unavailable |
| Conventional DRAM | ND | ND | Highly concentrated structure; exact current shares unavailable |
| NAND | ND | ND | More diversified than HBM/DRAM; exact shares unavailable |
| Advanced packaging | ND | ND | Product heterogeneity prevents simple revenue denominator |
| AI substrates/interposers | ND | ND | Qualified-capacity data unavailable |
| Merchant AI accelerators | ND | ND | Captive custom chips complicate market definition |
| AI servers | ND | ND | Conventional-server revenue is an invalid denominator |
| UK IaaS/PaaS, 2024 | CR2 = 60–80% | Strictly greater than 1,825 | Officially verifiable highly concentrated market |
| Frontier models | ND | ND | Market definition and usage denominator unresolved |
| Enterprise AI distribution | ND | ND | Must be separated by application category |
3.24 Final Judgment
The AI supply chain is not uniformly monopolised, but it is structurally concentrated around a small number of high-criticality nodes. The strongest concentration lies upstream in EUV lithography, leading-edge fabrication, EDA integration, HBM and advanced packaging; downstream concentration is most verifiable in public cloud infrastructure. The UK cloud case demonstrates that exact market shares are not necessary to establish a minimum structural conclusion: official ranges imply an HHI above the 1,800 high-concentration threshold even under the most concentration-minimising assumptions.
The architecture of scarcity is cumulative. Accelerator designers depend on foundries; foundries depend on lithography and EDA; accelerators depend on HBM and packaging; servers depend on networking and cooling; clouds depend on energy and server allocation; model developers depend on clouds; enterprise applications depend on models and installed software ecosystems. Scarcity at one layer can therefore transmit across the entire chain.
Profitability is likewise multi-causal. NVIDIA, TSMC and ASML report exceptional margins, but the evidence supports a combination of innovation rents, intellectual-property returns, scarcity rents, scale economies, risk compensation and potential market power. High margins alone do not establish unlawful conduct. A legal or competition finding would require defined markets, exclusionary conduct, counterfactual analysis and evidence that efficiencies do not explain the observed outcome.
The greatest five-year systemic vulnerabilities are EUV lithography, leading-edge foundry capacity, HBM, advanced packaging, EDA integration and accelerator software lock-in. Electricity and networking form the next-order constraints: even if semiconductor output expands, insufficient grid capacity or interconnection can prevent physical hardware from becoming usable compute. The central strategic conclusion is therefore that global AI capacity will not be determined by the number of chips designed, but by the smallest deployable quantity among all indispensable complements.
CHAPTER 4 — The Real Cost of Local AI
4.1 Scope, definitions and evidentiary boundary
The economic cost of local artificial intelligence cannot be inferred from the retail price of a computer, the nominal parameter count of a model, or the fact that a compressed checkpoint can technically be loaded into system memory. A local deployment becomes economically meaningful only when it can complete a specified workload at an acceptable level of quality, latency, throughput, concurrency, reliability, security and operational continuity. This chapter therefore defines usable local AI as a complete socio-technical system—not merely a model file—capable of satisfying explicit service-level and quality thresholds for its intended users. Its resource envelope includes accelerator memory, system memory, storage, interconnect bandwidth, electrical supply, cooling, networking, inference software, model licences, security controls, specialist labour, maintenance, replacement capacity, evaluation procedures and incident response. The analysis separates three evidentiary categories. First, hardware capacities and power limits attributed to named products are taken from live official manufacturer documentation. Second, memory, energy and cost relationships are derived from reproducible mathematical identities. Third, prices not publicly posted by manufacturers, future operating conditions and organisation-specific expenses are presented as transparent planning scenarios, not as observed market facts. This distinction is indispensable because enterprise accelerators, integrated servers, support contracts and secure infrastructure are frequently sold by quotation, while electricity tariffs, taxes, salaries, utilisation rates and compliance costs differ dramatically across jurisdictions. A false appearance of precision would be less scientific than a documented range accompanied by sensitivity analysis.
The physical constraints are already sufficient to disprove the proposition that every “large-memory computer” is interchangeable. An NVIDIA GeForce RTX 5090 provides 32 GB of GDDR7 memory, 1,792 GB/s of memory bandwidth and an official starting price of USD 1,999, whereas an NVIDIA H200 provides 141 GB of HBM3e and 4.8 TB/s of bandwidth. An eight-GPU NVIDIA DGX B200 integrates 1,440 GB of GPU memory, 64 TB/s of aggregate memory bandwidth and a maximum system power rating of approximately 14.3 kW. These are different economic objects: a consumer accelerator, a data-centre accelerator and a rack-scale integrated system. GeForce RTX 5090 Specifications – NVIDIA – January/2025 — NVIDIA GeForce RTX 5090 official specifications. NVIDIA H200 Tensor Core GPU – NVIDIA – November/2023 — NVIDIA H200 official product specification. NVIDIA DGX B200 – NVIDIA – 2026 — NVIDIA DGX B200 official specification. The memory capacity of a system therefore establishes only a feasibility boundary. It does not establish achievable tokens per second, time to first token, multi-user throughput, model quality, operational resilience or total cost.
4.2 The computational anatomy of a local LLM
4.2.1 Model-weight memory
For a dense model containing Np parameters stored at b bits per parameter, the raw weight-memory requirement is:
Mw = Np × b ÷ 8
where:
- Mw is raw model-weight memory in bytes;
- Np is the parameter count;
- b is the effective number of stored bits per parameter.
When Np is expressed in billions of parameters, the resulting number is conveniently expressed in decimal gigabytes:
Mw,GB = Np,B × b ÷ 8
The following table reports theoretical weight-only memory. It excludes quantisation scales, zero points, tensor metadata, KV cache, activations, execution workspaces, multimodal encoders, operating-system use and fragmentation.
| Dense-model size | FP32 | FP16/BF16 | FP8/INT8 | INT4 | 3-bit theoretical | 2-bit theoretical |
|---|---|---|---|---|---|---|
| 7B | 28 GB | 14 GB | 7 GB | 3.5 GB | 2.63 GB | 1.75 GB |
| 8B | 32 GB | 16 GB | 8 GB | 4 GB | 3 GB | 2 GB |
| 13B | 52 GB | 26 GB | 13 GB | 6.5 GB | 4.88 GB | 3.25 GB |
| 14B | 56 GB | 28 GB | 14 GB | 7 GB | 5.25 GB | 3.5 GB |
| 30B | 120 GB | 60 GB | 30 GB | 15 GB | 11.25 GB | 7.5 GB |
| 32B | 128 GB | 64 GB | 32 GB | 16 GB | 12 GB | 8 GB |
| 35B | 140 GB | 70 GB | 35 GB | 17.5 GB | 13.13 GB | 8.75 GB |
| 70B | 280 GB | 140 GB | 70 GB | 35 GB | 26.25 GB | 17.5 GB |
| 72B | 288 GB | 144 GB | 72 GB | 36 GB | 27 GB | 18 GB |
| 100B | 400 GB | 200 GB | 100 GB | 50 GB | 37.5 GB | 25 GB |
| 175B | 700 GB | 350 GB | 175 GB | 87.5 GB | 65.63 GB | 43.75 GB |
| 405B | 1,620 GB | 810 GB | 405 GB | 202.5 GB | 151.88 GB | 101.25 GB |
The table uses decimal gigabytes because hardware vendors commonly advertise capacity in decimal units. Conversion to binary gibibytes is:
MGiB = MGB ÷ 1.073741824
Consequently, 35 decimal GB of weights occupy approximately 32.60 GiB. This distinction matters near a device’s capacity limit, but it does not create usable headroom: a nominally 35 GB INT4 checkpoint cannot safely be treated as a complete 35 GB runtime.
4.2.2 Precision and quantisation regimes
| Representation | Nominal bits per parameter | Raw bytes per parameter | Principal advantage | Principal limitation |
|---|---|---|---|---|
| FP32 | 32 | 4.00 | High numerical range and reference-grade computation | Excessive memory and bandwidth for ordinary inference |
| FP16 | 16 | 2.00 | Broad accelerator support and high throughput | Smaller numerical range than BF16 |
| BF16 | 16 | 2.00 | FP32-like exponent range | Requires compatible hardware and kernels |
| FP8 | 8 | 1.00 | High accelerator throughput; major memory reduction | Hardware, calibration and kernel dependence |
| INT8 | 8 | 1.00 | Mature compression path with moderate quality risk | Speedup depends on native integer execution |
| INT4 | 4 | 0.50 | Makes 30–100B-class local inference materially more feasible | Quality and speed are method-, model- and kernel-dependent |
| 3-bit | 3 | 0.375 | Further capacity reduction | Metadata overhead becomes proportionally larger; quality risk rises |
| 2-bit | 2 | 0.25 | Extreme compression | Material model-specific degradation and limited kernel portability |
| Mixed precision | Variable | Variable | Preserves sensitive layers at higher precision | Effective bits per parameter must be measured, not assumed |
Nominal bit width does not equal effective storage cost. Group-wise quantisation stores scales and, depending on the format, zero points or codebooks. Embedding and output layers may remain at a higher precision. Alignment padding, tensor headers and duplicated buffers add further overhead. A defensible capacity calculation should therefore use:
beff = 8 × Scheckpoint ÷ Np
where Scheckpoint is the measured checkpoint size in bytes. Runtime memory must then be measured after model initialisation because some frameworks dequantise selected tensors, construct auxiliary lookup tables or reserve execution workspaces.
Quantisation is not an unconditional economic gain. It reduces memory traffic and may increase speed when the hardware and inference engine possess efficient kernels for the selected format. Conversely, a heavily compressed model can run more slowly when its kernels repeatedly dequantise weights into a higher-precision representation, when unsupported operations fall back to the CPU, or when heterogeneous offloading saturates the PCIe or system-memory interface. Quality loss is also workload-dependent: a compression method that preserves average language-model perplexity may still degrade code generation, multilingual reasoning, exact extraction, numerical work, tool use or rare-domain terminology. The correct comparison is therefore not “INT4 versus FP16” in the abstract, but the lowest-cost representation that passes a predefined task-specific evaluation suite.
4.3 Total operational memory
The correct operational identity is:
Mtotal = Mw + MKV + Mact + Mworkspace + Mmultimodal + MOS + Mfrag
This expands the requested expression by separating execution workspaces, multimodal components, operating-system reservation and fragmentation. The components cannot be represented by a universal fixed percentage because they respond differently to context length, batch size, model architecture, framework and concurrency.
4.3.1 KV-cache memory
For a decoder-only transformer, an architecture-aware approximation is:
MKV = 2 × L × HKV × Dhead × T × C × bKV ÷ 8
where:
- 2 represents the separately stored keys and values;
- L is the number of transformer layers;
- HKV is the number of key-value heads;
- Dhead is the head dimension;
- T is the number of cached tokens per sequence;
- C is the number of simultaneously resident sequences;
- bKV is the cache precision in bits.
If requests have unequal context lengths, the more general form is:
MKV = 2 × L × HKV × Dhead × bKV ÷ 8 × ΣTj
where ΣTj is the sum of cached tokens across all active sequences. This formulation avoids incorrectly multiplying a maximum context length by a nominal user count when many users are idle or have shorter prompts.
Two illustrative calculations show why context and concurrency can dominate memory planning. For a hypothetical 8B-class grouped-query model with 32 layers, eight KV heads, a head dimension of 128, a 32,768-token context, one active sequence and a 16-bit cache:
MKV = 2 × 32 × 8 × 128 × 32,768 × 1 × 16 ÷ 8
MKV = 4 GiB
For a hypothetical 70B-class grouped-query model with 80 layers and the same eight KV heads and 128-dimensional heads:
| Context and concurrency | KV-cache requirement |
|---|---|
| 8,192 tokens; one active sequence | 2.5 GiB |
| 32,768 tokens; one active sequence | 10 GiB |
| 131,072 tokens; one active sequence | 40 GiB |
| 32,768 tokens; four active sequences | 40 GiB |
| 32,768 tokens; eight active sequences | 80 GiB |
These are architectural illustrations, not universal values for every 8B or 70B model. Models with multi-head attention can require substantially more cache; multi-query attention can require less. KV-cache quantisation can reduce capacity consumption, but the resulting quality, latency and kernel compatibility must be tested.
4.3.2 Other operational allocations
| Allocation | Main determinants | Planning treatment |
|---|---|---|
| Activations, Mact | Batch size, prompt processing, hidden width, kernels, parallelism | Measure peak prefill and decode separately |
| Execution workspace, Mworkspace | Attention implementation, matrix kernels, graph capture | Empirical measurement is preferable to a fixed percentage |
| OS and display, MOS | Operating system, desktop display, background services, unified memory | Reserve capacity before model allocation |
| Fragmentation, Mfrag | Allocator, tensor sizes, model loading/unloading, long-running service | Stress-test over repeated requests; do not rely on a clean startup |
| Multimodal, Mmultimodal | Vision/audio encoders, projectors, image tokens, preprocessing | Add encoder weights and modality-dependent cache |
| Redundancy reserve | Failover model, rolling updates, duplicate workers | Required for production continuity but often omitted from hobby calculations |
A conservative deployment rule is:
Minstalled ≥ Mpeak,measured × (1 + h)
where h is an engineering headroom factor. A planning range of 0.10–0.25 may be reasonable for early design, but it is an assumption rather than a universal standard. Production acceptance should replace it with measured peak memory under the longest supported prompt, the maximum intended number of active sequences, simultaneous prefill events, tool calls and repeated model reloads.
4.3.3 Practical INT4 memory envelopes
| Model class | Raw INT4 weights | Indicative single-user operational envelope | Principal constraint |
|---|---|---|---|
| 7–8B | 3.5–4 GB | 8–12 GB | Long-context KV cache |
| 13–14B | 6.5–7 GB | 12–20 GB | Context, workspace and limited GPU headroom |
| 30–35B | 15–17.5 GB | 24–40 GB | Does not comfortably fit many 24 GB workloads at long context |
| 70–72B | 35–36 GB | 48–90 GB | KV cache, concurrency and multi-device communication |
| 100B | 50 GB | 70–130 GB | Sustained bandwidth and operational reserve |
| 175B | 87.5 GB | 120–240 GB | Multi-accelerator topology becomes decisive |
| 405B | 202.5 GB | 280–550 GB | Sharding, interconnect, reliability and energy |
| Large MoE | Based on total parameters | Often much larger than active-parameter count suggests | Expert storage and all-to-all communication |
The ranges are scenario envelopes, not specifications for individual models. They deliberately widen at larger sizes because model architecture, context, cache representation and parallelism produce increasingly large differences.
4.4 Mixture-of-experts, multimodal and frontier-scale systems
A mixture-of-experts model separates stored parameters from active parameters. If a model contains E experts but routes each token through only k experts, its arithmetic cost can approach that of the active subset while its weight-memory requirement remains related to total expert capacity:
Mw,MoE ≈ Ntotal × beff ÷ 8
Ctoken ≈ Cshared + k × Cexpert + Crouter
The active-parameter count therefore must never be substituted for total parameters when sizing storage or aggregate accelerator memory. Expert parallelism also introduces network communication: tokens must be routed to the accelerators holding the selected experts, creating all-to-all traffic, load imbalance and tail-latency risk. A nominally efficient MoE model can consequently perform poorly on consumer multi-GPU systems without high-bandwidth peer-to-peer links.
Multimodal deployment adds at least five components:
Mmultimodal = Mencoder + Mprojector + Mmodality-cache + Mpreprocess + Mgenerated-media
High-resolution images may produce hundreds or thousands of modality tokens. Video produces a sequence of frames and can multiply both preprocessing cost and context consumption. Speech systems add audio encoders, feature extraction and potentially speech synthesis. Image-generation models require diffusion or autoregressive decoders that can have a memory profile substantially different from text decoding. A computer that runs a text-only 70B model acceptably may therefore fail the equivalent multimodal service objective.
Near-frontier dense systems are physically feasible on sufficiently large integrated infrastructure but economically remote from ordinary users. A 405B dense model requires approximately 810 GB for FP16 weights alone or 202.5 GB at nominal INT4. The eight-GPU DGX B200’s official 1,440 GB aggregate GPU memory can contain 405B FP16 weights with material aggregate headroom, although actual feasibility still depends on the model, cache, kernels and parallelism. Its maximum rated power of approximately 14.3 kW implies that continuous operation would consume, before external cooling losses:
Aannual = 14.3 × 8,760 = 125,268 kWh
At a power-usage effectiveness factor of 1.3, facility energy would be approximately:
Afacility = 125,268 × 1.3 = 162,848 kWh per year
This is a maximum-rating scenario, not a measured annual load. It nevertheless demonstrates why frontier local AI moves from “computer ownership” into electrical, thermal and facility engineering.
4.5 What existing personal hardware can actually do
A system containing 128 GB of system RAM and a 12 GB GPU illustrates the difference between loadability and usability.
| Model class | Technical feasibility | Likely execution pattern | Usability judgment |
|---|---|---|---|
| 7–8B INT4/INT8 | Strong | Primarily GPU-resident | Suitable for interactive single-user work if model quality is adequate |
| 13–14B INT4 | Feasible but constrained | GPU-resident at short/moderate context or minor offload | Potentially usable; context and framework determine headroom |
| 30–35B INT4 | Weight file fits system RAM but not 12 GB VRAM | Heavy CPU execution or CPU/GPU offload | May be acceptable for patient single-user research; not assumed interactive |
| 70–72B INT4 | Fits system RAM in weight-only terms | Predominantly CPU/system-memory execution | Loadable, but sustained latency may fail practical thresholds |
| 100B INT4 | Raw weights fit 128 GB; full runtime may be marginal | Heavy CPU/offload | Experimental rather than dependable |
| 175B INT4 | Weight-only size may fit, runtime commonly does not | Inadequate headroom | Not operationally defensible |
| Large multimodal/MoE | Model-specific | Multiple memory pools and encoders | Requires exact architecture and workload testing |
No universal tokens-per-second value should be assigned to “a 12 GB GPU” because GPU model, memory bandwidth, CPU, system-memory bandwidth, PCIe generation, offload fraction, inference framework, quantisation kernel, context and sampling parameters all change performance. The correct procedure is to measure representative prompts and report the complete configuration.
4.6 Workload-specific definition of usable local AI
A deployment is usable only if it passes all applicable thresholds in the following vector:
U = f(Q, TPS, TTFT, C, T, A, R, P, G, O)
where:
- Q = task quality;
- TPS = output tokens per second;
- TTFT = time to first token;
- C = supported concurrency;
- T = supported effective context;
- A = availability;
- R = reliability and recovery;
- P = privacy and security controls;
- G = governance and auditability;
- O = operational update burden.
The following thresholds are proposed as falsifiable report criteria rather than universal regulatory standards.
| Workload | Minimum quality condition | Performance threshold | Context/concurrency threshold | Reliability and governance threshold |
|---|---|---|---|---|
| Personal assistant | Passes user-defined factual and instruction tests | Median ≥15 TPS; p95 TTFT ≤3 s | ≥8k effective tokens; one user | Recoverable service; encrypted storage |
| Student learning | Verified subject tests; citation fidelity measured | Median ≥12 TPS; p95 TTFT ≤4 s | ≥16k; one active user | Clear uncertainty and source checking |
| Coding assistant | Repository-specific test suite; accepted-patch rate measured | Median ≥20 TPS; p95 TTFT ≤3 s | ≥16k; one or two users | Sandboxed execution; rollback |
| Long-document analysis | Extraction recall and citation-location accuracy | Median ≥8 TPS; p95 TTFT ≤8 s | ≥32k or validated retrieval pipeline | Reproducible document ingestion |
| Independent research | Domain benchmark and provenance tests | Median ≥10 TPS | ≥32k; one to four users | Versioned model, prompts and corpora |
| Professional practice | Validated error and abstention thresholds | ≥12 TPS/user; p95 TTFT ≤4 s | Four or more concurrent sessions | Access control, audit log, backups |
| Clinical decision support | Prospective local validation; human review mandatory | Latency set by clinical workflow | Capacity under peak clinical load | High availability, traceability, incident response |
| Enterprise knowledge service | Task success and grounded-answer rate | ≥10 TPS/user; p95 TTFT ≤5 s | 50+ concurrent sessions or measured queue SLA | SSO, role control, monitoring, failover |
| Batch analytics | Validated batch accuracy | Completion within processing window | Throughput-based | Checkpointing and job recovery |
| Government-sensitive use | Mission-specific validation | Mission-defined worst-case latency | Surge capacity documented | Segmentation, supply-chain review, full auditability |
“Fits in memory” satisfies none of these criteria beyond basic physical loadability.
4.7 Capital architectures by user class
4.7.1 Reference deployment tiers
| Tier | Indicative architecture | Credible model envelope | Appropriate users |
|---|---|---|---|
| T1 — Entry local | 12–16 GB GPU; 64–128 GB RAM | 7–14B interactive; larger models through slower offload | Student, individual professional |
| T2 — Advanced workstation | 24–32 GB GPU; 128–256 GB RAM | 14–35B strong; 70B offloaded or constrained | Researcher, small practice |
| T3 — Large unified-memory or dual-accelerator workstation | 64–192 GB usable accelerator/unified pool | 35–100B quantised, depending on context | Research group, startup |
| T4 — Multi-GPU server | 192–640 GB aggregate accelerator memory | 70–405B quantised; smaller models at higher concurrency | SME, laboratory, hospital |
| T5 — Integrated enterprise system | 1 TB+ aggregate accelerator memory | Near-frontier dense inference and large MoE systems | Enterprise, government |
| T6 — Cluster | Multiple integrated systems with high-speed fabric | Large-scale serving, fine-tuning and resilience | Major enterprise, state or frontier laboratory |
Aggregate memory is not equivalent to a single coherent memory pool. Model weights must be sharded; communication can dominate latency; consumer cards may lack suitable peer-to-peer topology; failure probability rises with device count; and rolling updates can require duplicate capacity.
4.7.2 Planning ranges by institution
All monetary bands below are 2026 USD planning scenarios. They are not vendor quotations. They include broad implementation ranges because procurement geography, taxes, support, security and labour vary greatly.
| User or institution | Typical useful scale | Indicative initial CAPEX | Five-year cash TCO, excluding internal labour | Five-year full economic TCO, including specialist labour | Main economic constraint |
|---|---|---|---|---|---|
| Student | 7–14B; sometimes 30B offload | $1,500–$5,000 | $2,500–$9,000 | $4,000–$20,000 | Affordability and rapid obsolescence |
| Independent researcher | 14–35B; constrained 70B | $4,000–$15,000 | $7,000–$30,000 | $20,000–$100,000 | Time spent integrating and evaluating |
| Small professional practice | 14–70B | $8,000–$35,000 | $15,000–$80,000 | $50,000–$300,000 | Security, support and downtime |
| Startup | 35–100B; multi-user | $25,000–$200,000 | $70,000–$600,000 | $250,000–$2 million | Engineering payroll and utilisation risk |
| SME | 70–405B quantised or multiple smaller replicas | $75,000–$750,000 | $250,000–$3 million | $1–10 million | Reliability, concurrency and integration |
| University laboratory | 70B to near-frontier experiments | $100,000–$3 million | $400,000–$10 million | $2–30 million | Grants, facilities and specialist retention |
| Hospital/privacy-sensitive organisation | Validated 35–405B services | $250,000–$5 million | $1–20 million | $5–60 million | Governance, validation and high availability |
| Large enterprise | Multiple model classes and business units | $1–50 million | $5–200 million | $20–500 million | Platform engineering and governance |
| Government agency | Secure departmental system or sovereign cluster | $2–100+ million | $10–500+ million | $50 million–$1 billion+ | Security accreditation, resilience and sovereign supply |
These bands should not be combined into a single mean. The distributions are highly skewed, and institutional requirements create discontinuities. A hospital cannot simply purchase the student configuration multiplied by the number of clinicians; it requires access control, validated workflows, auditability, continuity, backups, change management and security monitoring. The HIPAA Security Rule, for example, requires regulated US entities to implement appropriate administrative, physical and technical safeguards for electronic protected health information. Summary of the HIPAA Security Rule – U.S. Department of Health and Human Services – August/2026 — HHS Summary of the HIPAA Security Rule. This does not prescribe a particular accelerator, but it changes the economic boundary of an acceptable deployment.
4.8 Five-year total cost of ownership
4.8.1 Correct discounted formulation
The requested TCO model becomes economically consistent when a residual value received in year five is discounted:
TCO5 = CAPEX + Σt=15 [(Et + Ct + Mt + St + Lt + Dt) ÷ (1+r)t] − RV5 ÷ (1+r)5
where:
- Et = electricity and cooling;
- Ct = connectivity and cloud overflow/support;
- Mt = maintenance, spares and warranties;
- St = software, licences and monitoring;
- Lt = specialised labour;
- Dt = downtime, failure and replacement cost;
- RV5 = year-five resale or residual value;
- r = real or nominal discount rate consistent with all cash flows.
If all cash flows are expressed in nominal dollars, r must be nominal and component prices should incorporate their expected inflation. If cash flows are constant-price real values, r must be real. Mixing nominal electricity escalation with a real discount rate overstates cost.
For a constant annual operating cost O and constant discount rate r:
PV(O) = O × [1 − (1+r)−5] ÷ r
At r = 7%, the five-year present-value factor is approximately 4.1002.
4.8.2 Electricity and cooling
Annual facility energy is:
KWhfacility,t = PIT,t × 8,760 × ut × PUEt
Et = KWhfacility,t × pe,t
where:
- PIT,t is average IT power while active, in kW;
- ut is the duty cycle;
- PUEt is facility power divided by IT power;
- pe,t is the electricity price per kWh.
For a personal workstation, a literal data-centre PUE may be inappropriate; cooling may already be embedded in household heating or air-conditioning. In that case, a cooling multiplier should be separately estimated. For a server facility, measured PUE is preferable.
US data centres consumed approximately 176 TWh, or 4.4% of US electricity, in 2023; the official federal assessment projected 325–580 TWh, or 6.7–12% of US electricity, by 2028. This is a system-level projection rather than a tariff forecast, but it identifies energy availability and grid connection as material cost risks. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – U.S. Department of Energy – December/2024 — U.S. Department of Energy data-centre electricity assessment.
4.8.3 Illustrative workstation TCO
Assume:
- CAPEX = $5,500;
- average active IT power = 0.65 kW;
- utilisation = 30%;
- cooling multiplier = 1.10;
- electricity = $0.15/kWh;
- maintenance and replacement reserve = $300/year;
- software and backup = $400/year;
- connectivity/cloud overflow = $200/year;
- downtime allowance = $250/year;
- specialist labour excluded from cash TCO;
- residual value in year five = 10% of CAPEX;
- discount rate = 7%.
Annual electricity is:
E = 0.65 × 8,760 × 0.30 × 1.10 × 0.15
E ≈ $282
Annual non-labour operating cost is:
O = 282 + 300 + 400 + 200 + 250
O = $1,432
Five-year present value is:
PV(O) = 1,432 × 4.1002
PV(O) ≈ $5,871
Present value of residual value is:
PV(RV5) = 550 ÷ 1.075
PV(RV5) ≈ $392
Therefore:
TCO5,cash = 5,500 + 5,871 − 392
TCO5,cash ≈ $10,979
If the owner spends 80 hours per year on model acquisition, quantisation compatibility, benchmarking, security, backups and troubleshooting, and time is valued at $50/hour, annual labour becomes $4,000. Its five-year present value is approximately $16,401, raising the full economic TCO to approximately $27,380. The example demonstrates that electricity is often not the dominant personal or small-organisation cost. Labour and underutilised capital can exceed it.
4.8.4 Sensitivity matrix
Using the same workstation, five-year discounted electricity cost varies as follows:
| Active utilisation | $0.10/kWh | $0.20/kWh | $0.40/kWh |
|---|---|---|---|
| 10% | $102 | $204 | $408 |
| 30% | $305 | $610 | $1,220 |
| 60% | $610 | $1,220 | $2,440 |
| 90% | $915 | $1,830 | $3,660 |
Assumptions: 0.65 kW active power, 1.10 cooling multiplier, 7% discount rate and constant real electricity prices. The matrix shows why consumer inference economics are often dominated by hardware and labour, while continuously utilised server economics become substantially more sensitive to energy and cooling.
4.9 Cost per million tokens
4.9.1 Energy intensity
If output is generated continuously at Rtok tokens per second:
KWhMtok = PIT × PUE × 1,000,000 ÷ (Rtok × 3,600)
For a 0.5 kW workstation producing 20 output tokens per second with a cooling multiplier of 1.10:
KWhMtok = 0.5 × 1.10 × 1,000,000 ÷ (20 × 3,600)
KWhMtok ≈ 7.64 kWh
At $0.15/kWh:
EnergyCostMtok ≈ $1.15
This excludes prompt-prefill energy, idle power, failed generations, embeddings, retrieval, preprocessing and system administration. It also assumes sustained decoding, which a lightly used personal system rarely achieves.
4.9.2 Levelised cost
A more complete levelised measure is:
LCMtok = TCOperiod ÷ [Tokenssuccessful,period ÷ 1,000,000]
“Successful” tokens should mean tokens belonging to completed workloads that pass quality and usefulness criteria. Counting rejected, hallucinated or unusable output understates economic cost.
For the illustrative workstation TCO of $10,979:
| Successful output over five years | Levelised cash cost |
|---|---|
| 100 million tokens | $109.79 per million |
| 500 million tokens | $21.96 per million |
| 1 billion tokens | $10.98 per million |
| 5 billion tokens | $2.20 per million |
Including the illustrative labour cost and using the $27,380 full economic TCO:
| Successful output over five years | Full economic cost |
|---|---|
| 100 million tokens | $273.80 per million |
| 500 million tokens | $54.76 per million |
| 1 billion tokens | $27.38 per million |
| 5 billion tokens | $5.48 per million |
Local AI becomes economically competitive only when utilisation is sufficiently high, the selected local model passes the required quality threshold, and privacy or control benefits are valued. A low-quality local answer is not a cheaper substitute for a successful frontier-model answer if it requires repetition, correction or professional review.
Local deployment can reduce the transmission of prompts and documents to an external inference provider, but it does not automatically provide privacy. The threat surface includes operating-system telemetry, model download channels, malicious model artefacts, compromised dependencies, unencrypted logs, cached prompts, vector databases, administrative accounts, remote-management tools, physical theft, backup media and inference-server APIs exposed without adequate authentication. NIST’s AI Risk Management Framework treats AI risk management as an organisational lifecycle covering governance, mapping, measurement and management rather than as a property delivered by hardware location alone. Artificial Intelligence Risk Management Framework 1.0 – National Institute of Standards and Technology – January/2023 — NIST AI RMF 1.0.
Production reliability creates a hidden capacity multiplier. If a service must remain available while a model is updated, a second runnable copy may be necessary. If the system cannot tolerate a failed GPU, it needs spare capacity or an external fallback. The availability relation for a repairable component is approximately:
A = MTBF ÷ (MTBF + MTTR)
where MTBF is mean time between failures and MTTR is mean time to repair. End-to-end availability is lower when all serial components must function. For n independent serial components:
Asystem = ΠAi
A system containing accelerators, storage, network switches, cooling and software services can therefore be less reliable than any one component. High-availability design often duplicates the inference endpoint, storage, network path and power supply; consequently, the capital cost of a dependable service can approach 1.5–2.5 times that of the minimum system capable of running the model.
4.11 Benchmark protocol
Every local system evaluated later in the report should disclose:
| Category | Mandatory fields |
|---|---|
| Hardware | Exact accelerator, count, memory, CPU, RAM channels, storage, interconnect |
| Software | OS, driver, inference framework, kernel/library versions |
| Model | Exact checkpoint, parameter count, architecture, licence, quantisation format |
| Workload | Prompt length, output length, context occupancy, sampling settings |
| Performance | Prefill tokens/s, decode tokens/s, median and p95 TTFT |
| Concurrency | Active sequences, queue depth, per-user throughput, p95 latency |
| Memory | Idle, model-loaded, peak prefill, peak decode, fragmentation after endurance run |
| Energy | Wall-socket IT energy, facility multiplier, idle and active power |
| Quality | Domain benchmark, task-success rate, abstention and factual-error rate |
| Reliability | Failure rate, restart time, update downtime, sustained-load duration |
| Economics | CAPEX, tax, electricity, support, labour, residual value, utilisation |
| Privacy | Data flow, logs, encryption, access control, telemetry, retention |
| Reproducibility | Seeds, scripts, test corpus version and measurement date |
A valid comparison must report both prefill and decode performance. Long prompts may be limited by prefill compute, while conversational output is often limited by memory bandwidth. Average latency must not replace p95 latency in multi-user services because batching and queueing can conceal severe tail delays.
4.12 Decision matrix by user class
| User class | Minimum defensible local target | When local deployment is rational | When it is not rational |
|---|---|---|---|
| Student | 7–14B quantised | Learning, offline notes, coding support, sensitive drafts | When frontier quality is essential and utilisation is low |
| Independent researcher | 14–35B; optional 70B offload | Reproducibility, confidential corpora, sustained experiments | When integration time exceeds research value |
| Small practice | 14–70B with managed security | Repeated domain workflow with high privacy value | When no staff member can own updates and evaluation |
| Startup | Replicated small models or 35–100B server | Stable workload, high utilisation, proprietary data | When demand is uncertain and capital must remain flexible |
| SME | Multi-user server plus cloud overflow | Predictable volume and integration with internal systems | When redundancy and staffing erase unit-cost savings |
| University lab | Multi-GPU research server | Model experimentation, controlled datasets, training access | When funding covers hardware but not operations |
| Hospital | Validated model service with redundancy | Sensitive data and clinically validated bounded tasks | When local model quality or governance is inadequate |
| Large enterprise | Portfolio of local, private-cloud and external services | High volume, sovereignty, differentiated workflows | When a single procurement is expected to solve heterogeneous needs |
| Government agency | Segmented secure inference infrastructure | Classified, sovereign or mission-critical processing | When supply-chain, staffing or lifecycle support is unresolved |
4.13 Five-year outlook, 2027–2031
Over the next five years, raw arithmetic efficiency and quantisation quality are likely to improve, but those improvements will not eliminate local-AI inequality. The binding resource is moving from raw parameter storage toward a composite of memory bandwidth, context capacity, concurrency, multimodal processing, software optimisation, evaluation competence and operational reliability. Smaller models will become increasingly capable for bounded workloads, allowing individuals to obtain useful private assistants without frontier-scale infrastructure. At the same time, the definition of competitive capability will move upward as remote frontier systems gain stronger reasoning, tool use, multimodality and long-context performance. This creates a moving-target effect: the hardware cost of reproducing today’s capability may fall while the cost of matching the current frontier remains high.
Three trajectories should be tested:
| Scenario | Hardware trajectory | Local-access effect | Institutional consequence |
|---|---|---|---|
| Diffusion | Memory capacity rises; efficient models improve rapidly | 7–70B capability becomes broadly affordable | Local AI becomes a normal workstation function |
| Bifurcation | Consumer capacity improves, but frontier requirements rise faster | Basic access broadens; frontier access remains concentrated | Two-tier capability market persists |
| Constriction | Memory, packaging, energy and export constraints remain severe | Useful local systems remain expensive outside wealthy markets | Sovereignty programmes and regional exclusion intensify |
The central inference is that local AI is not one market. It comprises at least four economic regimes: personal offline inference, professional single-node service, institutional multi-user deployment and frontier-scale sovereign infrastructure. The marginal cost of tokens can be low in a highly utilised system, but the fixed cost of entering each successive regime rises discontinuously because memory, redundancy, security and labour requirements arrive in indivisible blocks.
4.14 Testable conclusions
- Weight memory is a necessary but insufficient capacity measure. A 70B INT4 dense model has approximately 35 GB of raw weights, but long-context or concurrent operation can add tens of gigabytes of KV cache.
- Quantisation changes the feasible model class but not automatically the economically usable model class. Quality, kernel support and offload behaviour must be benchmarked.
- A 128 GB RAM system with a 12 GB GPU can load models much larger than its GPU memory, but loadability does not establish interactive usability. The economically defensible envelope is generally smaller than the physical RAM envelope.
- For individuals, capital and specialist time commonly dominate electricity. For continuously used institutional systems, energy, cooling, maintenance and redundancy become progressively more important.
- The cost per million useful tokens is utilisation-sensitive. Low utilisation can leave local inference more expensive than metered external access even when marginal electricity cost is small.
- Privacy-sensitive deployment adds costs rather than simply removing cloud costs. Security, logs, access control, model provenance, backups and incident response remain necessary.
- Near-frontier local inference is technically feasible only in the institutional meaning of “local.” It may run inside an organisation’s facility but still require rack-scale power, cooling, networking and specialist operations.
- The most scientifically valid definition of usable local AI is workload-specific. It must include quality, p95 latency, throughput, concurrency, effective context, reliability, security and maintenance.
- The five-year access divide should be measured against frontier-equivalent task success, not parameter count. Cheap small models may broaden basic access while leaving economically consequential frontier capability concentrated.
- Every later cost comparison must report both cash TCO and full economic TCO. Excluding internal labour is appropriate for cash-budget analysis but materially understates social and organisational resource consumption.
CHAPTER 5 — Cloud AI Economics and the Future Price of Tokens
5.1 Scope, analytical objective and evidentiary boundary
Cloud AI converts an exceptionally capital-intensive production system into an apparently simple metered service. The customer sees an API price per million input or output tokens, a monthly subscription, a reserved-capacity contract or a fee for completing a task. Behind that price lies a vertically layered cost structure comprising model research, training, post-training, evaluation, accelerators, high-bandwidth memory, host servers, networking, storage, data-centre construction, land, electricity, cooling, water, maintenance, specialist labour, software engineering, customer acquisition, financing and the economic cost of idle capacity. Token prices therefore cannot be interpreted as if they represented only the electricity consumed during inference. Nor can a low public API price be treated as proof that the corresponding model is inexpensive to produce. Providers can allocate training expenditure across future usage, recover costs through enterprise contracts, bundle AI with higher-margin software, cross-subsidise adoption from advertising or cloud profits, accept a temporary return below the cost of capital, or price strategically to create ecosystem dependence. Conversely, a provider can reduce prices without subsidising usage when hardware efficiency, quantisation, speculative decoding, batching, prompt caching, model routing or utilisation improves faster than capital costs rise. The empirical task is consequently not to assert that all present AI services are loss leaders, but to determine whether disclosed revenues, margins, capital expenditure, depreciation, service tiers and pricing differentials are consistent with competitive introductory pricing, cross-subsidisation or sustainable economic production.
This chapter uses three evidence classes. Observed prices are taken from live official provider documentation as verified on 15 August 2026. Corporate economics are drawn from official investor-relations disclosures, recognising that the major public companies do not publish a complete standalone profit-and-loss account for frontier-model inference. Unit-cost estimates and projections are analytical scenarios generated from explicitly stated parameters. The absence of a separately audited AI-inference segment prevents any scientifically valid claim that a named provider is definitively losing a specified amount per token. The chapter can test whether the loss-leader hypothesis is supported, contradicted or left unresolved; it cannot manufacture missing cost accounts.
5.2 Cloud AI as a production system
The basic economic chain is:
- Research and model architecture.
- Data acquisition, cleaning, licensing and governance.
- Pre-training.
- Post-training and alignment.
- Safety testing, red teaming and evaluation.
- Model optimisation and compilation.
- Deployment across inference clusters.
- Capacity orchestration, caching and routing.
- API, subscription and enterprise distribution.
- Customer support, compliance and incident management.
- Model replacement and infrastructure refresh.
This sequence creates two distinct but interacting cost pools:
FC = Cresearch + Ctraining + Cpost + Csafety + Csoftware + Cfacility-fixed + Ccommercial-fixed
VC = Caccelerator-time + Cmemory-time + Cenergy + Ccooling + Cnetwork + Cstorage-variable + Csupport-variable + Ctools
The distinction is not perfectly clean. Accelerator depreciation behaves as a committed fixed cost after a cluster is purchased, but it becomes a usage-attributable cost in product accounting. Specialist labour may be fixed within a planning horizon but scalable over longer periods. Data-centre capacity can be owned, financed, leased or reserved through another cloud provider. Model training is sunk after completion, yet continuous retraining and replacement transform it into a recurring portfolio expense. The appropriate classification therefore depends on the decision being studied:
| Decision | Economically relevant cost |
|---|---|
| Whether to serve one additional request on already-idle hardware | Short-run marginal cost |
| Whether to discount API prices for one quarter | Avoidable operating cost plus strategic objective |
| Whether to build another inference cluster | Incremental full cost and required return |
| Whether a model family is sustainable over five years | Fully allocated lifecycle cost |
| Whether an AI company earns economic profit | Revenue minus all operating costs and cost of capital |
| Whether a customer should use cloud or local AI | Customer-facing price plus integration, risk and switching costs |
A provider can rationally price above short-run marginal cost but below fully allocated average cost while spare capacity exists. Such pricing is not necessarily predatory or irrational; it may maximise contribution margin during demand ramp-up. It becomes structurally unsustainable if revenue remains insufficient to replace depreciating accelerators, finance facilities, develop successor models and compensate invested capital.
5.3 The complete cloud-AI cost stack
5.3.1 Cost-stack architecture
| Cost layer | Principal driver | Fixed, variable or mixed | Appropriate allocation base | Principal uncertainty |
|---|---|---|---|---|
| Accelerator purchase | Device count, memory, interconnect | Fixed after purchase | Accelerator-hours or successful tokens | Economic life and residual value |
| Host server | CPU, RAM, storage, NICs | Fixed | Server-hours | Shared use across models |
| Accelerator depreciation | Purchase cost, life, utilisation | Mixed in unit accounting | Productive accelerator-seconds | Obsolescence may precede physical failure |
| Networking | Fabric, switches, optics, WAN | Mixed | Bytes, requests, accelerator-hours | MoE and distributed-inference intensity |
| Storage | Weights, datasets, logs, caches | Mixed | GB-month, read/write operations | Replication and retention policies |
| Data-centre construction | Buildings, electrical and cooling systems | Fixed | IT capacity, rack-kW, useful life | Construction lead time and stranded capacity |
| Land | Location and grid proximity | Fixed | Site capacity | Connection availability and planning law |
| Electricity | IT load and tariff | Variable | kWh | Regional tariffs and peak charges |
| Cooling | Climate, PUE, water system | Variable/mixed | Facility energy and heat removed | Weather and water restrictions |
| Water | Cooling technology and climate | Variable | Litres per kWh or workload | Site-specific measurement |
| Maintenance | Failure rates, spares, warranties | Mixed | Installed assets | Component scarcity |
| Specialist labour | Engineers, researchers, security staff | Primarily fixed | Product, cluster or revenue | Shared R&D allocation |
| Pre-training | Compute, data, researchers | Sunk per model | Lifetime billable usage | Model lifetime and training disclosure |
| Post-training | Fine-tuning, preference data, reinforcement learning | Recurrent fixed | Model family or usage | Iteration frequency |
| Safety and evaluation | Red teams, benchmarks, governance | Recurrent fixed | Model or release | Scope increases with capability |
| Inference software | Kernels, compilers, schedulers | Recurrent fixed | Product or fleet | Proprietary efficiency gains |
| Customer acquisition | Sales, credits, advertising | Mixed | Customer cohort or revenue | Promotional credits and channel costs |
| Enterprise compliance | Certifications, legal, data residency | Mixed | Contract or region | Regulatory fragmentation |
| Financing | Debt, leases, equity capital | Fixed/required return | Invested capital | Weighted average cost of capital |
| Idle capacity | Unused but installed infrastructure | Economic overhead | Billable usage | Demand volatility |
| Failed work | Errors, retries, rejected outputs | Variable waste | Successful tasks | Rarely visible in token accounts |
The most frequently omitted item is idle capacity. Cloud inference must retain enough spare capacity to absorb demand variation, equipment failures and latency commitments. A fleet operated at 100% nominal utilisation would exhibit queues and unacceptable tail latency. If only 45% of theoretical accelerator time produces commercially billable tokens, each billed token must absorb more than twice the depreciation that would apply at 100% utilisation.
5.3.2 Accelerator depreciation
Let:
- Kacc = installed accelerator and server capital;
- RVn = expected residual value after n years;
- n = useful economic life;
- Hyear = 8,760 hours;
- uproductive = share of time producing billable work.
Straight-line annual depreciation is:
Dannual = (Kacc − RVn) ÷ n
Depreciation per productive accelerator-hour is:
Dhour = Dannual ÷ (8,760 × uproductive)
Consider an analytical cluster costing USD 40 million, with a four-year economic life and a 10% residual value.
Dannual = (40,000,000 − 4,000,000) ÷ 4
Dannual = USD 9,000,000
| Productive utilisation | Productive hours per year | Depreciation per productive cluster-hour |
|---|---|---|
| 25% | 2,190 | USD 4,110 |
| 40% | 3,504 | USD 2,568 |
| 60% | 5,256 | USD 1,712 |
| 75% | 6,570 | USD 1,370 |
| 90% | 7,884 | USD 1,142 |
Utilisation is therefore a first-order economic variable. A 20% improvement in raw accelerator speed is economically inferior to doubling productive utilisation if the faster system remains underused. Providers with large, diverse demand pools possess an important scale advantage because they can route traffic across customers, regions, models, interactive requests and batch jobs.
Microsoft reported quarterly capital expenditure of USD 31.9 billion in fiscal Q3 2026, with approximately two-thirds allocated to shorter-lived assets, principally GPUs and CPUs. The company simultaneously reported that continuing AI-infrastructure investment and growing AI usage reduced its cloud gross-margin percentage. Microsoft Fiscal Year 2026 Third Quarter Earnings Conference Call – Microsoft – April/2026 — Microsoft FY2026 Q3 investor disclosure. This is direct evidence that accelerator depreciation and AI usage are economically material. It does not disclose the cost of any particular model or token.
5.3.3 Data-centre construction, land, networking and finance
The total capital base supporting inference is broader than the accelerator fleet:
K = Kaccelerators + Kservers + Knetwork + Kstorage + Kelectrical + Kcooling + Kbuilding + Kland + Kconstruction-in-progress
Alphabet reported USD 35.7 billion of capital expenditure in Q1 2026, with the overwhelming majority directed to technical infrastructure supporting AI opportunities. Approximately 60% of that technical-infrastructure investment went to servers and 40% to data centres and networking equipment. Alphabet 2026 Q1 Earnings Call – Alphabet – April/2026 — Alphabet Q1 2026 investor disclosure. In Q2 2026, Alphabet reported USD 44.9 billion of capital expenditure and retained approximately the same 60:40 division between servers and data-centre/networking infrastructure. Alphabet 2026 Q2 Earnings Call – Alphabet – July/2026 — Alphabet Q2 2026 investor disclosure.
The disclosure establishes that analysing accelerator cost alone would omit roughly 40% of Alphabet’s recent technical-infrastructure allocation. It also demonstrates why a provider’s cash burden can rise before revenue is recognised: land, substations, buildings, cooling systems, networking and servers must be financed and constructed before capacity becomes billable. The resulting financing charge is:
Ccapital = WACC × Kemployed
where WACC is the weighted average cost of capital. Economic break-even must include this required return even if accounting operating profit is positive. A provider earning 4% on an infrastructure base while its risk-adjusted cost of capital is 9% is accounting-profitable but destroys economic value.
5.3.4 Electricity, cooling and water
Facility electricity attributable to a workload can be represented as:
Eworkload = PIT × t × PUE
Cenergy = Eworkload × pelectricity
where PUE is total facility energy divided by IT-equipment energy. Water use should be tracked separately:
Wworkload = EIT × WUE
where WUE is water-use effectiveness, normally expressed as litres per kWh of IT energy. Neither PUE nor WUE should be replaced by a universal global value. Climate, cooling design, water source, load, building age and accounting boundaries differ across facilities.
At the national level, US data centres consumed approximately 176 TWh, or 4.4% of total US electricity, in 2023; the official assessment projected 325–580 TWh, or 6.7–12% of US consumption, by 2028. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – U.S. Department of Energy – December/2024 — US Department of Energy data-centre electricity assessment. This does not prove that token prices must increase: efficiency, model compression and higher utilisation can offset electricity inflation. It does establish that grid access, power contracts and transmission capacity are becoming strategic inputs rather than negligible operating details.
5.3.5 Training, post-training and safety
The lifecycle cost attributable to a model family can be written as:
Cmodel-lifecycle = Cresearch + Cdata + Cpretrain + Cposttrain + Csafety + Cevaluation + Cdeployment + Cretirement
An allocation per successful commercial token would be:
Cmodel,token = Cmodel-lifecycle ÷ Tsuccessful,lifetime
The denominator is highly uncertain at launch. A successful model with trillions of billable tokens can amortise training expenditure widely. A model displaced within months may never recover its development cost. Frequent frontier releases consequently shorten the effective economic life of both model investments and optimised inference stacks.
Safety and evaluation cannot scientifically be treated as optional overhead. Their scope includes capability evaluations, domain testing, security reviews, red teaming, incident investigation, abuse monitoring, model-behaviour analysis and release governance. However, public provider price pages do not reveal how much of each token price is allocated to these activities.
5.4 Unit cost of inference
5.4.1 Core unit-cost equation
The requested inference-cost expression is:
Ctoken = [Ccompute + Cmemory + Cenergy + Cnetwork + Cfacility + Clabour + Ccapital] ÷ Tbillable
For commercial interpretation, this should be expanded to:
Csuccessful-token = [Cinference-direct + Callocated-model + Callocated-platform + Ccommercial + Crisk] ÷ Tsuccessful
where Tsuccessful excludes failed requests, free usage, promotional credits, internal tests and outputs rejected by the customer’s application.
A token-weighted blended price is:
Pblend = [PinTin + PcachedTcached + PoutTout + PreasonTreason + Ctools] ÷ Ttotal
A task-level cost is:
Ctask = PinTin + PoutTout + Pcache-writeTcache-write + Pcache-readTcache-read + ΣCtool,j + Cstorage + Cretrieval
This formulation demonstrates why advertised input-token rates cannot reliably predict the cost of an agentic workflow.
5.4.2 Provider break-even
The requested break-even relationship is:
PBE = [FC + VC + R × K] ÷ Q
where:
- FC = attributable fixed cost;
- VC = attributable variable cost;
- K = invested capital;
- R = required return on invested capital;
- Q = commercially billable usage.
If promotional or internal use consumes capacity, a capacity-adjusted formulation is preferable:
PBE = [FC + VC + R × K] ÷ [Qphysical × ubillable × ysuccess]
where:
- ubillable is the billable share of physical capacity;
- ysuccess is the share producing commercially successful work.
Illustrative break-even calculation
Assume an inference platform with:
| Parameter | Scenario value |
|---|---|
| Annual fixed costs | USD 240 million |
| Annual variable costs | USD 310 million |
| Invested capital | USD 2.5 billion |
| Required return | 10% |
| Annual billable usage | 100 trillion tokens |
Required annual revenue is:
RevenueBE = 240m + 310m + 0.10 × 2,500m
RevenueBE = USD 800 million
Break-even price per million tokens is:
PBE,MTok = 800,000,000 ÷ 100,000,000
PBE,MTok = USD 8.00
The 100 trillion-token denominator equals 100 million units of one million tokens. Sensitivity is substantial:
| Billable annual tokens | Break-even price per million tokens |
|---|---|
| 25 trillion | USD 32.00 |
| 50 trillion | USD 16.00 |
| 100 trillion | USD 8.00 |
| 200 trillion | USD 4.00 |
| 400 trillion | USD 2.00 |
The example does not estimate any named provider. It establishes the mathematical significance of scale: identical infrastructure economics can support either USD 32 or USD 2 per million tokens depending on billable throughput.
5.5 Verified public pricing architecture as of 15 August 2026
5.5.1 Selected direct API prices
The following comparison reports official list prices per one million tokens. It is not a quality-equivalence table: models differ in capability, tokenizer, speed, context, reliability and supported tools.
| Provider and model | Standard input | Cached input or cache hit | Standard output | Batch input | Batch output | Important modifier |
|---|---|---|---|---|---|---|
| OpenAI GPT-5.6 Sol | USD 5.00 | USD 0.50 | USD 30.00 | USD 2.50 | USD 15.00 | Long context: USD 10 input and USD 45 output |
| OpenAI GPT-5.6 Terra | USD 2.00 | USD 0.20 | USD 12.00 | USD 1.00 | USD 6.00 | Long context: USD 4 input and USD 18 output |
| OpenAI GPT-5.6 Luna | USD 0.20 | USD 0.02 | USD 1.20 | USD 0.10 | USD 0.60 | Long context: USD 0.40 input and USD 1.80 output |
| OpenAI GPT-5.5 | USD 5.00 | USD 0.50 | USD 30.00 | Provider table identifies separate batch rates | Provider table identifies separate batch rates | Above 272k input: 2× input and 1.5× output |
| OpenAI GPT-5.5 Pro | USD 30.00 | No cached-input discount | USD 180.00 | Supported | Supported | Higher-compute professional tier |
| Anthropic Claude Opus 4.8 | USD 5.00 | USD 0.50 cache hit | USD 25.00 | USD 2.50 | USD 12.50 | Five-minute cache write USD 6.25 |
| Anthropic Claude Sonnet 5 | USD 2.00 | USD 0.20 cache hit | USD 10.00 | USD 1.00 | USD 5.00 | Standard price retained after introductory period |
| Anthropic Claude Sonnet 4.6 | USD 3.00 | USD 0.30 cache hit | USD 15.00 | USD 1.50 | USD 7.50 | Regional endpoints may add 10% |
| Anthropic Claude Haiku 4.5 | USD 1.00 | USD 0.10 cache hit | USD 5.00 | USD 0.50 | USD 2.50 | One-hour cache write costs more than five-minute write |
| Google Gemini 3.6 Flash, 2026 standard | USD 0.75 | USD 0.075 | USD 3.75 | USD 0.375 | USD 1.875 | Officially scheduled doubling from 1 January 2027 |
| Google Gemini 3.6 Flash, 2026 priority | USD 1.35 | USD 0.135 | USD 6.75 | Not the priority mode | Not the priority mode | Latency/priority premium |
| Google Gemini 3.5 Flash | USD 1.50 | USD 0.15 | USD 9.00 | USD 0.75 | USD 4.50 | Output includes thinking tokens |
Pricing – OpenAI API – August/2026 — Official OpenAI API pricing.
GPT-5.5 Model – OpenAI API – August/2026 — Official GPT-5.5 model pricing.
GPT-5.5 Pro Model – OpenAI API – August/2026 — Official GPT-5.5 Pro model pricing.
Pricing – Claude Platform Docs – August/2026 — Official Anthropic Claude pricing.
Gemini Developer API Pricing – Google – August/2026 — Official Gemini API pricing.
Three empirical observations follow.
First, output tokens are consistently more expensive than ordinary input tokens. The ratios in the selected standard tiers range from approximately 5:1 to 8:1. Output decoding is sequential and less readily parallelised than prompt processing; reasoning may add internal generation; and providers also price according to value and demand, not merely physical cost.
Second, cached inputs can cost approximately one-tenth of uncached inputs, while writing a cache can cost more than ordinary input. OpenAI states that GPT-5.6-family cache writes are billed at 1.25 times the uncached-input rate. Prompt Caching – OpenAI API – August/2026 — Official OpenAI prompt-caching documentation. Anthropic similarly distinguishes five-minute and one-hour cache writes from cheaper cache hits. The tariff therefore encourages repeated use of stable context but discourages indiscriminate creation of caches that are never reused.
Third, batch processing commonly receives a 50% discount. Anthropic and Google publish half-price batch rates for selected models, and Amazon Bedrock states that selected foundation models receive a 50% batch-inference discount relative to on-demand pricing. Amazon Bedrock Pricing – Amazon Web Services – August/2026 — Official Amazon Bedrock pricing. A consistent 50% discount is economically meaningful: it indicates that scheduling flexibility, smoother utilisation and relaxed latency materially reduce the provider’s opportunity cost.
5.5.2 Long-context pricing
Context length affects cost through at least four mechanisms:
- More input tokens are billed.
- Prefill computation increases.
- KV-cache memory remains occupied during generation.
- Long requests can reduce batching and scheduling efficiency.
A general context multiplier can be represented as:
Pcontext = Pbase × mlength
where mlength may be continuous, tiered or embedded in a model-specific price.
OpenAI’s official August 2026 table distinguishes short- and long-context prices for GPT-5.6 models. For GPT-5.6 Sol, the listed short-context rates are USD 5 input and USD 30 output per million tokens; the corresponding long-context rates are USD 10 and USD 45. Google and Anthropic apply different structures: Google lists model- and service-tier prices, while Anthropic states that Claude 4.6 and later models include the full one-million-token context window at standard token rates. These differences show that “long-context premium” can appear as an explicit multiplier, a higher model tier, cache-storage charges or simply the larger number of tokens processed.
Illustrative long-document task
Assume:
- 500,000 uncached input tokens;
- 20,000 output tokens;
- no tool calls.
| Model/tier | Input cost | Output cost | Total |
|---|---|---|---|
| GPT-5.6 Sol long-context | USD 5.00 | USD 0.90 | USD 5.90 |
| GPT-5.6 Terra long-context | USD 2.00 | USD 0.36 | USD 2.36 |
| Claude Sonnet 5 standard | USD 1.00 | USD 0.20 | USD 1.20 |
| Gemini 3.6 Flash 2026 standard | USD 0.375 | USD 0.075 | USD 0.45 |
The comparison does not establish equivalent answer quality. A lower-cost model becomes economically more expensive if it fails the task, requires repeated calls or produces an answer needing extensive human correction.
5.5.3 Subscriptions versus metered API access
Consumer subscriptions transform variable usage into an apparently fixed monthly payment:
ARPUsubscription = Feemonthly ÷ User
The provider’s contribution per subscriber is:
CMuser = Feemonthly − Cinference,user − Csupport,user − Cacquisition,user − Callocated-platform,user
A fixed subscription creates a usage-risk distribution. Light users subsidise heavy users within the subscriber pool unless rate limits, model routing or usage credits control the tail. Consequently, consumer plans generally include fair-use policies, flexible access, message limits, lower-priority service during congestion or automatic routing among models.
Subscriptions also perform strategic functions that a pure API tariff does not:
- accelerate habit formation;
- increase switching costs;
- create persistent conversation and file ecosystems;
- generate usage data for product improvement;
- provide predictable recurring revenue;
- bundle high-cost and low-cost functions;
- obscure the marginal price of individual tasks;
- segment customers by willingness to pay.
Enterprise contracts add identity management, data controls, audit functions, support, contractual commitments, regional processing, service levels and procurement integration. Their effective unit price cannot be inferred from public consumer fees.
5.5.4 Reserved, priority and batch capacity
| Commercial mode | Customer receives | Provider receives | Economic effect |
|---|---|---|---|
| On-demand standard | Flexibility without commitment | Variable demand and uncertain utilisation | Standard list price |
| Batch | Lower price; delayed completion | Scheduling freedom and higher utilisation | Discount |
| Flex or low-priority | Lower price with variable latency | Ability to use residual capacity | Discount or variable service |
| Priority/fast | Lower tail latency and preferential scheduling | Higher revenue per constrained capacity unit | Premium |
| Reserved capacity | Guaranteed or predictable throughput | Committed revenue and demand visibility | Contracted capacity price |
| Regional processing | Geographic control | Reduced routing flexibility and regional duplication | Geographic premium |
| Dedicated deployment | Isolation and control | Long commitment and customer concentration | Negotiated price |
| Surge access | Immediate capacity during scarcity | Scarcity rent | Dynamic premium |
| Spot/interruptible | Very low price with interruption risk | Monetisation of otherwise idle capacity | Deep discount |
Amazon Bedrock officially identifies Reserved, Priority, Standard and Flex service tiers. Service Tiers for Optimizing Performance and Cost – Amazon Web Services – August/2026 — Official Amazon Bedrock service-tier documentation. This represents an observable movement away from one uniform token price toward a yield-management model similar to cloud computing, telecommunications and transport.
5.6 Fine-tuning, retrieval, storage and agentic-tool costs
5.6.1 Fine-tuning
The full cost of a fine-tuned model is:
CFT,total = Ctraining-tokens + Cdata-preparation + Cevaluation + Chosting + Cinference + Cmaintenance
Fine-tuning can lower prompt length by internalising instructions or domain patterns, but it introduces dataset construction, validation, versioning and regression-testing costs. The break-even condition is:
Savingsprompt + Savingsquality + Savingshuman-review > CFT,total
A fine-tuned model that reduces each request by 10,000 input tokens but requires costly retraining after every base-model update may never achieve positive lifecycle value.
5.6.2 Retrieval and storage
Retrieval-augmented generation generates several charges beyond text generation:
CRAG = Cembedding + Cvector-storage + Cindexing + Cretrieval + Creranking + Cretrieved-input + Cgeneration
The customer can reduce model hallucination and avoid very long prompts, but only if retrieval precision is high. Poor retrieval increases both token expenditure and error rates.
5.6.3 Agentic tool use
An agent can call search, code execution, databases, browsers, external APIs or other models. Its cost is recursive:
Cagent = Σi=1n Cmodel,i + Σj=1m Ctool,j + Cstorage + Corchestration + Cverification
OpenAI’s official pricing lists web search at USD 10 per 1,000 calls, in addition to applicable model-token charges. Pricing – OpenAI API – August/2026 — Official OpenAI API and tool pricing.
Illustrative agentic workflow
Assume a research task uses:
- 200,000 input tokens;
- 20,000 output tokens;
- 30 web-search calls;
- GPT-5.6 Sol standard short-context pricing.
Model cost:
Cmodel = 0.2 × 5 + 0.02 × 30
Cmodel = USD 1.60
Search cost:
Csearch = 30 × 10 ÷ 1,000
Csearch = USD 0.30
Total direct provider charge:
Ctask = 1.60 + 0.30
Ctask = USD 1.90
If the agent enters an error loop and repeats the sequence five times, direct cost rises to USD 9.50 even though the final deliverable remains one task. Future enterprise pricing is therefore likely to emphasise budgets per completed task and controls on recursive tool use.
5.7 Is current AI pricing subsidised?
5.7.1 Definition of subsidisation
The term “subsidised” should be divided into four distinct propositions:
| Hypothesis | Test |
|---|---|
| Short-run loss leader | Price is below avoidable marginal cost |
| Full-cost under-recovery | Price covers marginal cost but not allocated training, infrastructure and commercial costs |
| Below-required-return pricing | Accounting profit exists but return on capital is below WACC |
| Cross-subsidisation | Loss or low return in AI is financed by another product, segment or investor capital |
These propositions require different evidence. Low price alone proves none of them.
5.7.2 Evidence consistent with subsidisation
Several observations are consistent with introductory or strategic underpricing:
- Capital expenditure has expanded before the full corresponding revenue stream is realised.
- AI infrastructure and usage have pressured cloud gross margins.
- Providers offer free tiers, credits, discounted batch processing and low-priced entry models.
- Subscription bundles can permit heavy users to consume more inference than their monthly fee would purchase at retail API prices.
- Model providers have strong incentives to acquire developers, establish default APIs and create ecosystem dependence.
- Training and post-training expenditure is not separately itemised in token prices.
- Rapid model replacement can shorten the amortisation period of research and infrastructure optimisation.
Microsoft reported a 66% Microsoft Cloud gross margin in fiscal Q3 2026, down year over year because of continued AI investment, while company capital expenditure reached USD 31.9 billion. Microsoft Fiscal Year 2026 Third Quarter Earnings Conference Call – Microsoft – April/2026 — Microsoft FY2026 Q3 investor disclosure. This supports the proposition that AI is capital-intensive and margin-dilutive at the disclosed cloud level.
5.7.3 Evidence against a universal loss-leader conclusion
The same evidence does not establish that every token is sold below cost:
- Microsoft Cloud remained strongly gross-profitable at the aggregate level.
- Google Cloud has reported positive operating income while expanding AI infrastructure.
- Batch and cached-input discounts may reflect genuine cost savings from scheduling and reuse.
- Output-token premiums may produce high contribution margins on selected workloads.
- Model routing can direct simple tasks toward materially cheaper models.
- Large providers can use their demand aggregation to achieve utilisation unavailable to smaller firms.
- Enterprise contracts and reserved capacity may recover costs not visible in consumer pricing.
- Advertising, productivity software and cloud-platform revenue can monetise AI indirectly without making the AI service economically irrational.
5.7.4 Formal test result
| Proposition | Finding as of August 2026 |
|---|---|
| AI infrastructure is capital-intensive | Strongly supported |
| AI investment pressures some disclosed margins | Supported |
| Some access is intentionally discounted or free | Supported |
| Batch/flex prices reflect utilisation economics | Strongly supported |
| All consumer AI subscriptions are loss-making | Not established |
| All API tokens are sold below marginal cost | Not established |
| Frontier-model lifecycle costs are fully recovered by present token prices | Not verifiable from public disclosures |
| Strategic cross-subsidisation exists somewhere in the market | Plausible and consistent with business structures, but not quantifiable from disclosed segment accounts |
| Material future repricing is possible | Supported as a risk, not a certainty |
The scientifically defensible conclusion is therefore narrower than the popular “AI companies lose money on every query” claim. Public filings support capital intensity, margin pressure and strategic investment ahead of demand. They do not disclose the model-level cost accounts required to prove universal loss-leading.
5.8 The emerging pricing system
5.8.1 Price per input token
Input-token pricing approximates prompt-processing demand, but it ignores semantic value. A million repetitive tokens and a million highly specialised tokens can cost the same despite producing different business outcomes. Caching will increasingly split input into:
- uncached input;
- cache-write input;
- cache-hit input;
- persistent-context storage.
5.8.2 Price per output token
Output pricing reflects sequential generation and captures willingness to pay. It can penalise verbose models and create incentives to produce concise answers. The effective output cost should be:
Ceffective-output = Cgenerated-output ÷ Taccepted-output
A model generating twice as much text for the same accepted result is economically less efficient even at the same listed token rate.
5.8.3 Price per reasoning token
Reasoning-token pricing monetises internal computation more directly. Its risks include limited customer observability, incentives for providers to allow inefficient reasoning and difficulty comparing models with different internal processes. A defensible system would disclose:
- visible output tokens;
- billed reasoning tokens;
- maximum reasoning budget;
- completed-task success;
- refunds or caps for failed reasoning.
5.8.4 Price per completed task
Task pricing is likely for bounded, verifiable work:
Ptask = BaseFee + mcomplexity + mlatency + mrisk + Ctools
Examples include document classification, invoice extraction, code repair, translation, compliance review or customer-service resolution. Task pricing transfers execution-efficiency risk from the customer to the provider: if an agent uses excessive tokens, the provider absorbs the difference.
5.8.5 Price per unit of compute
A compute-based tariff could use accelerator-seconds, FLOP-equivalents or memory-bandwidth consumption:
Pcompute = aHaccelerator + bBmemory + cNnetwork
This is transparent for infrastructure buyers but unsuitable for non-technical users because identical compute can produce radically different quality.
5.8.6 Priority and latency multipliers
Ppriority = Pstandard × mpriority
The official Google table already distinguishes standard, batch, flex and priority prices, and Amazon Bedrock identifies Reserved, Priority, Standard and Flex tiers. This is direct evidence that latency is becoming separately monetised.
5.8.7 Context-length multipliers
Plong-context = Pbase × mcontext
Long-context pricing can be explicit or embedded in model tiers. Future systems may price according to maximum resident KV-cache allocation rather than only tokens actually processed.
5.8.8 Model-quality tiers
Providers already differentiate economical, balanced, frontier and high-reasoning systems. The market may evolve toward:
| Tier | Economic function |
|---|---|
| Commodity | Classification, extraction, routing |
| Professional | General business and coding |
| Frontier | Complex reasoning and research |
| Expert | Domain-calibrated regulated work |
| Sovereign | Geographic, security and control guarantees |
5.8.9 Reserved-capacity contracts
A reserved contract can be expressed as:
Creserved = Fcapacity + PoverageQoverage
The customer exchanges flexibility for guaranteed throughput; the provider reduces demand uncertainty and financing risk.
5.8.10 Surge pricing
Psurge,t = Pbase × [1 + λSt]
where St measures scarcity or queue pressure and λ is price sensitivity. Surge pricing could ration scarce frontier capacity during major product launches, financial events, elections, cyber incidents or regional outages. It would improve allocation efficiency but deepen inequality by allowing wealthy customers to purchase priority during scarcity.
5.8.11 Outcome-based pricing
Poutcome = Fbase + θVverified
where Vverified is measurable customer value and θ is the provider’s share. Outcome pricing is plausible for recovered revenue, resolved support cases, qualified sales leads or completed software tasks. It is difficult where causation, attribution or quality is disputed.
5.8.12 Auctioned compute capacity
An auction could allocate a fixed block of frontier inference:
Pclearing = bid of the marginal accepted buyer
Auctions could emerge for very large training runs, sovereign capacity, urgent simulation or guaranteed frontier-agent service. However, they would introduce volatility, strategic bidding and market-manipulation risks.
5.9 Five-year pricing scenarios, 2027–2031
5.9.1 Scenario framework
| Scenario | Probability range | Core mechanism | Token-price direction | Access effect |
|---|---|---|---|---|
| Efficiency deflation | 25–35% | Hardware and software efficiency outpace demand | Commodity-model prices fall sharply | Broad basic access |
| Bifurcated market | 35–50% | Small-model costs fall while frontier compute remains scarce | Low tiers fall; frontier tiers remain high or rise | Capability inequality persists |
| Infrastructure repricing | 15–30% | Depreciation, energy and financing require fuller recovery | Broad price increases or tighter limits | SMEs and lower-income users pressured |
| Capacity shock | 5–15% | Export control, conflict, grid shortage or supply disruption | Surge and regional premiums | Severe geographic inequality |
| Outcome transition | 20–40% | Agents are sold by task rather than token | Token prices lose relevance | Greater price discrimination |
Ranges overlap because scenarios can coexist across models and regions.
5.9.2 Most likely structure
The central five-year case is not a universal increase in the nominal price of every token. It is a bifurcation:
- small and medium models become cheaper;
- cached and asynchronous work becomes heavily discounted;
- high-priority frontier reasoning remains expensive;
- long-context, multimodal and agentic workloads acquire additional meters;
- enterprise data residency and reserved capacity carry premiums;
- consumer subscriptions become more explicitly usage-limited;
- complex tasks migrate toward outcome or credit-based pricing;
- free tiers rely increasingly on lower-cost model routing;
- regional scarcity creates geographic price differences.
The user-facing unit may therefore change from “one million tokens” to an opaque composite credit representing model quality, reasoning depth, tools, latency and context.
5.10 Is AI compute becoming like Bitcoin?
5.10.1 Where the analogy is useful
The analogy is useful in five limited respects.
First, both AI compute and Bitcoin-related activity depend on scarce specialised hardware, electricity and access to infrastructure. Second, both can exhibit scarcity premiums when demand expands faster than supply. Third, both can be geographically concentrated near favourable electricity, regulation or supply chains. Fourth, expectations about future value can accelerate present investment and speculative behaviour. Fifth, capacity can be represented in tradable contracts, reservations, futures or financial claims even when the underlying resource is physical.
An “AI compute index” could therefore measure the spot or forward price of a standardised capacity bundle:
ACIt = P(Hstandard, Mmemory, Bbandwidth, Llatency, Rregion, Aavailability)
Such an index could support budgeting, hedging and infrastructure contracts.
5.10.2 Where the analogy fails
| Property | Bitcoin | AI compute |
|---|---|---|
| Supply rule | Protocol-defined issuance and maximum supply | Produced through manufacturing and construction |
| Fungibility | Units are protocol-standardised | GPUs, TPUs, memory, models and locations differ |
| Durability | Digital ledger asset does not physically wear out | Hardware depreciates and becomes obsolete |
| Storage | Can be held without productive use | Idle compute still incurs depreciation and facility cost |
| Transportability | Transferable across a global ledger | Bound to data centres, grids, networks and jurisdiction |
| Store of value | Can be held as an asset | Unused past compute cannot generally be stored |
| Quality | One unit is equivalent to another at protocol level | One accelerator-hour can vary radically in capability |
| Location dependence | Ownership transfer is largely location-independent | Latency, sovereignty and export rules are location-sensitive |
| Production response | Issuance cannot exceed protocol | Supply can expand, though with long lead times |
| Expiry | Asset does not expire technologically | Reserved capacity expires; hardware ages |
| Demand basis | Monetary/speculative/network demand | Productive demand for computation and intelligence |
| Unit standardisation | Native unit exists | No universal equivalent across models and workloads |
Compute is closer to a combination of electricity, cloud capacity, airline seats and industrial machinery than to Bitcoin. Like electricity, it must be generated and transmitted through infrastructure. Like cloud capacity, it is heterogeneous and service-dependent. Like airline seats, unused time expires and priority can be dynamically priced. Like machinery, it requires capital, maintenance and replacement.
The most important distinction is temporal non-storability. A Bitcoin held today remains a Bitcoin tomorrow. An unused accelerator-second at 14:00 cannot be sold at 15:00. Providers are therefore incentivised to discount batch and flexible workloads to fill otherwise idle capacity. This is yield management, not monetary scarcity.
5.11 Empirical tests for the next five years
| Test ID | Proposition | Required data | Method | Falsification condition |
|---|---|---|---|---|
| C5-T1 | Present prices under-recover full lifecycle cost | Provider cost accounts, capex, depreciation, usage | Fully allocated unit economics | Revenue per successful token persistently exceeds cost plus required return |
| C5-T2 | Batch discounts reflect utilisation benefits | Batch and on-demand throughput/cost | Difference-in-differences | No material provider cost reduction from scheduling flexibility |
| C5-T3 | Output tokens have higher physical cost | Measured prefill/decode resource use | Accelerator telemetry | Output premium remains despite equalised physical cost and market power controls |
| C5-T4 | Frontier prices remain high while commodity prices fall | Historical model-quality-adjusted prices | Hedonic panel model | Frontier-equivalent price falls at same rate as basic-model price |
| C5-T5 | Subscriptions cross-subsidise heavy users | User-level usage and fees | Cohort contribution analysis | Every usage decile produces non-negative fully allocated contribution |
| C5-T6 | Data residency raises cost | Regional versus global endpoints | Matched price comparison | No premium after quality and SLA controls |
| C5-T7 | Agentic systems shift billing toward tasks | Provider pricing catalogues | Product-panel analysis | Token-only billing remains dominant through 2031 |
| C5-T8 | Scarcity increases price discrimination | Priority, surge and reserved prices | Panel regression | Tier dispersion falls as utilisation rises |
| C5-T9 | Infrastructure costs pressure margins | Capex, depreciation and cloud margins | Distributed-lag regression | Higher AI capital intensity consistently raises margins immediately |
| C5-T10 | Access inequality grows at the frontier | Regional income and frontier usage | Income-elasticity analysis | Lower-income regions achieve convergent frontier usage per capita |
A provider-level margin model can be estimated as:
GMj,t = α + β1AIKj,t + β2Uj,t + β3Ej,t + β4Dj,t + μj + τt + εj,t
where:
- GMj,t = gross margin of provider j at time t;
- AIKj,t = AI-related capital intensity;
- Uj,t = productive utilisation;
- Ej,t = energy cost exposure;
- Dj,t = depreciation burden;
- μj = provider fixed effects;
- τt = period effects.
Because companies do not consistently disclose AI-only capital expenditure, the model must report measurement uncertainty and avoid interpreting cloud-wide correlations as pure inference economics.
5.12 Strategic implications by user class
| User | Principal cloud advantage | Principal cloud risk | Recommended economic control |
|---|---|---|---|
| Student | Frontier quality without hardware CAPEX | Subscription increases and usage limits | Mixed free, subscription and local-small-model strategy |
| Researcher | Access to multiple frontier models | Reproducibility and price volatility | Log exact model, tokens, tools and cost per experiment |
| Startup | Elastic capacity and rapid deployment | Margin dependence on provider tariffs | Model routing, caching and contractual price protection |
| SME | Avoids specialist infrastructure | Lock-in and uncontrolled agentic expenditure | Budget caps, multi-provider abstraction and task costing |
| University | Access to high capability | Grant budgets exposed to metered usage | Reserved research allocations and institutional procurement |
| Hospital | Managed enterprise controls | Sensitive-data and dependency risks | Regional processing, contractual audit rights and local fallback |
| Large enterprise | Scale and global availability | Concentration and switching cost | Reserved capacity plus portable evaluation layer |
| Government | Immediate access to frontier systems | Sovereignty, export and continuity risks | Sovereign capacity, diversified supply and emergency local capability |
5.13 Final assessment
The available evidence supports a nuanced conclusion. Cloud AI is not simply “cheap computation sold at a loss,” nor is it an ordinary mature software service with negligible marginal cost. It is a capital-intensive infrastructure business combined with an unusually rapid research cycle and a platform-competition strategy. Public filings demonstrate extraordinary infrastructure investment and observable margin pressure. Official pricing demonstrates strong segmentation by model quality, input/output direction, caching, context, latency, geography and scheduling flexibility. What public evidence does not demonstrate is that all current inference is priced below short-run marginal cost or that every subscription is loss-making.
The greatest five-year risk is not necessarily a uniform multiplication of every posted token price. Providers may continue reducing the cost of basic inference while shifting cost recovery into frontier reasoning, output generation, priority service, long context, data residency, tools, storage, agentic loops and guaranteed capacity. The market can therefore appear deflationary at the headline price-per-token level while becoming more expensive for economically valuable completed work.
The probable endpoint is a multi-dimensional tariff:
PAI = f(M, Tin, Tout, Treason, L, C, G, A, R, X, V)
where:
- M = model-quality tier;
- Tin = input tokens;
- Tout = output tokens;
- Treason = reasoning tokens or effort;
- L = latency or priority;
- C = context and cache allocation;
- G = geographic processing requirement;
- A = guaranteed availability;
- R = reserved capacity;
- X = external tools and retrieval;
- V = verified task value or outcome.
AI access will consequently resemble a metered strategic utility more than a single software subscription. Commodity intelligence may become abundant, while reliable frontier intelligence—delivered with long context, low latency, tool access, privacy, geographic control and contractual guarantees—remains scarce and expensive. That division, rather than the nominal price of one million generic tokens, will determine the economic and social distribution of AI capability between 2027 and 2031.
CHAPTER 6 — Enterprise Exposure and the Cost of Mandatory AI Adoption
6.1 Analytical scope and central proposition
Enterprise AI adoption is entering a phase in which expenditure may become competitively necessary before its financial return can be demonstrated conclusively. A company may purchase licences, API capacity, cloud infrastructure, local accelerators, data-engineering services, cybersecurity controls and specialist personnel not because a particular project has already produced a positive net present value, but because customers, competitors, employees, regulators and investors increasingly expect AI-enabled speed, personalisation, analytics and automation. This produces a new category of expenditure: the defensive competitive input. Like cybersecurity, digital connectivity or enterprise software, AI may become difficult to avoid even when its attributable productivity gain remains uncertain. The economic risk is asymmetrical. A company that adopts too slowly may lose customers, talent and operational efficiency; a company that adopts too quickly may accumulate duplicated licences, fragmented pilots, proprietary lock-in, compliance liabilities and infrastructure whose utilisation remains too low to recover its cost.
Observed adoption is rising rapidly but remains highly unequal by enterprise size. In 2025, 20.0% of EU enterprises with at least ten employees reported using AI technologies, compared with 13.5% in 2024, 8.1% in 2023 and 7.7% in 2021. Large-enterprise adoption reached 55.03%, materially exceeding adoption among smaller firms. Use of Artificial Intelligence in Enterprises – Eurostat – December/2025 — Eurostat enterprise AI statistics. OECD data similarly indicate that reported firm-level AI use increased to 20.2% in 2025 from 14.2% in 2024 and 8.7% in 2023 across countries with available data. AI Use by Individuals Surges Across the OECD as Adoption by Firms Continues to Expand – OECD – January/2026 — OECD firm-level AI adoption statistics.
These figures establish diffusion, not profitability. A survey response indicating use of at least one AI technology does not disclose implementation depth, expenditure, productivity, risk-adjusted return or whether the system is embedded in a critical workflow. This chapter therefore constructs representative financial archetypes. Every monetary result is an analytical planning scenario expressed in 2026 USD unless otherwise stated; it is not an observed account for a named company.
6.2 Enterprise-size definitions and exposure
The European Commission defines an SME through staff headcount and either turnover or balance-sheet total. SMEs employ fewer than 250 people and ordinarily have annual turnover not exceeding EUR 50 million or a balance-sheet total not exceeding EUR 43 million. SME Definition – European Commission – August/2026 — European Commission SME definition.
For modelling, this chapter uses the following operational classes:
| Enterprise class | Employees | Illustrative annual revenue | AI-adoption characteristics |
|---|---|---|---|
| Microenterprise | 1–9 | USD 0.25–2 million | Subscription-led; little internal technical capacity |
| Small enterprise | 10–49 | USD 2–15 million | SaaS and managed-service adoption |
| Medium enterprise | 50–249 | USD 15–100 million | Hybrid integration; emerging governance function |
| Large corporation | 250+ | USD 100 million–10 billion+ | Portfolio adoption, proprietary integration and dedicated governance |
| Systemically or strategically critical organisation | Variable | Budget or revenue can exceed USD 1 billion | High assurance, resilience, sovereignty and regulatory obligations |
The economic importance of small firms makes unequal adoption a macroeconomic problem rather than a niche technology issue. Eurostat reported that micro and small businesses represented 99% of EU enterprises in 2022, employed 77.5 million people and generated EUR 11.9 trillion of turnover. Large firms represented only 0.2% of enterprises but generated 51% of turnover. Micro and Small Businesses Make Up 99% of Enterprises in the EU – Eurostat – October/2024 — Eurostat enterprise-size statistics. If AI entails large fixed costs for governance, data preparation and integration, those costs will consume a larger share of small-firm revenue even where API prices are identical.
6.3 Core enterprise financial metrics
6.3.1 AI Cost Ratio
AI Cost Ratio = Annual AI Expenditure ÷ Revenue
The ratio measures the share of revenue absorbed by AI expenditure. It should be calculated both gross and net of capitalised development:
AI Cost Ratiocash = AI Cash Expenditure ÷ Revenue
AI Cost RatioP&L = AI Expense Recognised in Period ÷ Revenue
A company purchasing infrastructure may record a large cash outflow but recognise depreciation over several years. The cash ratio measures financing pressure; the income-statement ratio measures current accounting burden.
6.3.2 AI Labour Burden
AI Labour Burden = Annual AI Expenditure ÷ Annual Labour Cost
This ratio compares AI expenditure with the labour base from which productivity benefits are commonly expected. It is economically informative because an AI programme costing 15% of annual labour expense must create a very large improvement in labour productivity, revenue or risk reduction to break even.
6.3.3 Net AI Value
Net AI Value = Productivity Gains + Incremental Revenue − Risk Costs − AI Expenditure
For financial accuracy, incremental revenue should be converted into incremental contribution, not treated as entirely profit:
Net AI Value = Labour Savings + Non-labour Savings + m × Incremental Revenue + Avoided Losses − Risk Costs − AI Expenditure
where m is the contribution margin on incremental revenue.
6.3.4 Return on investment
ROIAI = [NPV(Benefits) − NPV(Costs)] ÷ NPV(Costs)
The net present value of benefits is:
NPV(Benefits) = Σt=0n Bt ÷ (1+r)t
The net present value of costs is:
NPV(Costs) = Σt=0n Ct ÷ (1+r)t
The project creates financial value only when:
NPV(Benefits) > NPV(Costs)
and:
ROIAI > 0
For a risk-adjusted investment hurdle h, acceptance requires:
ROIAI ≥ h
A positive ROI below the company’s risk-adjusted hurdle rate can still represent value destruction relative to alternative investments.
6.4 The complete enterprise AI cost stack
6.4.1 Direct and indirect costs
| Cost category | Included expenditure | Primary driver | Frequently omitted element |
|---|---|---|---|
| User licences | Per-seat AI assistants and embedded software | Number of enabled users | Paid but inactive seats |
| API consumption | Input, output, reasoning and multimodal usage | Workload volume and model tier | Failed calls and agent loops |
| Cloud infrastructure | Compute, databases, storage and networking | Data volume and availability | Egress and regional-processing premiums |
| Local infrastructure | Servers, accelerators, storage and facilities | Model size, concurrency and resilience | Spares, replacement and cooling |
| Cybersecurity | Identity, monitoring, testing and incident response | Threat profile and system criticality | Model-specific red teaming |
| Compliance | Legal assessment, documentation and conformity work | Sector, jurisdiction and risk classification | Continuous evidence maintenance |
| Data preparation | Cleaning, labelling, permissions and lineage | Data fragmentation and quality | Rights remediation |
| Governance | Policies, inventory, committees and controls | Number and criticality of systems | Shadow AI discovery |
| Model evaluation | Benchmarks, validation, bias and robustness tests | Use-case risk | Regression testing after model updates |
| Human supervision | Review, escalation and override | Error severity and regulatory exposure | Reviewer fatigue |
| Integration | APIs, workflow redesign, testing and migration | Legacy-system complexity | Process redesign |
| Specialist personnel | AI engineers, data scientists, security and legal staff | Scope and assurance | Recruitment and retention premium |
| Vendor switching | Re-engineering, data migration and evaluation | Proprietary dependence | Temporary parallel operation |
| Downtime | Lost output, service interruption and recovery | Availability and fallback capacity | Queue accumulation after recovery |
| Model error | Correction, rework, customer harm and liability | Error probability and consequence | Correlated systematic failure |
| Privacy-preserving deployment | Isolation, encryption, local inference and auditing | Data sensitivity | Lower utilisation from dedicated capacity |
| Training | AI literacy and role-specific competence | Workforce size | Refresher training after updates |
| Change management | Process redesign, communication and adoption | Organisational complexity | Productivity dip during transition |
| Opportunity cost | Capital and staff diverted from other projects | Portfolio constraints | Abandoned alternatives |
Annual AI expenditure should therefore be calculated as:
AIE = Llicence + AAPI + Ccloud + Ilocal + Ssecurity + Ccompliance + Ddata + Ggovernance + Eevaluation + Hsupervision + Iintegration + Ppersonnel + Vswitching + Ddowntime + Rerror
A narrow estimate including only licences and API charges can understate full economic expenditure by several multiples.
6.4.2 Risk-adjusted error costs
The expected annual cost of model errors is:
E(Risk Cost) = Σj=1m pj × Lj
where:
- pj is the annual probability of error class j;
- Lj is its total loss if realised.
Loss should include:
Lj = Ccorrection + Cdowntime + Ccustomer + Clegal + Cregulatory + Creputation + Cremediation
A low-frequency error can dominate AI economics when its consequence is large. An AI system producing USD 2 million in annual process savings but carrying a 2% probability of a USD 100 million loss has an expected loss of USD 2 million before risk aversion, capital requirements or tail-risk limits are considered.
For risk-averse institutions, expected value is insufficient. A more appropriate decision rule is:
Risk-Adjusted Net AI Value = Expected Benefits − Expected Costs − λ × Tail Risk
where λ reflects the organisation’s risk tolerance and Tail Risk can be measured through value at risk, expected shortfall or scenario loss.
6.5 Representative enterprise archetypes
6.5.1 Baseline annual scenarios
The following table applies one consistent annual model. “Productivity gains” represent realised cash savings or additional productive capacity assigned a defensible monetary value. Incremental revenue is converted to contribution using the margins embedded in each scenario. Risk cost includes expected error, downtime, privacy and liability losses.
| Archetype | Revenue or operating budget | Labour cost | Annual AI expenditure | Productivity gain | Incremental contribution | Risk cost | Net AI value |
|---|---|---|---|---|---|---|---|
| Micro professional firm | USD 0.80m | USD 0.32m | USD 0.022m | USD 0.018m | USD 0.008m | USD 0.004m | USD 0.000m |
| Small enterprise | USD 5m | USD 1.5m | USD 0.150m | USD 0.120m | USD 0.060m | USD 0.030m | USD 0.000m |
| Medium enterprise | USD 30m | USD 7m | USD 0.600m | USD 0.560m | USD 0.240m | USD 0.100m | USD 0.100m |
| Large diversified company | USD 5bn | USD 1.2bn | USD 65m | USD 75m | USD 40m | USD 15m | USD 35m |
| Bank | USD 10bn | USD 3bn | USD 180m | USD 200m | USD 90m | USD 70m | USD 40m |
| Insurer | USD 5bn | USD 1bn | USD 90m | USD 100m | USD 50m | USD 45m | USD 15m |
| Healthcare organisation | USD 2bn | USD 1.1bn | USD 55m | USD 70m | USD 10m | USD 35m | −USD 10m |
| Legal/professional firm | USD 100m | USD 55m | USD 4m | USD 5.5m | USD 1.5m | USD 1m | USD 2m |
| Manufacturer | USD 1bn | USD 180m | USD 18m | USD 20m | USD 12m | USD 6m | USD 8m |
| Media company | USD 200m | USD 70m | USD 6m | USD 7m | USD 4m | USD 3m | USD 2m |
| Telecommunications operator | USD 5bn | USD 900m | USD 100m | USD 120m | USD 80m | USD 25m | USD 75m |
| Defence/critical-infrastructure operator | USD 3bn | USD 700m | USD 75m | USD 65m | USD 30m | USD 35m | −USD 15m |
| Public administration | USD 2bn budget | USD 900m | USD 60m | USD 70m | USD 0m | USD 30m | −USD 20m |
| University/research system | USD 1bn budget | USD 600m | USD 35m | USD 35m | USD 5m | USD 20m | −USD 15m |
These figures are not sector averages. They are transparent archetypes designed to identify financial thresholds. Negative results do not mean AI is socially undesirable or strategically avoidable. A hospital, defence organisation or public authority may rationally accept a negative accounting return to obtain safety, sovereignty, service quality or strategic capability. The correct conclusion is that such deployments require a broader public-value or mission-value framework rather than a false claim of immediate commercial ROI.
6.5.2 AI Cost Ratio and AI Labour Burden
| Archetype | AI Cost Ratio | AI Labour Burden | Annual return on AI expenditure before discounting |
|---|---|---|---|
| Micro professional firm | 2.75% | 6.88% | 0.0% |
| Small enterprise | 3.00% | 10.00% | 0.0% |
| Medium enterprise | 2.00% | 8.57% | 16.7% |
| Large diversified company | 1.30% | 5.42% | 53.8% |
| Bank | 1.80% | 6.00% | 22.2% |
| Insurer | 1.80% | 9.00% | 16.7% |
| Healthcare organisation | 2.75% | 5.00% | −18.2% |
| Legal/professional firm | 4.00% | 7.27% | 50.0% |
| Manufacturer | 1.80% | 10.00% | 44.4% |
| Media company | 3.00% | 8.57% | 33.3% |
| Telecommunications operator | 2.00% | 11.11% | 75.0% |
| Defence/critical infrastructure | 2.50% | 10.71% | −20.0% |
| Public administration | 3.00% of budget | 6.67% | −33.3% |
| University/research system | 3.50% of budget | 5.83% | −42.9% |
The calculation demonstrates why the same nominal expenditure has different consequences across sectors. Legal services may support a relatively high AI Cost Ratio because labour represents a large part of cost and document-intensive work is potentially augmentable. Manufacturing may have a lower AI Cost Ratio but still a high AI Labour Burden because revenue includes substantial material and energy costs. Public administration and education may create social benefits not recorded as incremental revenue, making private-sector ROI formulas incomplete.
6.6 Cost structure by organisation size
6.6.1 Microenterprise
A representative five-person professional business may purchase:
| Cost item | Annual scenario |
|---|---|
| Five AI subscriptions | USD 3,000 |
| Accounting, document and CRM AI features | USD 2,500 |
| API and automation | USD 1,500 |
| External integration | USD 4,000 |
| Data cleanup and templates | USD 2,000 |
| Cybersecurity and privacy controls | USD 2,500 |
| Training and governance | USD 2,000 |
| Human review and error correction | USD 2,500 |
| Contingency and switching | USD 2,000 |
| Total | USD 22,000 |
The main risk is not token consumption but fixed implementation overhead. At USD 0.8 million revenue, the expenditure equals 2.75% of turnover. If net operating margin before AI is 10%, annual operating profit is USD 80,000; AI consumes 27.5% of pre-AI profit. The relevant ratio is therefore:
AI Profit Burden = Annual AI Expenditure ÷ Pre-AI Operating Profit
AI Profit Burden = 22,000 ÷ 80,000 = 27.5%
A cost equal to only 2.75% of revenue can be strategically significant when margins are thin.
6.6.2 SME
SMEs face an intermediate problem. Their workflows are complex enough to require integration and governance, but their scale may be insufficient to spread those costs widely. The OECD reported that AI adoption among firms remains lower than adoption of other digital technologies and that significant gaps persist between SMEs and large firms. AI Adoption by Small and Medium-Sized Enterprises – OECD – December/2025 — OECD report on SME AI adoption.
A USD 5 million-revenue small enterprise spending USD 150,000 annually on AI must achieve one or a combination of:
- a 10% gross saving on its USD 1.5 million labour bill;
- a 3% reduction in total non-AI operating costs if those equal USD 5 million;
- USD 500,000 of incremental revenue at a 30% contribution margin;
- avoided annual losses of USD 150,000;
- a mixed benefit portfolio with equivalent value.
If the firm realises only half the expected productivity gain, AI destroys value unless adoption prevents a larger competitive loss.
6.6.3 Large corporation
Large firms possess several advantages:
- fixed governance costs are spread across more revenue;
- internal data and use cases are more abundant;
- enterprise bargaining power can reduce unit prices;
- specialist staff can be employed internally;
- workloads can be routed across models;
- larger usage can justify reserved capacity;
- benefits can be diversified across many business functions.
They also face disadvantages:
- integration with legacy systems;
- duplicated departmental pilots;
- larger cyberattack surfaces;
- complex regulatory exposure;
- slower change management;
- high cost of correlated errors;
- vendor concentration at enormous scale.
The large-firm advantage is therefore not simply “more money.” It is the capacity to convert fixed AI costs into a lower cost per employee, transaction and unit of revenue.
6.7 Sector-by-sector exposure
6.7.1 Banking
AI use is already widespread among significant European banks. The European Central Bank reported in June 2026 that more than 85% of banks under European banking supervision used AI. Strengthening Operational Resilience for the Age of AI – European Central Bank – June/2026 — ECB assessment of AI and banking resilience.
Banks can deploy AI in:
- fraud detection;
- transaction monitoring;
- credit analysis;
- customer service;
- coding;
- document processing;
- regulatory reporting;
- cybersecurity;
- risk modelling;
- collections;
- knowledge management.
The expected-value equation is:
Net AI Valuebank = Operating Savings + Fraud Losses Avoided + Credit Losses Avoided + Incremental Margin − Compliance Cost − Model Risk − AI Expenditure
Banking has high benefit potential but also high tail risk. An error in an internal drafting assistant is not equivalent to a systematic credit-scoring error. Data governance, explainability, validation, monitoring and third-party concentration must be allocated by use case. The ECB has warned that AI may improve operational efficiency while increasing operational risk and third-party dependence. The Rise of Artificial Intelligence: Benefits and Risks for Financial Stability – European Central Bank – May/2024 — ECB financial-stability analysis.
6.7.2 Insurance
Insurance AI can improve underwriting, claims triage, fraud detection, pricing and customer interaction. Its value equation is:
Net AI Valueinsurance = Expense Savings + Claims Leakage Reduction + Pricing Improvement + Incremental Premium Margin − Conduct Risk − Model Risk − AI Expenditure
The principal danger is that historical claims data encode past exclusions or biases. A small average improvement in loss ratio can be economically valuable, but a systematic error across thousands of policies can generate correlated liability.
6.7.3 Healthcare
Healthcare benefits include clinical-documentation support, scheduling, coding, imaging assistance, research and operational planning. Costs include validation, security, clinical supervision, integration with electronic health records and high-availability infrastructure.
Net AI Valuehealth = Administrative Savings + Capacity Value + Avoided Error + Health Outcome Value − Clinical Risk − Privacy Risk − AI Expenditure
A purely financial calculation may understate public benefit. Conversely, monetising every minute “saved” is invalid if clinicians use the time for additional verification or if staffing levels do not change. Productivity capacity becomes cash savings only when it reduces overtime, avoids hiring, increases reimbursable activity or improves outcomes sufficiently to produce measurable value.
6.7.4 Legal and professional services
Legal and advisory firms have high labour intensity and document-heavy workflows, giving them strong theoretical exposure to productivity gains. However, time saved can reduce billable hours under hourly pricing.
Net AI Valuelegal = Cost Savings + New Matters + Higher Realisation + Capacity Value − Lost Billable Hours − Liability − AI Expenditure
The business-model effect is decisive:
| Billing model | AI productivity effect |
|---|---|
| Hourly billing | Time savings can reduce revenue unless volume rises |
| Fixed fee | Efficiency increases margin |
| Subscription/retainer | Efficiency can increase capacity |
| Outcome fee | AI value depends on result quality |
| Internal legal department | Efficiency reduces cost or unmet demand |
The firm may need to shift pricing before productivity becomes financially valuable.
6.7.5 Manufacturing
Manufacturing use cases include predictive maintenance, quality inspection, process control, supply-chain planning, engineering support, procurement and technical documentation.
Net AI Valuemanufacturing = Downtime Avoided + Scrap Reduction + Yield Improvement + Labour Capacity + Inventory Reduction + Incremental Margin − Safety Risk − Integration Cost − AI Expenditure
The benefit should be calculated from physical variables:
Benefitscrap = Units × Scrap Reduction × Cost per Unit
Benefitdowntime = Hours Avoided × Contribution per Production Hour
Benefitinventory = Inventory Reduction × Carrying Cost Rate
Manufacturing AI can create value without reducing headcount. Yield, energy, scrap and uptime may dominate office-productivity savings.
6.7.6 Media
Media AI can reduce transcription, translation, editing, tagging, personalisation and production costs. It also creates copyright, provenance, trust and substitution risks.
Net AI Valuemedia = Production Savings + Audience Revenue + Archive Monetisation − Rights Risk − Brand Damage − Displacement Loss − AI Expenditure
The EU’s Article 50 transparency obligations became applicable from 2 August 2026 and address marking and detection of certain AI-generated content, including deepfakes and specified publications. Code of Practice on Transparency of AI-Generated Content – European Commission – July/2026 — European Commission AI-content transparency framework.
6.7.7 Telecommunications
Telecommunications operators can apply AI to network optimisation, predictive maintenance, customer service, fraud, churn, sales, cybersecurity and capacity planning. Their scale and recurring data flows make them strong candidates for positive AI economics. However, telecommunications systems are critical infrastructure, and dependence on external models can create availability and sovereignty risks.
Net AI Valuetelecom = Network Savings + Churn Reduction + Fraud Avoidance + Incremental ARPU + Service Automation − Outage Risk − Security Risk − AI Expenditure
A one-percentage-point reduction in churn can be worth more than large administrative savings; the correct unit of value may be retained customer lifetime value rather than labour hours.
6.7.8 Defence and critical infrastructure
Defence and critical infrastructure require higher assurance, isolated environments, sovereign control, testing, redundancy and supply-chain review. The financial return may appear negative because mission value is not fully recorded as revenue.
Mission-Adjusted AI Value = Financial Benefits + Resilience Value + Sovereignty Value + Capability Value − Mission Risk − AI Expenditure
A system that improves intelligence processing but creates an unacceptable adversarial-attack surface should not be approved merely because its expected financial return is positive.
6.7.9 Public administration
Public administration can use AI for document handling, translation, citizen support, fraud analysis and policy research. Benefits should include service quality and waiting-time reduction:
Public AI Value = Administrative Savings + Citizen Time Saved + Service Quality + Fraud Avoided − Rights Risk − Error Cost − AI Expenditure
Public bodies cannot assume that staff time saved becomes fiscal savings. If no positions, overtime or procurement costs are reduced, the benefit is improved capacity rather than cash.
6.7.10 Education and research
Education and research institutions face a dual burden: purchasing AI access while also redesigning assessment, research integrity, data governance and teaching.
Education AI Value = Teaching Capacity + Research Acceleration + Student Support + Accessibility − Integrity Risk − Inequality Cost − AI Expenditure
If affluent institutions purchase frontier models while poorer institutions rely on lower-capability systems, AI can amplify the research and education divide described in earlier chapters.
6.8 Regulatory and governance burden
The EU AI Act entered into force on 1 August 2024 and became generally applicable on 2 August 2026, subject to phased exceptions. Governance and general-purpose-model obligations became applicable earlier; certain high-risk provisions have extended timelines following the 2026 AI Omnibus. AI Act Regulatory Framework – European Commission – August/2026 — Official European Commission AI Act timeline.
Compliance costs depend on role and risk. A company can be:
- a provider;
- a deployer;
- an importer;
- a distributor;
- a product manufacturer;
- a downstream modifier.
The annual compliance burden can be represented as:
Ccompliance = Cinventory + Cclassification + Cdocumentation + Ctesting + Cmonitoring + Ctraining + Clegal + Cincident
Regulation is not the only driver. Even where a system is not legally classified as high-risk, contractual liability, professional standards, cybersecurity and customer expectations can require comparable controls.
NIST’s Generative AI Profile treats risk management as a lifecycle covering governance, mapping, measurement and management. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile – NIST – July/2024 — NIST Generative AI Profile. This supports a governance model in which evaluation and monitoring are recurring operating expenses rather than one-time launch tasks.
6.9 Deployment-model comparison
6.9.1 Structural comparison
| Deployment model | Initial CAPEX | Variable cost | Privacy control | Frontier capability | Switching cost | Internal skill requirement |
|---|---|---|---|---|---|---|
| Full cloud proprietary | Low | High/usage-based | Medium | Very high | High | Medium |
| Fully local proprietary | Very high | Moderate | High | Limited by available hardware/licence | High | Very high |
| Fully local open-weight | High | Moderate | Very high | Model-dependent | Medium | Very high |
| Hybrid | Medium/high | Medium | High for sensitive workloads | High | Medium | High |
| Specialised small model | Low/medium | Low | High if local | High for bounded task | Low/medium | Medium/high |
| Retrieval-augmented proprietary | Medium | Medium/high | Medium | High | Medium/high | High |
| Retrieval-augmented open-weight | Medium/high | Low/moderate | High | Model-dependent | Low/medium | High |
| Shared/cooperative compute | Shared CAPEX | Shared | Potentially high | Medium/high | Governance-dependent | Shared specialist team |
6.9.2 Financial scoring
Scores range from 1, least favourable, to 5, most favourable.
| Model | Cost predictability | Low entry cost | Privacy | Scalability | Portability | High-assurance suitability |
|---|---|---|---|---|---|---|
| Cloud proprietary | 3 | 5 | 2 | 5 | 2 | 3 |
| Fully local | 4 | 1 | 5 | 2 | 3 | 5 |
| Hybrid | 4 | 3 | 4 | 4 | 4 | 5 |
| Open-weight local | 4 | 2 | 5 | 3 | 5 | 4 |
| Specialised small model | 5 | 4 | 4 | 3 | 4 | 4 |
| RAG with frontier API | 3 | 4 | 3 | 5 | 3 | 3 |
| Cooperative compute | 4 | 3 | 4 | 3 | 4 | 4 |
No model dominates every dimension. Hybrid deployment is frequently economically attractive because it routes sensitive or predictable workloads locally and sends high-complexity or burst demand to the cloud.
6.10 Cloud, local and hybrid break-even
Let:
Ccloud(Q) = Fcloud + pcloudQ
Clocal(Q) = Flocal + plocalQ
Cloud and local costs are equal at:
Q* = [Flocal − Fcloud] ÷ [pcloud − plocal]
Assume:
- cloud fixed integration cost = USD 100,000;
- local fixed annualised cost = USD 700,000;
- cloud variable cost = USD 12 per million successful tokens;
- local variable cost = USD 3 per million successful tokens.
Then:
Q = (700,000 − 100,000) ÷ (12 − 3)
Q = 66,667 million tokens
Q* ≈ 66.7 billion successful tokens annually
Below approximately 66.7 billion successful tokens, the cloud scenario is cheaper under these assumptions. Above it, local infrastructure becomes cheaper if quality, utilisation and reliability are equivalent. They frequently are not equivalent, so a quality-adjusted threshold is needed:
Q*QA = [Flocal − Fcloud] ÷ [pcloudycloud − plocalylocal]
where y represents the cost adjustment required to obtain a successful workload.
6.11 Mandatory competitive adoption
6.11.1 The adoption game
Consider two competing firms. Each can adopt or not adopt AI.
| Firm A / Firm B | B adopts | B does not adopt |
|---|---|---|
| A adopts | Both incur AI costs; relative advantage may disappear | A may gain speed, cost or product advantage |
| A does not adopt | A risks losing share or talent | Neither pays; existing equilibrium continues |
If both firms adopt and obtain identical productivity gains that competition passes to customers through lower prices, neither may retain the benefit as profit. AI then becomes a Red Queen investment: each firm must invest merely to maintain its relative position.
The private decision rule becomes:
Adopt if:
Net AI Value + Competitive Loss Avoided > 0
Define:
Dcompetitive = Revenue or margin lost if the company does not adopt
Then:
Strategic Net AI Value = Net AI Value + Dcompetitive
A project with direct Net AI Value of −USD 2 million can be rational if non-adoption is expected to destroy USD 5 million of contribution margin. Its strategic value is positive USD 3 million.
6.11.2 Mandatory-AI tax
The de facto competitive burden can be expressed as:
Mandatory AI Tax = Minimum AI Expenditure Required to Maintain Competitive Parity ÷ Revenue
| Company type | Illustrative minimum parity expenditure | Revenue | Mandatory AI Tax |
|---|---|---|---|
| Micro professional firm | USD 10,000 | USD 0.8m | 1.25% |
| Small enterprise | USD 75,000 | USD 5m | 1.50% |
| Medium enterprise | USD 300,000 | USD 30m | 1.00% |
| Large enterprise | USD 25m | USD 5bn | 0.50% |
| Bank | USD 80m | USD 10bn | 0.80% |
| Media company | USD 3m | USD 200m | 1.50% |
The absolute burden is larger for large firms, but the proportional burden can be greater for small firms because governance, integration and security contain indivisible fixed components.
6.12 Value-destruction thresholds
6.12.1 Minimum labour-productivity gain
Let:
- A = annual AI expenditure;
- RC = annual risk cost;
- L = annual labour cost;
- ΔR = incremental revenue;
- m = contribution margin.
Break-even requires:
gLL + mΔR ≥ A + RC
The minimum required labour-productivity gain is:
gL* = [A + RC − mΔR] ÷ L
Archetype thresholds
| Archetype | AI expenditure plus risk cost | Incremental contribution | Labour cost | Minimum labour productivity gain |
|---|---|---|---|---|
| Micro firm | USD 26k | USD 8k | USD 320k | 5.63% |
| Small enterprise | USD 180k | USD 60k | USD 1.5m | 8.00% |
| Medium enterprise | USD 700k | USD 240k | USD 7m | 6.57% |
| Large company | USD 80m | USD 40m | USD 1.2bn | 3.33% |
| Bank | USD 250m | USD 90m | USD 3bn | 5.33% |
| Healthcare organisation | USD 90m | USD 10m | USD 1.1bn | 7.27% |
| Legal firm | USD 5m | USD 1.5m | USD 55m | 6.36% |
| Manufacturer | USD 24m | USD 12m | USD 180m | 6.67% |
| Telecommunications operator | USD 125m | USD 80m | USD 900m | 5.00% |
| Defence operator | USD 110m | USD 30m | USD 700m | 11.43% |
| Public administration | USD 90m | USD 0m | USD 900m | 10.00% |
| University system | USD 55m | USD 5m | USD 600m | 8.33% |
A “productivity gain” must become economically usable. If employees save 8% of their time but workload, staffing, revenue and service output remain unchanged, the realised financial gain may be close to zero.
6.12.2 Minimum incremental revenue
If no labour saving occurs:
ΔR* = [A + RC − S] ÷ m
where S is other verified cost savings.
For a small enterprise with USD 150,000 AI expenditure, USD 30,000 risk cost, USD 50,000 other savings and a 30% contribution margin:
ΔR = (150,000 + 30,000 − 50,000) ÷ 0.30
ΔR = USD 433,333
The company must generate more than USD 433,333 of additional annual revenue merely to break even.
6.12.3 Minimum operating margin
If annual AI expenditure is a proportion a of revenue and pre-AI operating margin is m0:
Post-AI margin before benefits = m0 − a
A company with a 4% operating margin and AI costs equal to 3% of revenue loses 75% of operating profit before benefits.
AI Profit Burden = a ÷ m0
| Pre-AI operating margin | AI Cost Ratio | Share of operating profit consumed |
|---|---|---|
| 3% | 1% | 33.3% |
| 3% | 2% | 66.7% |
| 3% | 3% | 100.0% |
| 5% | 2% | 40.0% |
| 5% | 3% | 60.0% |
| 10% | 2% | 20.0% |
| 20% | 2% | 10.0% |
| 30% | 2% | 6.7% |
Low-margin industries are therefore highly exposed even when AI costs appear modest relative to revenue.
6.13 Five-year NPV scenario
Consider a medium enterprise with:
- initial integration and data-preparation cost: USD 1.2 million;
- recurring year-one cost: USD 600,000;
- recurring cost growth: 8% annually;
- year-one benefit: USD 700,000;
- benefit growth: 15% annually;
- discount rate: 9%;
- terminal value excluded;
- expected error and downtime costs included in recurring cost.
| Year | Cost | Benefit | Net cash flow |
|---|---|---|---|
| 0 | USD 1.200m | USD 0 | −USD 1.200m |
| 1 | USD 0.600m | USD 0.700m | USD 0.100m |
| 2 | USD 0.648m | USD 0.805m | USD 0.157m |
| 3 | USD 0.700m | USD 0.926m | USD 0.226m |
| 4 | USD 0.756m | USD 1.065m | USD 0.309m |
| 5 | USD 0.816m | USD 1.225m | USD 0.409m |
Approximate discounted values:
| Measure | Value |
|---|---|
| NPV of benefits | USD 3.55m |
| NPV of recurring costs | USD 2.69m |
| Initial cost | USD 1.20m |
| Total NPV of costs | USD 3.89m |
| Net project NPV | −USD 0.34m |
| ROIAI | −8.7% |
Although annual net cash flow is positive from year one, the project fails to recover its initial integration cost within the five-year discounted horizon. This illustrates why pilot-level operating savings do not prove enterprise value.
6.14 Sensitivity analysis
Using the medium-enterprise baseline, the five-year result is most sensitive to adoption, benefit realisation and risk.
| Benefit realisation | Cost overrun | Risk-cost multiplier | Likely financial result |
|---|---|---|---|
| 50% | 25% | 2.0× | Severe value destruction |
| 75% | 15% | 1.5× | Negative |
| 100% | 0% | 1.0× | Near break-even |
| 125% | 0% | 1.0× | Positive |
| 125% | −10% | 0.75× | Strongly positive |
| 150% | −15% | 0.50× | Transformational |
The realised benefit ratio is:
BRR = Realised Benefits ÷ Forecast Benefits
The project should trigger remediation when:
BRR < BRRminimum
If forecast benefits are USD 1 million and break-even requires USD 750,000, then:
BRRminimum = 750,000 ÷ 1,000,000 = 75%
6.15 Open-weight, small-model and cooperative alternatives
6.15.1 Open-weight systems
Open-weight models reduce dependence on a single inference provider and allow local inspection, fine-tuning and deployment. They do not eliminate cost. The enterprise assumes responsibility for:
- infrastructure;
- optimisation;
- security;
- model evaluation;
- updates;
- monitoring;
- licensing analysis;
- incident response;
- specialist staffing.
The appropriate comparison is:
TCOopen versus TCOproprietary
not “free model” versus “paid API.”
6.15.2 Small-model specialisation
A smaller model can create superior enterprise value when:
Qualitysmall,task ≥ Qualityrequired
and:
TCOsmall < TCOfrontier
Task-specific evaluation may reveal that a small model with retrieval outperforms a general frontier model on controlled extraction, classification or internal terminology while costing less and exposing less data.
6.15.3 Retrieval-augmented generation
RAG is economically attractive when the value of improved grounding and shorter model context exceeds indexing and retrieval costs:
ValueRAG = Error Reduction + Prompt Savings + Update Flexibility − Retrieval Cost − Integration Cost
RAG does not solve poor data governance. It can retrieve incorrect, obsolete or unauthorised content with high confidence.
Universities, hospitals, municipalities, professional associations and SMEs can pool infrastructure:
Cost per Member = [Shared Fixed Cost + Shared Operating Cost] ÷ Members + Member-Specific Cost
Cooperation can reduce fixed-cost inequality but introduces governance questions:
- capacity allocation;
- confidentiality;
- liability;
- model selection;
- maintenance responsibility;
- admission and exit;
- priority during scarcity;
- cost distribution;
- data separation.
A cooperative model is particularly attractive where members share assurance requirements but do not individually possess sufficient usage to justify dedicated infrastructure.
6.16 Enterprise decision framework
An AI investment should proceed only after answering:
| Test | Required evidence |
|---|---|
| Strategic necessity | Quantified competitive loss from non-adoption |
| Task suitability | Baseline and model performance on representative work |
| Financial case | Five-year NPV, ROI and payback |
| Benefit convertibility | Mechanism translating time savings into cash or output |
| Risk tolerance | Expected and tail-loss assessment |
| Data readiness | Rights, quality, lineage and access |
| Deployment choice | Cloud, local, hybrid or cooperative comparison |
| Exit feasibility | Switching cost and portable data/model layer |
| Governance | Named owner, inventory, monitoring and escalation |
| Human control | Review threshold and accountable decision-maker |
| Resilience | Fallback, outage process and recovery target |
| Measurement | Post-deployment control group or credible counterfactual |
The company should establish a counterfactual:
Incremental AI Value = OutcomeAI − Outcomeno-AI
Without a control group, phased rollout, matched workflow or credible baseline, ordinary business improvement may be misattributed to AI.
6.17 Five-year enterprise outlook, 2027–2031
The most probable outcome is not universal replacement of labour but the progressive conversion of AI into a mandatory layer of enterprise infrastructure. Adoption costs will move from visible subscriptions toward less visible integration, governance, evaluation and risk-management expenditure. Large firms will increasingly build model-routing platforms combining proprietary frontier models, specialised small models, retrieval systems and local infrastructure. Smaller firms will remain more dependent on bundled SaaS products because they cannot efficiently internalise governance and engineering. This will transfer bargaining power to software vendors and create a persistent risk that AI becomes an unavoidable surcharge embedded across accounting, productivity, CRM, design, legal, security and communications products.
Three enterprise equilibria are plausible:
| Equilibrium | Description | Distributional consequence |
|---|---|---|
| Productivity diffusion | Benefits exceed costs across firm sizes | Broad increase in output and lower prices |
| Competitive compulsion | Firms adopt to maintain parity; gains pass to customers | AI becomes mandatory overhead |
| Capability concentration | Large firms capture scale economies and better models | SME margins and market share decline |
The second and third outcomes can coexist. Customers may receive faster or cheaper services while enterprise profits become more concentrated among AI infrastructure and platform providers.
6.18 Final findings
- AI adoption is accelerating, but use rates do not establish profitability. EU enterprise adoption rose to approximately 20% in 2025, while large-enterprise use exceeded 55%.
- The relevant enterprise cost is substantially larger than licence and API expenditure. Data, integration, cybersecurity, governance, evaluation, human supervision and switching must be included.
- Small firms face a disproportionate fixed-cost burden. Identical governance work consumes a much larger share of small-firm revenue and profit.
- Low-margin companies are exposed even at apparently small AI Cost Ratios. AI expenditure equal to 3% of revenue eliminates pre-AI operating profit when the margin is also 3%.
- Time saved is not automatically financial value. It becomes cash benefit only through lower cost, avoided hiring, increased output, higher revenue or improved public service.
- High-risk sectors require tail-risk analysis. Expected productivity gains cannot justify deployment when a low-probability failure threatens catastrophic loss.
- Cloud deployment minimises initial capital but can maximise dependency. Fully local deployment improves control but requires high utilisation and specialist staff.
- Hybrid systems are likely to dominate mature enterprise architecture. Sensitive and predictable workloads can remain local while frontier and burst demand use external capacity.
- Open-weight models eliminate neither infrastructure nor governance costs. Their economic advantage is control and portability, not zero TCO.
- Small specialised models can outperform frontier systems economically on bounded tasks. The relevant criterion is successful workload cost, not nominal parameter count.
- AI may become a Red Queen investment. Companies may pay merely to preserve competitive parity, with productivity gains passed to customers rather than retained as profit.
- The minimum productivity threshold can be calculated explicitly. It rises with AI expenditure, risk costs and low contribution margins and falls with verifiable revenue or non-labour benefits.
- Positive annual savings do not guarantee positive five-year NPV. Initial integration and recurring cost escalation can leave a project value-destructive.
- Defence, healthcare, government and education require mission-adjusted value. Commercial ROI alone omits resilience, sovereignty, safety and public benefit.
- The decisive divide from 2027 to 2031 will be organisational capability, not simple access to a chatbot. Firms able to govern data, evaluate models, route workloads and negotiate infrastructure will capture disproportionate value; firms purchasing fragmented AI tools without measurement may incur a new mandatory cost without receiving a durable productivity return.
CHAPTER 7 — The AI Capability Divide: Students, Researchers and Social Mobility
7.1 Scope and central research question
The emerging AI divide is not adequately described by whether an individual has ever used a chatbot. Meaningful access has at least six dimensions: availability, affordability, model capability, usage allowance, hardware and connectivity, and the user’s ability to integrate AI into productive work. A nominally free service may impose restrictive rate limits, use a lower-capability model, lack long-context document analysis, provide no persistent project environment, withhold advanced tools or become unavailable during periods of congestion. Paid consumer access can remove some constraints while remaining materially below enterprise access in throughput, governance, data integration and guaranteed availability. Local high-performance access can provide privacy and experimentation but requires substantial capital, electricity, technical competence and maintenance. The correct analytical variable is therefore quality-adjusted effective access, not account ownership.
This distinction matters because AI can augment activities that accumulate human and institutional capital: studying, coding, translating, publishing, analysing documents, preparing applications, designing businesses, conducting experiments and searching scientific literature. Small initial differences in access may compound. A student who receives higher-quality explanations, faster feedback and stronger coding assistance can complete more projects; those projects can improve employability; higher income can finance better AI access; better access can then produce further advantage. The opposite sequence can affect students and researchers confined to outdated hardware, unstable free tiers or low-capability local models.
Adoption is already strongly stratified. In 2025, 20.0% of EU enterprises with at least ten employees reported using AI, but the rate among large enterprises reached 55.03%. Use of Artificial Intelligence in Enterprises – Eurostat – December/2025 — Eurostat enterprise AI statistics. Although enterprise statistics do not directly measure student access, they reveal the institutional environments into which graduates enter: researchers and workers associated with wealthy organisations can obtain better models, data, tools and compute than equally capable individuals outside them.
7.2 From digital access to capability access
7.2.1 Four access groups
| Group | Typical access | Effective capability | Principal limitation |
|---|---|---|---|
| G0 — No regular AI access | No stable account, device, connectivity or permission | None or occasional indirect assistance | Complete exclusion |
| G1 — Free-tier access | Free consumer model on a phone or ordinary computer | Basic explanation, drafting, translation and limited coding | Rate limits, model routing, weak continuity and uncertain availability |
| G2 — Paid consumer access | Monthly subscription and adequate general-purpose device | Stronger reasoning, files, longer context and more tools | Fair-use limits, no guaranteed capacity and limited institutional integration |
| G3 — Frontier enterprise or high-performance local access | Frontier APIs, enterprise plans, specialised agents, proprietary data or high-memory local hardware | Sustained research, coding, long-document processing, automation and confidential workflows | High monetary, technical and organisational cost |
The groups should not be treated as permanent identities. A person may move between them by task: free consumer access for general questions, university enterprise access for research and no permissible AI access for confidential clinical data. Measurement should therefore record the percentage of relevant study or work time during which each access class is available.
7.2.2 Quality-adjusted AI access
Define quality-adjusted access as:
QAAu,t = Aavailability × Qtask × Uallowance × Ttools × Ddata × Sskills
where:
- Aavailability is the probability that the system is accessible when required;
- Qtask is task-specific quality;
- Uallowance is usable volume relative to demand;
- Ttools represents access to retrieval, code execution, file analysis and external tools;
- Ddata represents lawful access to relevant data;
- Sskills represents user competence.
Because the variables have multiplicative interaction, one near-zero component can neutralise the others. A highly capable model provides little effective access if the user can submit only a few requests, cannot upload required documents, lacks connectivity or cannot evaluate unreliable output.
7.3 Dimensions of the AI capability divide
7.3.1 Educational attainment
AI can affect educational attainment through:
- personalised explanations;
- repeated feedback without additional tutoring fees;
- adaptive practice;
- translation;
- accessibility assistance;
- coding and mathematical support;
- summarisation and note generation;
- exam preparation;
- administrative guidance;
- application and scholarship support.
The causal estimand is not the difference between AI users and non-users, because users may differ in prior attainment, income, motivation and institutional support. A valid model is:
Yi,t = α + βQAAi,t + γYi,t−1 + δXi,t + μinstitution + τt + εi,t
where:
- Yi,t is attainment;
- QAAi,t is quality-adjusted AI access;
- Xi,t contains income, prior grades, language, disability, device and connectivity controls;
- μinstitution captures institutional differences;
- τt captures period effects.
A positive β would support the access-advantage hypothesis only if AI access precedes the measured outcome and confounding is credibly controlled.
7.3.2 Research productivity
Research benefits may include:
- literature discovery;
- document classification;
- code generation;
- data cleaning;
- translation;
- hypothesis generation;
- simulation support;
- manuscript revision;
- grant preparation;
- administrative automation.
Research productivity should be measured through multiple outcomes:
RPi,t = w1Publications + w2Citations + w3Datasets + w4Software + w5Replication + w6ResearchQuality
Raw publication count is insufficient because AI can increase text production without improving scientific validity. Retractions, irreproducible analyses, fabricated references and homogeneous research agendas are negative outputs, not productivity.
7.3.3 Employability
AI affects employability through two opposing mechanisms. It can increase a worker’s ability to code, communicate, analyse and automate. It can also reduce the market value of routine skills. The employability advantage is therefore:
EAi = Productivity Complementarity + AI Literacy + Portfolio Quality − Skill Substitution − Verification Failure
A worker becomes more employable when AI complements scarce domain expertise. A worker whose output is easily substituted may face greater competition even after acquiring basic AI skills.
7.3.4 Entrepreneurship
Entrepreneurial opportunity can expand because AI reduces the cost of:
- market research;
- translation;
- software prototyping;
- design;
- customer service;
- legal-document preparation;
- marketing;
- data analysis;
- business planning.
However, frontier access can create a new minimum efficient scale. If competitors use integrated agents, proprietary data and reserved inference capacity, free-tier users may be able to create prototypes but not reliable production systems.
7.3.5 Software development
Coding is among the most measurable domains because performance can be evaluated through tests, issue resolution, defect rate and accepted code. Appropriate outcomes include:
- time to accepted pull request;
- percentage of issues resolved;
- unit-test pass rate;
- security defects;
- maintainability;
- human review time;
- cost per successfully completed issue.
Anthropic reported a 72.5% result for Claude Opus 4 on SWE-bench under its stated evaluation configuration. Introducing Claude 4 – Anthropic – May/2025 — Anthropic Claude 4 evaluation disclosure. This is evidence that frontier systems can resolve a substantial share of benchmarked software-engineering tasks under specified conditions. It is not proof that every enterprise codebase will obtain the same result.
7.3.6 Language access
Multilingual AI can reduce the penalty borne by people working outside dominant scientific and commercial languages. It can translate instructions, research, technical documentation and professional correspondence. Capability remains uneven across languages, dialects and specialist terminology.
The language-access advantage can be represented as:
LAl = Qtranslation,l × Qreasoning,l × Coveragel − ErrorCostl
A system that translates fluently but reasons less reliably in the target language may create false inclusion. Anthropic’s Sonnet 4.6 system card reports evaluation across 57 academic subjects and 14 non-English languages and publishes the model’s MMMLU methodology and results. Claude Sonnet 4.6 System Card – Anthropic – February/2026 — Anthropic Claude Sonnet 4.6 System Card. Such results remain provider-reported and should be independently reproduced before being used for policy conclusions.
7.4 AI Affordability Index
7.4.1 Definition
AAIc,t = Annual Cost of Adequate AI Accessc,t ÷ Median Disposable Incomec,t
The numerator must represent adequate—not merely nominal—access. It should include:
Costadequate = Subscription + API + DeviceAnnualisation + Connectivity + Storage + Software + Maintenance + Tax
An affordability index based solely on a USD 20 monthly subscription understates the cost for users who need a capable computer, broadband, storage or additional API capacity.
7.4.2 Proposed interpretation
| AAI value | Interpretation |
|---|---|
| Below 1% | Broadly affordable for median-income users |
| 1–3% | Noticeable but generally manageable |
| 3–5% | Material household decision |
| 5–10% | Restrictive |
| 10–20% | Severely restrictive |
| Above 20% | Economically exclusionary without subsidy |
These thresholds are analytical classifications, not international standards.
7.4.3 Adequate-access baskets
| Basket | Components | Annual scenario cost |
|---|---|---|
| B0 — Nominal free access | Existing device, free tier, existing connectivity | USD 0 incremental, but no adequacy guarantee |
| B1 — Basic paid access | Consumer subscription, ordinary device annualisation, incremental connectivity | USD 660 |
| B2 — Professional frontier access | Higher usage, frontier subscription/API, storage and capable device | USD 2,400 |
| B3 — Advanced research access | High API use, tools, storage and workstation annualisation | USD 4,000 |
| B4 — High-performance private access | Local workstation/server, maintenance, electricity and model operations | USD 8,000–25,000 |
The baskets are standardised analytical scenarios. Taxes, exchange rates and local hardware prices must be applied separately.
7.4.4 Income-level affordability scenarios
| Median annual disposable resources | B1: USD 660 | B2: USD 2,400 | B3: USD 4,000 | B4 low: USD 8,000 |
|---|---|---|---|---|
| USD 3,000 | 22.0% | 80.0% | 133.3% | 266.7% |
| USD 6,000 | 11.0% | 40.0% | 66.7% | 133.3% |
| USD 12,000 | 5.5% | 20.0% | 33.3% | 66.7% |
| USD 24,000 | 2.75% | 10.0% | 16.7% | 33.3% |
| USD 45,000 | 1.47% | 5.33% | 8.89% | 17.78% |
| USD 75,000 | 0.88% | 3.20% | 5.33% | 10.67% |
Eurostat reported median annual disposable income of 21,245 purchasing-power standards per EU inhabitant in 2024. Living Conditions in Europe: Income Distribution and Income Inequality – Eurostat – 2026 — Eurostat median disposable-income statistics. At an illustrative income of 21,245 units, the standardised B1, B2 and B3 baskets would absorb approximately 3.1%, 11.3% and 18.8% respectively if the basket units were price-level comparable. A rigorous country table must convert the basket into local prices rather than equating one purchasing-power standard mechanically with one US dollar.
7.5 Student AI Affordability Index
7.5.1 Definition
SAIc,t = Annual AI and Hardware Costc,t ÷ Annual Student Disposable Resourcesc,t
Student resources should include income actually available after tuition and essential housing:
Student Resources = Grants + Scholarships + Family Support + Net Employment Income − Tuition − Essential Housing − Essential Living Costs
Using household median income as the denominator can severely understate student hardship.
7.5.2 Student affordability scenarios
| Annual disposable student resources | Basic paid access: USD 660 | Professional access: USD 2,400 | Advanced research: USD 4,000 |
|---|---|---|---|
| USD 1,500 | 44.0% | 160.0% | 266.7% |
| USD 3,000 | 22.0% | 80.0% | 133.3% |
| USD 6,000 | 11.0% | 40.0% | 66.7% |
| USD 12,000 | 5.5% | 20.0% | 33.3% |
| USD 24,000 | 2.75% | 10.0% | 16.7% |
This explains why a service that appears inexpensive to a professional can be exclusionary for a student. A USD 20 monthly subscription equals USD 240 annually before hardware and connectivity. For a student with only USD 1,500 of genuinely disposable resources, the subscription alone consumes 16%.
7.6 Workstation affordability
7.6.1 Standardised workstation
For cross-country modelling, define a reference AI workstation as:
- modern multi-core CPU;
- 128 GB system memory;
- 24–32 GB accelerator memory;
- 2–4 TB solid-state storage;
- adequate power supply and cooling;
- five-year economic life.
The reference planning price is set at:
Preference = USD 4,000 before local tax, tariff and currency adjustment
This is an analytical basket rather than a named retail product. As an official hardware anchor, NVIDIA lists the GeForce RTX 5090 with 32 GB GDDR7 memory and a starting price of USD 1,999 before the remainder of the workstation. GeForce RTX 5090 Specifications – NVIDIA – January/2025 — NVIDIA GeForce RTX 5090 official specifications.
7.6.2 Wage-month formula
Wage Monthsc = Local Workstation Pricec ÷ Median Monthly Wagec
The ratio must use local after-tax wages for an affordability study. Gross earnings can be reported as a secondary indicator but should not be confused with disposable income.
7.6.3 Verified and indicative calculations
| Country | Official earnings input | Workstation basket | Wage months | Evidentiary qualification |
|---|---|---|---|---|
| United States | Median USD 1,251/week; approximately USD 5,421/month | USD 4,000 | 0.74 | Median gross full-time earnings, Q2 2026 |
| Germany | Implied median EUR 21.48/hour from official low-wage threshold; approximately EUR 3,723/month at 173.3 hours | EUR 4,000 | 1.07 | Estimated monthly median from official hourly median |
| Italy | Average hourly earnings approximately EUR 16.35 in the reported 2022 structure survey; approximately EUR 2,834/month at 173.3 hours | EUR 4,000 | 1.41 | Average, not median; indicative only |
| India | Average employee earnings INR 22,220/month in 2025 | INR 350,000 analytical local basket | 15.75 | Average, not median; basket requires local price verification |
| South Africa | Average employee earnings ZAR 7,980/month in the latest ILO profile observation, dated 2020 | ZAR 80,000 analytical basket | 10.03 | Stale average; unsuitable for current ranking |
US median full-time earnings were USD 1,251 per week in Q2 2026. Usual Weekly Earnings of Wage and Salary Workers – U.S. Bureau of Labor Statistics – July/2026 — BLS Q2 2026 median earnings. Germany’s Federal Statistical Office reported a median hourly earnings benchmark of EUR 21.48 for April 2025 through its official low-wage-threshold table. Low-Wage Thresholds in Germany – Destatis – May/2025 — Destatis German wage thresholds. ILOSTAT reports average monthly employee earnings of INR 22,220 for India in 2025. India Country Profile – ILOSTAT – 2026 — ILOSTAT India labour profile.
The table must not be interpreted as a definitive cross-country league table because wage concepts, dates, taxes and workstation prices are not perfectly harmonised. Its robust finding is ordinal: hardware representing less than one month of full-time US median earnings can represent many months of employee income in lower-income markets.
7.7 Measuring the capability difference between access groups
7.7.1 Outcome matrix
| Capability | G0: no access | G1: free tier | G2: paid consumer | G3: frontier/enterprise/local high-performance |
|---|---|---|---|---|
| Basic explanation | None | Moderate | High | High |
| Complex reasoning | None | Limited/variable | High | Highest available |
| Coding | None | Basic to moderate | Strong | Strong with tools and sustained workloads |
| Factual research | Manual only | Variable | Better tools and context | Retrieval, APIs, databases and auditability |
| Long documents | Manual | Restricted | Moderate/high | High throughput and persistent projects |
| Multilingual support | Manual tools | Broad but variable | Higher quality | Domain-integrated and scalable |
| Scientific workflow | Manual | Limited | Moderate | Code, tools, private data and automation |
| Confidential work | No AI | Often unsuitable | Contract-dependent | Enterprise controls or local privacy |
| Concurrency | None | Low | Moderate | High |
| Reproducibility | Manual | Weak | Moderate | API snapshots, logging and controlled deployment |
| Adaptation | None | Minimal | User-level | RAG, fine-tuning, agents and local models |
The table expresses expected structural access rather than fixed model performance. Free tiers can temporarily expose frontier models, while paid users may select weaker models. The classification must therefore be recorded at the level of the actual model, service tier and task.
7.8 Reproducible capability benchmark
7.8.1 Test principle
Parameter count must not be used as a proxy for utility. Performance depends on:
- architecture;
- total and active parameters;
- training data;
- post-training;
- quantisation;
- context;
- retrieval;
- tools;
- inference framework;
- prompt format;
- sampling;
- language;
- task specialisation;
- compute budget.
The benchmark unit should be:
Successful Workloads ÷ Total Cost
rather than raw benchmark score alone.
7.8.2 Required benchmark suite
| Dimension | Test construction | Primary metric | Failure metric |
|---|---|---|---|
| Reasoning | Novel multi-step problems with contamination controls | Exact success rate | Unsupported intermediate claims |
| Coding | Repository issues with tests | Resolved issue rate | New defects and insecure code |
| Factual reliability | Time-stamped questions with verified corpus | Supported-claim precision | Hallucination rate |
| Multilingual | Parallel domain tasks across languages | Quality parity ratio | Performance degradation by language |
| Long context | Multiple relevant and adversarial passages | Retrieval and synthesis accuracy | Context-position failure |
| Document processing | Tables, scans, footnotes and conflicting passages | Extraction F1 and citation accuracy | Invented fields |
| Scientific tasks | Reproducible calculations and literature analysis | Correct result and provenance | Fabricated references |
| Energy | Wall-socket or facility energy | Joules per successful task | Energy per failed output |
| Cost | Full provider or local TCO | Cost per successful task | Cost per generated token without quality |
| Privacy | Data-flow and retention audit | Compliance rate | Unauthorised transmission |
7.8.3 Controlled experimental configuration
Every run must disclose:
- exact model and version;
- quantisation;
- hardware;
- inference framework;
- prompt;
- temperature;
- reasoning effort;
- tool access;
- retrieval corpus;
- context length;
- maximum output;
- number of trials;
- confidence interval;
- date;
- total cost;
- energy measurement.
OpenAI’s official model guidance recommends comparing configurations on representative tasks instead of assuming that the highest reasoning effort always provides the best trade-off. Model Guidance – OpenAI API – August/2026 — Official OpenAI model-evaluation guidance.
7.8.4 Statistical analysis
For model m and task k:
Successm,k ∈ {0,1}
Mean success is:
p̂m = ΣSuccessm,k ÷ n
The quality-adjusted cost is:
QACm = Total Costm ÷ ΣSuccessm,k
Energy per successful task is:
Esuccess,m = Total Energym ÷ ΣSuccessm,k
Hallucination rate is:
HRm = Unsupported Factual Claimsm ÷ All Factual Claimsm
Multilingual parity for language l is:
MPm,l = Scorem,l ÷ Scorem,reference
Values below one indicate lower performance than the reference language.
7.9 Do smaller local models create a measurable disadvantage?
The hypothesis is:
H7.1: After controlling for tools, task specialisation, quantisation, context and cost, smaller local models produce a lower successful-workload rate on complex general tasks than frontier cloud models.
The counter-hypothesis is:
H7.2: On bounded and specialised tasks, a smaller local model combined with retrieval can equal or exceed a general frontier model at lower cost and higher privacy.
Both can be true.
7.9.1 Expected task pattern
| Task | Likely frontier advantage | Potential small-model advantage |
|---|---|---|
| Novel multi-step reasoning | High | Limited |
| Complex repository coding | High | Specialised code model may compete |
| Routine classification | Low | Strong cost and latency advantage |
| Structured extraction | Low/moderate | Strong with schema and validation |
| Multilingual specialist work | Language-dependent | Strong only after targeted adaptation |
| Long-document synthesis | High without retrieval | RAG can reduce gap |
| Private document processing | Cloud quality may be higher | Local privacy and control |
| Scientific reasoning | High | Domain-tuned model may compete narrowly |
| High-concurrency automation | Infrastructure-dependent | Smaller models may scale more economically |
Google’s Gemini 3.1 Pro model card reports evaluation across reasoning, multimodal, multilingual, agentic and long-context tasks. Gemini 3.1 Pro Model Card – Google DeepMind – February/2026 — Google DeepMind Gemini 3.1 Pro Model Card. Such model cards demonstrate that capability is multidimensional. They do not provide a substitute for an independently controlled local-versus-cloud experiment.
7.9.2 Quantisation controls
A valid local comparison should test at least:
- BF16 or FP16;
- INT8;
- INT4;
- selected lower-bit method if documented.
The quantisation penalty is:
QPm,q,k = Scorem,reference,k − Scorem,q,k
A lower-bit model should be considered economically superior only when:
QACm,q < QACm,reference
and:
Scorem,q,k ≥ Minimum Acceptable Scorek
A model that becomes cheaper but falls below the task’s minimum quality threshold is not an adequate substitute.
7.10 Scientific-publication and research divide
Research institutions with enterprise AI access can provide:
- bulk document ingestion;
- licensed databases;
- secure code execution;
- private datasets;
- long-context models;
- shared prompt libraries;
- evaluation support;
- dedicated local compute;
- legal and research-integrity guidance.
An unaffiliated researcher may have only a consumer subscription and personal computer. The difference is not merely model quality. It includes data rights, concurrency, reproducibility, storage, collaboration and the ability to process sensitive material.
Define the Research AI Resource Index:
RARIi = w1Compute + w2ModelQuality + w3DataAccess + w4ToolAccess + w5Support + w6UsageAllowance
A publication model can be estimated as:
Publicationsi,t+1 = α + βRARIi,t + γPriorPublicationsi,t + δFundingi,t + μfield + εi,t
Citation and quality measures should be analysed separately to detect whether AI increases quantity without validity.
7.11 Cumulative advantage
7.11.1 Requested model
At+1 = At + αQt − βCt
where:
- At = accumulated human or institutional advantage;
- Qt = quality-adjusted AI access;
- Ct = access and adaptation constraints;
- α = conversion of AI access into advantage;
- β = damage caused by constraints.
After n periods:
An = A0 + Σt=0n−1[αQt − βCt]
The access threshold required merely to avoid declining relative advantage is:
Qt* = βCt ÷ α
If Qt < Qt*, accumulated advantage falls relative to the competitive environment.
7.11.2 Two-user simulation
Assume:
- α = 0.8;
- β = 0.5;
- high-access user: Q = 1.0 and C = 0.2;
- constrained user: Q = 0.3 and C = 0.8;
- both begin with A0 = 10.
Annual increments are:
ΔAhigh = 0.8 × 1.0 − 0.5 × 0.2 = 0.70
ΔAconstrained = 0.8 × 0.3 − 0.5 × 0.8 = −0.16
| Year | High-access advantage | Constrained-user advantage | Gap |
|---|---|---|---|
| 0 | 10.00 | 10.00 | 0.00 |
| 1 | 10.70 | 9.84 | 0.86 |
| 2 | 11.40 | 9.68 | 1.72 |
| 3 | 12.10 | 9.52 | 2.58 |
| 4 | 12.80 | 9.36 | 3.44 |
| 5 | 13.50 | 9.20 | 4.30 |
Even with equal starting capability, persistent access differences create divergence.
7.11.3 Feedback mechanism
In reality, accumulated advantage can finance future access:
Qt+1 = q0 + λAt − φPt
where:
- λ is the degree to which advantage improves future access;
- Pt is the effective price;
- φ is price sensitivity.
Substitution produces:
At+1 = At + α[q0 + λAt − φPt] − βCt
At+1 = (1 + αλ)At + αq0 − αφPt − βCt
When αλ is positive, advantage compounds rather than accumulating linearly.
7.12 The intelligence poverty trap
An intelligence poverty trap exists when low income or institutional resources cause weak AI access, weak access suppresses productivity and opportunity, and reduced opportunity prevents investment in better access.
The cycle is:
Low Resources → Low-Quality Access → Lower Productivity → Lower Income or Funding → Low Resources
The trap condition can be formalised as:
Return from AI Accesslow-resource < Cost of Adequate Access
while:
Return from AI Accesshigh-resource > Cost of Adequate Access
This can occur even when the same nominal subscription price applies to both groups because:
- the price consumes a larger income share;
- poorer users have weaker hardware;
- connectivity is less reliable;
- free time for adaptation is limited;
- schools and firms provide less support;
- users cannot absorb experimental failure;
- local language performance may be lower;
- payment systems or regional availability may restrict access.
7.12.1 Trap threshold
Let income evolve as:
Yt+1 = Yt + θAt
and access be:
Qt = max[0, q − P ÷ Yt]
A low-income equilibrium exists when P ÷ Yt is large enough to suppress Q, preventing A and Y from rising. A subsidy lowering P, a university providing shared access or a specialised low-cost model increasing q can move the user above the threshold.
7.13 Interventions and their economic logic
| Intervention | Mechanism | Principal risk |
|---|---|---|
| Student AI vouchers | Reduces subscription cost | Benefits vendor without ensuring quality |
| University enterprise access | Pools procurement and governance | Institutional lock-in |
| Public compute centres | Provides local capacity | Underutilisation and maintenance |
| Cooperative research infrastructure | Shares fixed cost | Complex allocation and confidentiality |
| Open-weight model support | Reduces supplier dependence | Technical skill burden |
| Device grants | Reduces hardware barrier | Rapid obsolescence |
| Multilingual evaluation | Exposes language gaps | Benchmark may not reflect real use |
| AI-literacy programmes | Raises conversion coefficient α | Training without continued access has low value |
| Research API credits | Supports reproducible experimentation | Credits expire or favour certain providers |
| Public-interest model routing | Chooses cheapest adequate model | Requires independent evaluation |
| Offline and low-bandwidth systems | Reduces connectivity barrier | Capability may remain below frontier |
| Outcome-based grants | Funds demonstrated research value | Can disadvantage exploratory work |
The optimal policy does not give every user the most expensive frontier model for every task. It ensures access to the lowest-cost system that meets the task’s quality threshold, together with escalation to frontier capacity when smaller systems fail.
7.14 Research design for causal measurement
7.14.1 Randomised access study
Participants should be randomly assigned to:
- no AI;
- free-tier AI;
- paid consumer AI;
- frontier enterprise AI.
All groups receive the same tasks, time windows and initial training except where training itself is part of the treatment. Outcomes include:
- test score;
- completion time;
- retention after four weeks;
- error rate;
- transfer to new tasks;
- confidence calibration;
- dependency;
- cost;
- energy.
The treatment effect is:
ATE = E[Y|Access = j] − E[Y|Access = 0]
7.14.2 Longitudinal study
A longitudinal panel should track:
- prior attainment;
- socioeconomic status;
- AI access history;
- model tier;
- actual usage;
- grades;
- publications;
- employment;
- income;
- entrepreneurship;
- skills.
A fixed-effects specification is:
Yi,t = βQAAi,t + γXi,t + μi + τt + εi,t
Individual fixed effects μi control for stable unobserved differences.
7.14.3 Institutional natural experiments
Possible quasi-experiments include:
- staggered university licence rollout;
- temporary free-credit programmes;
- regional connectivity changes;
- laboratory hardware grants;
- service-tier changes;
- model withdrawal;
- price increases.
A difference-in-differences estimator is:
Effect = [Ytreated,after − Ytreated,before] − [Ycontrol,after − Ycontrol,before]
Parallel pre-treatment trends must be tested.
7.15 Five-year scenarios, 2027–2031
| Scenario | Access development | Educational and social effect |
|---|---|---|
| Broad diffusion | Capable small models and institutional access expand | Divide narrows for routine tasks |
| Tiered intelligence | Free models improve, but frontier models advance faster | Nominal access broadens while capability gap persists |
| Subscription escalation | High-end access becomes more expensive and metered | Students and independent researchers lose relative access |
| Public-infrastructure response | Universities and governments provide shared capacity | Geographic divide narrows where institutions are effective |
| Sovereign fragmentation | Model and hardware access differs by region | Nationality and location become capability determinants |
| Agentic concentration | Advanced agents require enterprise data and tools | Organisational affiliation dominates individual talent |
The central scenario is tiered intelligence. Basic AI becomes ubiquitous, but frontier-quality reasoning, long-context analysis, private data integration, sustained agentic work and high-volume usage remain concentrated among wealthy individuals and institutions. The resulting divide is less visible than total exclusion because both rich and poor users can claim to “have AI,” while the quality and productive capacity of that access differ substantially.
7.16 Falsifiable hypotheses
| Hypothesis | Testable proposition | Required evidence |
|---|---|---|
| H7.1 | Paid AI access improves task completion relative to no access | Randomised user study |
| H7.2 | Frontier access produces greater gains than free-tier access on complex tasks | Multi-treatment experiment |
| H7.3 | Smaller local models underperform frontier models on general complex tasks | Controlled benchmark |
| H7.4 | Specialised small models can match frontier systems on bounded tasks | Task-specific quality-adjusted cost study |
| H7.5 | AI access increases research output | Longitudinal researcher panel |
| H7.6 | AI access increases publication quantity more than quality | Publication, citation and replication analysis |
| H7.7 | Multilingual capability gaps reduce gains outside dominant languages | Parallel multilingual experiment |
| H7.8 | AI affordability is negatively associated with national income | Country-level AAI panel |
| H7.9 | Access effects compound over time | Longitudinal cumulative-advantage model |
| H7.10 | Institutional access mediates the effect of personal income | Multilevel student and researcher model |
| H7.11 | High AI prices reduce entrepreneurial entry | Regional or cohort study |
| H7.12 | Free-tier access does not eliminate the frontier capability divide | Quality-adjusted access comparison |
7.17 Final findings
- AI access is not binary. No access, free access, paid consumer access and frontier institutional access represent materially different productive environments.
- Affordability must include hardware, connectivity, storage and usage. Subscription price alone substantially understates adequate-access cost.
- Student affordability is structurally worse than household affordability. Students often possess far less disposable income than the national median.
- An AI-capable workstation can cost less than one median monthly wage in a high-income country but many months of earnings in lower-income markets. Tariffs and local scarcity can widen the difference.
- Smaller local models do not always create a disadvantage. They can be economically superior for bounded, specialised and privacy-sensitive tasks.
- Frontier models retain likely advantages in novel reasoning, difficult coding, long-context synthesis and multidisciplinary scientific work.
- Parameter count is not an adequate utility measure. Architecture, post-training, quantisation, tools, retrieval, context and task design must be controlled.
- Benchmark cost must be divided by successful workloads, not generated tokens. Cheap incorrect output is not affordable intelligence.
- AI can improve language access while preserving hidden quality disparities. Fluency does not guarantee equivalent reasoning or factual reliability.
- Institutional affiliation is becoming a determinant of effective capability. Researchers inside wealthy universities or companies can access data, tools, compute and support unavailable to independent peers.
- Cumulative advantage is mathematically plausible. Persistent differences in quality-adjusted access generate widening human-capital and institutional gaps even when starting ability is equal.
- A self-reinforcing intelligence poverty trap is possible. Low resources reduce access; weak access reduces productivity and opportunity; reduced opportunity prevents future investment.
- Universal free access may not eliminate the divide. If frontier systems improve faster than free systems, nominal inclusion can coexist with growing capability inequality.
- The correct policy target is adequate task-specific access. Public support should supply the least-cost system meeting validated quality thresholds, with escalation to frontier capacity where necessary.
- Between 2027 and 2031, social mobility may increasingly depend on access not merely to AI, but to sufficiently capable, sustained, verifiable and economically usable AI.
CHAPTER 8 — Countries, Development and Global AI Inequality
8.1 Scope, purpose and evidentiary limits
Global AI inequality cannot be measured through the number of chatbot accounts, national AI strategies or data centres alone. Effective access depends on a connected system of household purchasing power, hardware affordability, electricity, broadband, cloud infrastructure, advanced semiconductors, foreign-currency availability, digital skills, local-language capability, research institutions, regulation and geopolitical permission to acquire technology. A country may possess strong mobile connectivity but little research compute; inexpensive electricity but an unreliable grid; cloud access but no domestic region; capable engineers but severe foreign-currency constraints; or extensive infrastructure whose benefits remain concentrated in a small urban elite. A multidimensional index must capture these complementarities without concealing them inside an opaque ranking.
The World Bank’s 2025 framework identifies four foundations of AI development: connectivity, compute, context and competency. It concludes that high-income countries continue to dominate AI innovation, compute infrastructure and startup funding, while low- and middle-income economies face persistent gaps in connectivity, locally relevant data, skills and compute. It also identifies “Small AI”—lower-cost, task-specific systems capable of running on ordinary devices—as a potential route to wider diffusion. Digital Progress and Trends Report 2025: Strengthening AI Foundations – World Bank – November/2025 — World Bank AI foundations report.
This chapter constructs a Global AI Access Index, or GAI, and applies it to a pilot set of 25 country and regional cases. The numerical scores are transparent analytical estimates constructed for scenario analysis; they are not an official World Bank, ITU or national-government ranking. The index must be recalculated from a frozen country-year dataset before publication as a definitive statistical table. This qualification is necessary because several required variables—particularly local-language model quality, latency to AI-capable cloud capacity, research-compute availability and hardware street prices—are not yet available through a single harmonised international dataset.
8.2 Global connectivity and compute baseline
Almost three-quarters of the global population was online in 2025, but approximately 2.2 billion people remained offline. Internet use reached 94% in high-income economies but only 23% in low-income economies. Approximately 96% of offline people lived in low- and middle-income countries. Urban internet use reached 85%, compared with 58% in rural areas, while a typical user in a high-income country generated nearly eight times as much mobile data as a user in a low-income economy. Measuring Digital Development: Facts and Figures 2025 – International Telecommunication Union – November/2025 — ITU Facts and Figures 2025.
These differences matter because nominal coverage does not establish usable AI connectivity. AI use can require:
- sustained rather than intermittent service;
- sufficiently low latency;
- adequate upload capacity for documents and images;
- affordable data volume;
- secure access;
- compatible devices;
- reliable payment mechanisms;
- cloud services available in the jurisdiction;
- continuous electricity.
ITU’s concept of universal and meaningful connectivity includes quality, availability, affordability, devices, skills and security. Global Connectivity Report 2025 – International Telecommunication Union – 2025 — ITU Global Connectivity Report 2025. This is more appropriate to AI-access analysis than a binary connected/not-connected variable.
8.3 Country and regional comparison
8.3.1 Structural comparison
| Country or group | Principal strengths | Principal constraints | Strategic position |
|---|---|---|---|
| United States | Frontier providers, hyperscale cloud, accelerators, capital, universities | Regional inequality, high enterprise concentration, energy bottlenecks | Frontier producer and dominant platform jurisdiction |
| Canada | Research institutions, cloud access, skills, electricity | Smaller domestic market and dependence on foreign platforms | Advanced adopter with selected research strengths |
| European Union | Industrial base, research, regulation, multiple cloud regions | Fragmentation, limited frontier-platform ownership, energy cost | Large regulated market with incomplete compute sovereignty |
| United Kingdom | Strong research, finance, cloud access and language advantage | Dependence on imported hardware and foreign providers | Advanced service and research hub |
| China | Large market, data centres, domestic models and industrial policy | Export controls, leading-edge hardware constraints | Parallel ecosystem with substantial scale |
| Japan | High income, advanced industry, stable electricity and research | Ageing workforce, imported accelerator dependence | Advanced adopter and industrial AI power |
| South Korea | Semiconductors, broadband, cloud, digital skills | Imported frontier accelerators and geopolitical exposure | High-readiness technology producer |
| Israel | Advanced research, cybersecurity, startups and cloud investment | Small market, geopolitical risk and import dependence | High-capability specialist ecosystem |
| India | Large technical workforce, cloud regions, software sector and scale | Income, electricity and rural connectivity inequality | Rapid adopter with large internal divide |
| Brazil | Large market, cloud region, research and Portuguese-language scale | Hardware cost, taxation, income inequality and currency risk | Regional leader with affordability constraints |
| Mexico | Manufacturing integration, proximity to US, growing cloud access | Research-compute and skills inequality | Nearshore AI adopter |
| Croatia | EU membership, connectivity and regulatory integration | Small market and limited research compute | Higher-readiness Balkan adopter |
| Serbia | Technical workforce and regional connectivity | Non-EU position, smaller compute base and currency exposure | Emerging regional hub |
| Albania | Improving digital government and connectivity | Limited research capacity and income | Service-led emerging adopter |
| Bosnia and Herzegovina | Educated diaspora and regional access | Institutional fragmentation and limited investment | Capacity-constrained adopter |
| Morocco | Cloud and data-centre investment potential, multilingual workforce | Hardware affordability and research scale | Leading North African services platform |
| Tunisia | Technical education and nearshore potential | Financing, currency and infrastructure constraints | Talent-rich but capital-constrained |
| Egypt | Large market, cables, Arabic-language demand | Income, currency and infrastructure pressure | Potential regional platform with affordability barriers |
| South Africa | Strongest Sub-Saharan cloud footprint and research base | Electricity reliability, inequality and high hardware cost | Continental infrastructure leader with internal divide |
| Kenya | Digital services, entrepreneurship and regional connectivity | Research compute, electricity and income constraints | East African digital-services hub |
| Nigeria | Large population and entrepreneurship | Electricity reliability, foreign currency and infrastructure | Large demand base with severe access constraints |
| Ghana | Institutional stability and regional digital-services potential | Small compute base and affordability | Emerging public-service adopter |
| Rwanda | Digital-government orientation and policy coordination | Small market and imported infrastructure | Institutionally agile, compute-constrained |
| Ethiopia | Large population and development potential | Connectivity, electricity, foreign currency and skills constraints | Early-stage AI-access environment |
| Bangladesh | Large workforce and digital-services potential | Energy, hardware affordability and research-compute limitations | Labour-scale adopter with constrained infrastructure |
8.4 Global AI Access Index specification
8.4.1 Core equation
GAIi = Σj=1m wjzij
where:
- GAIi is the Global AI Access Index for country i;
- zij is the normalised value of indicator j;
- wj is the indicator weight;
- Σwj = 1.
The index is scaled from 0 to 100:
GAIi,100 = 100 × GAIi
A high score represents greater affordable access, infrastructure readiness and capability.
8.2 Indicator architecture
| Indicator | Symbol | Preferred empirical measure | Direction |
|---|---|---|---|
| Household income | INC | Median disposable income per capita, PPP-adjusted | Positive |
| Hardware affordability | HWA | Reference workstation cost divided by median income | Negative |
| Import duties | DUT | Effective tariff and tax burden on reference hardware | Negative |
| Electricity reliability | REL | Outage frequency/duration or reliable-access measure | Positive |
| Electricity cost | EPC | Commercial/household price per kWh | Negative |
| Broadband availability | BBA | Fixed/mobile meaningful-connectivity measure | Positive |
| Data-centre latency | LAT | Median latency to nearest AI-capable cloud region | Negative |
| Cloud-region availability | CRA | Number and capability of domestic/nearby hyperscale regions | Positive |
| Local-language quality | LLQ | Reproducible model-quality parity across national languages | Positive |
| Digital skills | DSK | Advanced ICT-skill rate or harmonised proxy | Positive |
| Research compute | RCO | Quality-adjusted public research accelerator capacity | Positive |
| Foreign-currency access | FXA | Convertibility and ability to finance technology imports | Positive |
| Sanctions/export exposure | SEE | Legal and practical restriction on advanced hardware/services | Negative |
| Local data-centre capacity | DCC | Operational IT capacity per capita or GDP | Positive |
| Regulatory environment | REG | Predictability, data governance and competition balance | Positive |
The ITU notes that comparable ICT-skills data remain scarce: only 90 countries had submitted relevant data since 2020, and only about 40 provided broadly comparable skill-level information. ITU ICT SDG Indicators – International Telecommunication Union – 2026 — ITU ICT-skills data limitations. Missing data must therefore be disclosed and not silently replaced by subjective assumptions.
8.5 Normalisation
8.5.1 Min-max normalisation
For a positive indicator:
zij = [xij − min(xj)] ÷ [max(xj) − min(xj)]
For a negative indicator:
zij = [max(xj) − xij] ÷ [max(xj) − min(xj)]
Min-max normalisation is intuitive but sensitive to outliers.
8.5.2 Winsorised min-max
Values below the fifth percentile are replaced with the fifth-percentile value; those above the ninety-fifth percentile are replaced with the ninety-fifth-percentile value:
xijW = min[max(xij, P5,j), P95,j]
The min-max calculation is then applied to xijW.
8.5.3 Z-score normalisation
zij = [xij − μj] ÷ σj
Z-scores preserve distance from the mean but can produce negative index components. A cumulative normal transformation can map them to 0–1.
8.5.4 Rank normalisation
zij = [rank(xij) − 1] ÷ (n − 1)
Rank normalisation reduces outlier influence but discards cardinal differences. A country narrowly above another receives the same rank interval as a country separated by a very large absolute gap.
8.6 Weighting systems
8.6.1 Equal weights
With 15 indicators:
wj = 1 ÷ 15 = 0.0667
Equal weights are transparent but imply that every indicator has identical conceptual importance.
8.6.2 Expert weights
| Indicator | Expert weight |
|---|---|
| Household income | 7% |
| Hardware affordability | 9% |
| Import duties | 4% |
| Electricity reliability | 8% |
| Electricity cost | 5% |
| Broadband availability | 8% |
| Data-centre latency | 5% |
| Cloud-region availability | 8% |
| Local-language quality | 7% |
| Digital skills | 8% |
| Research compute | 10% |
| Foreign-currency access | 4% |
| Sanctions/export exposure | 7% |
| Local data-centre capacity | 7% |
| Regulatory environment | 3% |
| Total | 100% |
The higher weights on hardware, research compute, broadband and electricity reliability reflect their role as binding complements. Regulation receives a lower direct weight because regulatory quality should not compensate arithmetically for the absence of electricity or compute.
8.6.3 Principal-component-derived weights
PCA weights should be derived from the first components explaining a predetermined share—such as 70%—of standardised variance. If loading ajk represents indicator j on component k and λk its eigenvalue:
wj,PCA = Σk=1K|ajk|λk ÷ Σj=1mΣk=1K|ajk|λk
PCA identifies empirical covariance, not moral or developmental importance. It can assign low weight to an essential variable if that variable has little cross-country variance.
8.7 Missing-data protocol
A country should not receive a favourable score because data are unavailable. The protocol should be:
- Report the missing indicator.
- Use multiple imputation only where predictor relationships are credible.
- Calculate an uncertainty interval across imputed datasets.
- Exclude countries missing more than 30% of weighted indicators.
- Publish a completeness score:
Completenessi = ΣwjI(xij observed)
- Report:
GAIadjusted,i = GAIi × Completenessiκ
where κ is a disclosed penalty parameter.
The unpenalised and adjusted scores should both be published.
8.8 Provisional pilot index
The following pilot scores use the expert-weight framework and a 0–100 scale. They represent analytical synthesis from the official structural evidence reviewed in this report, not a frozen final statistical dataset.
| Country or regional case | Provisional GAI | Access tier | Principal binding constraint |
|---|---|---|---|
| United States | 89 | Frontier | Internal affordability and geographic inequality |
| South Korea | 84 | Advanced | Imported frontier accelerator exposure |
| Israel | 83 | Advanced | Scale and geopolitical exposure |
| Canada | 82 | Advanced | Foreign-platform dependence |
| United Kingdom | 82 | Advanced | Imported hardware and platform dependence |
| European Union aggregate | 80 | Advanced | Internal fragmentation and frontier-provider deficit |
| Japan | 80 | Advanced | Imported accelerator dependence |
| China | 78 | Advanced/parallel | Export controls and leading-edge hardware |
| India | 58 | Emerging-scale | Income and internal infrastructure inequality |
| Croatia | 58 | Emerging/advanced | Small research-compute base |
| Brazil | 55 | Emerging-scale | Hardware affordability and taxation |
| South Africa | 52 | Emerging regional hub | Electricity reliability and inequality |
| Mexico | 51 | Emerging | Research compute and skills dispersion |
| Morocco | 50 | Emerging regional hub | Research scale and hardware affordability |
| Serbia | 49 | Emerging | Cloud depth, market size and currency exposure |
| Egypt | 45 | Constrained scale | Currency, income and infrastructure pressure |
| Tunisia | 43 | Constrained talent hub | Financing and foreign-currency access |
| Kenya | 43 | Emerging digital hub | Research compute and household affordability |
| Albania | 42 | Emerging | Small market and limited compute |
| Rwanda | 40 | Policy-led emerging | Scale and imported infrastructure |
| Bosnia and Herzegovina | 38 | Constrained | Institutional fragmentation and investment |
| Ghana | 38 | Constrained emerging | Compute and affordability |
| Nigeria | 35 | High-potential constrained | Electricity, currency and infrastructure |
| Bangladesh | 32 | Constrained scale | Compute, energy and affordability |
| Ethiopia | 22 | Severely constrained | Connectivity, electricity and foreign currency |
An aggregate EU score conceals major internal variation. Countries with domestic cloud regions, high incomes and strong research systems will score materially above lower-income member states with smaller compute ecosystems. The same applies to the United States, China, India, Brazil, South Africa and Nigeria: national averages conceal extreme urban-rural and income-based divergence.
8.9 Sensitivity analysis
8.9.1 Alternative weighting effects
| Country group | Equal weights | Expert weights | Expected PCA effect |
|---|---|---|---|
| United States, Canada, UK | Remain high | Remain high | Strong infrastructure-income component sustains rank |
| EU aggregate | High | High | Regulatory component may matter less under PCA |
| China | High | Moderately reduced by export exposure | Industrial scale may raise PCA rank |
| South Korea and Japan | High | High | Semiconductor and connectivity strength reinforced |
| Israel | High | High | Small market may reduce scale-based PCA score |
| India | Middle | Middle | Digital scale raises rank; income reduces it |
| Brazil and Mexico | Middle | Middle | Cloud and market scale offset affordability |
| Balkans | Middle/lower | Highly country-specific | EU membership and connectivity differentiate |
| North Africa | Lower-middle | Lower-middle | Currency and research compute reduce expert score |
| Sub-Saharan Africa | Lower | Lower | Connectivity and income covariance reinforce PCA penalty |
8.9.2 Rank robustness standard
For each country:
RankRangei = [min Ranki,s, max Ranki,s]
where s covers:
- equal weights;
- expert weights;
- PCA weights;
- min-max;
- winsorised min-max;
- z-score;
- rank normalisation;
- alternative missing-data treatments.
A country should be assigned a robust tier only if it remains in that tier across at least 75% of specifications.
8.9.3 Monte Carlo sensitivity
Weights can be drawn from a Dirichlet distribution centred on expert weights:
w(b) ∼ Dirichlet(κwexpert)
For each of B simulations:
GAIi(b) = Σwj(b)zij
The report should publish:
- median score;
- fifth and ninety-fifth percentiles;
- probability of belonging to each tier;
- probability of outranking selected peers.
8.10 Inequality measurement
8.10.1 Gini coefficient
For access values xi:
G = [Σi=1nΣj=1n|xi − xj|] ÷ [2n2x̄]
Using the 25 provisional GAI scores:
- n = 25;
- mean GAI = 56.36;
- Gini coefficient = 0.193.
The value indicates substantial but not extreme dispersion among country averages. It understates inequality because:
- country averages weight Ethiopia and the United States equally rather than by population;
- within-country inequality is omitted;
- a score of zero is bounded;
- access quality may have nonlinear economic returns.
8.10.2 Theil index
The Theil T index is:
T = (1 ÷ n)Σi=1n(xi ÷ x̄)ln(xi ÷ x̄)
For the provisional score distribution:
T = 0.0597
The Theil index is decomposable:
T = Tbetween + Twithin
Dividing the pilot cases into advanced, emerging and constrained groups produces:
| Component | Theil contribution | Share |
|---|---|---|
| Between-group inequality | 0.0546 | 91.4% |
| Within-group inequality | 0.0051 | 8.6% |
| Total | 0.0597 | 100% |
This result reflects the constructed pilot grouping and should not be generalised to the world population. It nevertheless demonstrates that development-tier differences dominate this illustrative country-average distribution.
8.10.3 Percentile ratios
| Measure | Pilot value |
|---|---|
| P90/P10 | 2.28 |
| P80/P20 | 2.03 |
| Maximum/minimum | 4.05 |
The maximum pilot score, 89, is slightly more than four times the minimum, 22.
8.10.4 Palma ratio
The conventional Palma ratio is the share held by the top 10% divided by that of the bottom 40%:
Palma = Sharetop10 ÷ Sharebottom40
Applied mechanically to the equally weighted country-score distribution, the pilot Palma ratio is approximately 0.57. A value below one is possible because the bottom 40% contains four times as many country observations as the top 10%, while GAI scores are bounded and much less concentrated than income. For AI analysis, a population-weighted Palma and a compute-capacity Palma would be more informative.
8.11 Population weighting and within-country inequality
A country-level index can conceal more inequality than it reveals. Population-weighted access should use:
x̄population = ΣPixi ÷ ΣPi
Within a country, access can be modelled by household decile, region, gender, rurality, education and institutional affiliation:
GAIi,h = f(Incomeh, Deviceh, Connectivityh, Skillsh, InstitutionalAccessh)
The national score is:
GAIi = Σh=1HphGAIi,h
A high national average can coexist with severe exclusion among rural households, informal workers, minority-language communities and underfunded institutions.
8.12 Cloud geography and latency inequality
Cloud geography matters because local regions can reduce latency, support data residency, improve availability and reduce international bandwidth dependence. The global footprint remains geographically concentrated.
AWS reported 123 Availability Zones across 39 geographic regions, with additional regions planned. AWS Global Infrastructure – Amazon Web Services – August/2026 — AWS global infrastructure. Microsoft reported more than 80 Azure regions and more than 500 data centres on its current global-infrastructure page. Azure Global Infrastructure – Microsoft – August/2026 — Microsoft Azure global infrastructure. Google Cloud reported 43 regions and 130 zones across six continents. Global Locations: Regions and Zones – Google Cloud – August/2026 — Google Cloud global locations.
A domestic region should not be coded simply as one. The cloud-region indicator should measure:
CRAi = Availability × AIServiceDepth × AcceleratorAvailability × Redundancy × Residency
A region lacking frontier accelerators or the required managed AI service is not equivalent to a fully provisioned US region.
Latency should be measured through repeated active tests:
LATi = median RTT from representative population nodes to nearest eligible AI endpoint
Eligibility must include:
- model availability;
- legal access;
- data-residency compliance;
- sufficient capacity;
- payment availability.
8.13 Export controls and geopolitical access
Advanced-compute access is partly a legal and geopolitical variable. US controls apply to specified advanced computing semiconductors, semiconductor-manufacturing equipment, high-bandwidth memory and associated entities and destinations. Commerce Strengthens Restrictions on Advanced Computing Semiconductors and Semiconductor Manufacturing Equipment – Bureau of Industry and Security – December/2024 — US BIS advanced-semiconductor controls.
In January 2026, BIS moved to case-by-case review for specified H200-, MI325X- and similar-class exports to China subject to stated conditions. Department of Commerce Revises License Review Policy for Semiconductors Exported to China – Bureau of Industry and Security – January/2026 — US BIS China semiconductor licensing policy.
The sanctions/export indicator should distinguish:
| Score component | Question |
|---|---|
| Legal eligibility | Can the country legally import the hardware? |
| Entity eligibility | Are principal firms or universities restricted? |
| Licence probability | Are licences routinely approved, conditional or presumptively denied? |
| Cloud substitutability | Can equivalent capacity be accessed remotely? |
| Payment access | Can customers legally and practically pay? |
| Diversion risk | Do compliance concerns discourage suppliers? |
| Domestic alternative | Can local technology substitute? |
China differs fundamentally from a low-income sanctioned economy: it possesses a large domestic industrial, cloud and model ecosystem capable of partial substitution. The same numerical penalty should not be applied without a domestic-capability offset.
8.14 Electricity inequality
AI readiness depends on both electricity access and quality:
EnergyReadinessi = Access × Reliability × Capacity × Affordability × Expandability
A country may report near-universal electricity access while experiencing outages that make continuous local inference unreliable. Another may offer cheap tariffs but lack generation or grid capacity for new data centres. A third may possess reliable electricity whose price makes domestic compute uncompetitive.
For local AI:
Costenergy = PIT × 8,760 × u × PUE × pelectricity
For national AI infrastructure, connection lead time and firm power availability may matter more than average tariff.
8.15 Local-language and cultural representation
8.15.1 Language-quality indicator
Local-language model quality should be measured as:
LLQi = Σl=1Lsl[Scorel ÷ Scorereference]
where sl is the population share using language l.
Tests should cover:
- factual questions;
- public-service terminology;
- legal and medical language;
- educational content;
- dialects;
- culturally specific reasoning;
- toxicity and refusal;
- speech recognition;
- document OCR;
- translation.
A model may support a language conversationally while performing poorly on technical, scientific or administrative tasks. Countries with large language markets—Chinese, Japanese, Korean, Arabic, Portuguese and Hindi—possess stronger commercial incentives for localisation than small-language communities.
8.15.2 Cultural dependence
Dependence on foreign platforms can shape:
- which sources are retrieved;
- which dialects are recognised;
- moderation boundaries;
- historical representation;
- public-service terminology;
- political and cultural assumptions;
- data retained outside the country.
Scientific sovereignty therefore includes the ability to evaluate, adapt and govern models—not merely host them.
8.16 Development consequences
8.16.1 Economic development
AI can support development by lowering the cost of knowledge-intensive services, enabling small firms to reach international markets and improving logistics, agriculture, finance and administration. The development contribution can be represented as:
ΔYi = αAi + βKi + γSi + δDi − λMi
where:
- Ai = AI adoption;
- Ki = complementary capital;
- Si = skills;
- Di = data and institutional quality;
- Mi = import and monopoly leakage.
If AI services are imported while little domestic value is created, productivity may rise without generating a large domestic AI industry.
8.16.2 Productivity convergence
AI promotes convergence if lower-income countries obtain capability at a lower cost than developing equivalent expertise internally:
Convergence effect > 0 when:
Imported AI Productivity Gain > Subscription Leakage + Displacement + Adaptation Cost
AI promotes divergence when leading countries combine better models with greater capital, skills and data:
Divergence effect > 0 when:
Complementarityfrontier × Existing Capitalrich > Catch-up Gainpoor
The likely result is heterogeneous. Basic administrative and translation tasks may converge, while frontier science and advanced industrial design diverge.
8.16.3 Education
Low-cost AI can distribute tutoring and translation to underserved regions. It can also create a new inequality between students using free generic systems and those with frontier access, institutional data, high-performance devices and expert supervision.
8.16.4 Public services
Governments can deploy AI for:
- document processing;
- fraud detection;
- tax administration;
- citizen support;
- translation;
- resource allocation;
- health triage.
The World Bank’s 2025 GovTech Maturity Index covers 198 economies and reports a global average increase from 0.552 in 2022 to 0.589 in 2025, with widening differences between the strongest and weakest maturity groups. GovTech Maturity Index 2025 – World Bank – December/2025 — World Bank GovTech Maturity Index.
8.16.5 Healthcare
AI can extend scarce specialist capacity through decision support, imaging assistance, translation and administrative automation. Risks arise when imported models are not validated on local populations, diseases, languages or clinical practice.
8.16.6 Entrepreneurship
Cloud APIs reduce entry costs but expose startups to:
- currency depreciation;
- foreign payment restrictions;
- sudden repricing;
- model withdrawal;
- data-residency limits;
- API lock-in;
- unequal access to enterprise discounts.
A startup paying dollar-denominated AI costs while earning local-currency revenue carries an embedded exchange-rate mismatch.
8.16.7 Skilled migration
Talent may move toward countries and institutions offering:
- frontier compute;
- higher salaries;
- research datasets;
- strong laboratories;
- startup capital;
- cloud capacity.
The migration feedback is:
Research Compute → Talent Attraction → Publications and Startups → Funding → More Research Compute
The reverse can become a scientific-capability trap.
8.16.8 National security
Dependence on foreign AI platforms affects:
- intelligence analysis;
- defence supply chains;
- critical-infrastructure operations;
- cyber defence;
- government data;
- continuity during geopolitical crisis.
Sovereignty does not require domestic production of every chip. It requires credible continuity options, diversified providers, portable systems, local evaluation and priority access during crisis.
8.17 Evidence that AI can reduce development gaps
| Equalising mechanism | Development effect |
|---|---|
| Small AI on ordinary devices | Reduces hardware requirements |
| Cloud access | Eliminates minimum local data-centre investment |
| Translation | Broadens access to global knowledge |
| Open-weight models | Reduces licence and platform dependence |
| Shared public compute | Spreads fixed cost |
| Mobile distribution | Reaches users without personal computers |
| Agricultural and health specialisation | Targets high-value local problems |
| Coding assistance | Expands software-production capability |
| Digital public infrastructure | Enables scalable government services |
| International research credits | Provides temporary frontier access |
The World Bank reports that middle-income countries accounted for more than 40% of global ChatGPT traffic by mid-2025, led by countries including Brazil, India, Indonesia and Viet Nam. It also reports rapid growth in generative-AI-related vacancies in middle-income economies. Strengthening AI Foundations: Emerging Opportunities for Developing Countries – World Bank – November/2025 — World Bank developing-country AI factsheet.
8.18 Evidence that AI can widen development gaps
| Divergence mechanism | Consequence |
|---|---|
| Expensive accelerators and memory | Capital-intensive capability remains concentrated |
| Import taxes and currency depreciation | Hardware costs rise relative to local income |
| Electricity unreliability | Local inference and data centres become costly |
| Lack of cloud regions | Higher latency and weaker sovereignty |
| Export controls | Advanced capacity differs by geopolitical alignment |
| Dollar-denominated subscriptions | Currency shocks affect access |
| Weak local-language performance | Lower productivity gain |
| Research-compute concentration | Frontier science remains geographically concentrated |
| Platform lock-in | Domestic value leaks to foreign providers |
| Skilled migration | Weak ecosystems lose their most capable researchers |
| Data scarcity | Models perform worse on local conditions |
| Regulatory uncertainty | Investment and deployment are delayed |
| Free-versus-frontier tiering | Nominal inclusion masks capability inequality |
UNCTAD reported that data centres accounted for more than one-fifth of global greenfield investment in 2025, illustrating the growing capital intensity of digital infrastructure. Data Centres Are Reshaping the Global Investment Landscape – UN Trade and Development – January/2026 — UNCTAD data-centre investment analysis.
8.19 Regional findings
8.19.1 North America
The United States remains the principal frontier producer. Canada benefits from geographic integration, research and cloud access but remains dependent on foreign platforms and hardware supply. Internal inequality is the principal risk: national abundance does not guarantee affordable frontier access for every student, household or small firm.
8.19.2 European Union and United Kingdom
Europe possesses strong research, industrial capacity, connectivity and regulatory institutions but lacks equivalent ownership of the dominant frontier platforms and advanced accelerator design ecosystem. The EU aggregate conceals major differences between northern and western member states and lower-income or smaller members. The UK retains strong research and financial-sector demand but shares Europe’s dependence on imported hardware and foreign hyperscalers.
8.19.3 East Asia
China, Japan and South Korea all possess high capability but different constraints. China has market scale and domestic platforms but faces export controls. Japan has strong industrial and research capacity but depends on imported leading accelerators. South Korea combines world-leading memory and semiconductor capabilities with high connectivity, although frontier-accelerator access remains geopolitically exposed.
8.19.4 Israel
Israel’s strengths in research, cybersecurity, semiconductors and startups yield high readiness. Its constraints are market size, regional security, infrastructure concentration and dependence on imported leading-edge hardware.
8.19.5 India
India demonstrates the difference between national capability and household access. It possesses world-scale software talent, domestic cloud regions and an expanding AI market, while income, language, rural connectivity and educational inequality produce a very wide internal distribution.
8.19.6 Latin America
Brazil and Mexico possess large markets and cloud access but face hardware affordability, taxation, exchange-rate and research-compute constraints. Brazil benefits from Portuguese-language scale; Mexico benefits from proximity and industrial integration with the United States.
8.19.7 Balkans
Croatia benefits from EU integration. Serbia has technical talent and potential regional importance but a smaller domestic compute base. Albania has made progress in digital government, while Bosnia and Herzegovina faces institutional fragmentation. Regional cooperative compute could be economically superior to duplicated small national clusters.
8.19.8 North Africa
Morocco has strong nearshore and multilingual potential; Tunisia possesses significant technical talent but faces capital and currency constraints; Egypt combines population scale, cable geography and Arabic-language demand with severe affordability pressures. These countries could become localisation and service hubs if electricity, cloud and research compute expand.
8.19.9 Sub-Saharan Africa
South Africa has the region’s strongest hyperscale and research position but suffers electricity and income inequality. Kenya is an East African digital-services hub. Nigeria has enormous market potential but severe power and currency constraints. Ghana and Rwanda have policy and digital-government strengths but limited compute. Ethiopia remains constrained across several complementary foundations.
8.20 Policy architecture
| Policy objective | Instrument | Measurable outcome |
|---|---|---|
| Reduce hardware burden | Tariff relief and pooled procurement | Workstation wage-month ratio |
| Improve power reliability | Grid and dedicated research-power investment | Outage-adjusted compute availability |
| Expand cloud access | Regional investment and competition policy | Latency and eligible service depth |
| Protect affordability | Student, SME and researcher credits | AAI and SAI reduction |
| Build scientific sovereignty | National or regional research clusters | Research accelerator-hours per researcher |
| Improve language inclusion | Local-language datasets and evaluations | LLQ parity score |
| Reduce lock-in | Portability and interoperability rules | Switching time and cost |
| Address currency risk | Local-currency billing and hedging facilities | AI price volatility |
| Expand skills | Applied AI and verification training | Successful-task productivity |
| Support small AI | Grants for local task-specific systems | Cost per validated local workload |
| Enable regional cooperation | Shared compute and data infrastructures | Utilisation and member access |
| Protect security | Supply diversification and fallback capacity | Recovery time during provider disruption |
8.21 Five-year scenarios, 2027–2031
| Scenario | Core development | Inequality result |
|---|---|---|
| Inclusive diffusion | Small AI, open models and public infrastructure scale | Global divide narrows |
| Platform dependence | Cloud access expands without domestic capability | Consumption rises; sovereignty gap widens |
| Frontier bifurcation | Basic AI becomes cheap while frontier compute remains scarce | Nominal access converges; productive capability diverges |
| Geopolitical fragmentation | Export controls and sovereign ecosystems expand | Access follows alliances and jurisdiction |
| Infrastructure leapfrog | Selected emerging economies build energy and cloud hubs | New regional leaders emerge |
| Currency and debt shock | Import and subscription affordability deteriorates | Low-income access contracts |
| Regional cooperation | Shared Balkan, African or Latin American capacity grows | Small-country scale disadvantage falls |
The most plausible central scenario is frontier bifurcation. Low-cost AI becomes widespread through mobile services, smaller models and cloud APIs, generating real benefits in education, translation, agriculture, administration and entrepreneurship. Simultaneously, frontier research, high-assurance enterprise systems, advanced multimodal agents and sovereign local deployment remain concentrated in countries possessing abundant capital, reliable electricity, leading hardware, cloud regions and research institutions.
8.22 Falsifiable hypotheses
| Hypothesis | Test |
|---|---|
| H8.1 | Higher GAI predicts higher firm-level AI adoption after controlling for income |
| H8.2 | Hardware wage-month ratios predict local-model adoption |
| H8.3 | Domestic cloud regions increase enterprise AI adoption and reduce latency |
| H8.4 | Electricity unreliability reduces local compute investment |
| H8.5 | Export restrictions reduce research-compute growth in affected jurisdictions |
| H8.6 | Local-language quality predicts public and SME adoption |
| H8.7 | Foreign-currency shocks reduce API consumption in developing economies |
| H8.8 | Public research compute reduces skilled migration |
| H8.9 | Small AI produces larger proportional gains in low-income settings for bounded tasks |
| H8.10 | Frontier capability remains more concentrated than basic AI use |
| H8.11 | Between-country inequality is smaller than combined between- and within-country inequality |
| H8.12 | Regional cooperative infrastructure improves access in small economies |
8.23 Final findings
- AI readiness is a system of complements. Connectivity without electricity, skills without compute or cloud access without foreign currency cannot produce full capability.
- The global digital divide remains large. Approximately 2.2 billion people remained offline in 2025, overwhelmingly in low- and middle-income economies.
- National averages conceal severe internal inequality. India, Brazil, China, the United States, South Africa and Nigeria contain both globally competitive centres and severely constrained communities.
- Cloud geography is a development variable. Domestic and nearby regions affect latency, availability, residency and bargaining power.
- Advanced hardware access is geopolitical. Export controls and sanctions can alter national compute capacity independently of income or technical skill.
- Hardware affordability should be measured in wage months, not nominal dollars. Identical equipment can represent less than one month of earnings in one country and many months in another.
- Local-language quality is part of infrastructure. A country cannot be considered fully AI-ready if its population receives materially weaker performance in national languages.
- Research compute is a sovereignty asset. Its absence can drive publication gaps, dependence on foreign platforms and skilled migration.
- AI can reduce development gaps. Small models, mobile access, translation, cloud APIs and shared infrastructure can distribute useful capability at low marginal cost.
- AI can simultaneously widen frontier gaps. Advanced accelerators, energy, data centres, research funding and enterprise tools remain highly concentrated.
- The provisional pilot Gini of 0.193 understates total inequality. It omits population weighting and within-country access disparities.
- The pilot Theil decomposition attributes 91.4% of measured dispersion to differences between broad readiness groups. This result is illustrative and depends on the constructed sample and grouping.
- The correct policy objective is not national ownership of every technology. It is affordable access, continuity, evaluation capacity, portability and the ability to adapt AI to local needs.
- Regional cooperative compute is especially important for small countries. It can spread fixed costs and reduce dependence, provided governance and capacity allocation are credible.
- The central risk through 2031 is a world in which basic AI becomes nearly universal while frontier-quality, private, reliable and scientifically useful AI remains concentrated. Such a system would reduce some development gaps while institutionalising a deeper global hierarchy of intelligence capability.
CHAPTER 9 — Five-Year Forecasts and Quantitative Scenarios, 2026–2031
9.1 Forecast objective, base year and evidentiary status
Forecasting AI infrastructure differs from forecasting a mature commodity. The system is undergoing simultaneous structural changes in model architecture, accelerator design, memory intensity, advanced packaging, data-centre construction, electricity procurement, inference optimisation, service pricing, regulation and geopolitical access. Historical series are short, product generations are not quality-equivalent, reported prices frequently differ from negotiated prices, and several variables—particularly billable inference tokens and advanced-packaging output—are not disclosed through complete global datasets. A single extrapolated trend would therefore create false confidence.
This chapter uses a model ensemble. Where an official physical baseline exists, such as US data-centre electricity consumption, the forecast reports physical units. Where the underlying global level is not observed consistently, it reports an index with 2026=100. Nominal-price forecasts are separated from quality-adjusted prices. Central, low and high trajectories are scenario paths rather than claims of deterministic outcomes. The 80% and 95% ranges are probabilistic simulation intervals only where a parameter distribution has been specified; they are not presented as classical time-series confidence intervals when the historical sample is inadequate.
The official evidence establishes a high-growth but capacity-constrained starting point:
- US data-centre electricity consumption increased from 58 TWh in 2014 to 176 TWh in 2023 and was officially projected at 325–580 TWh by 2028. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers – U.S. Department of Energy – December/2024 — US Department of Energy data-centre electricity forecast.
- TSMC stated in April 2025 that it was working to double CoWoS capacity during 2025. TSMC First Quarter 2025 Earnings Conference – TSMC – April/2025 — TSMC Q1 2025 investor transcript.
- Micron stated in March 2026 that DRAM and NAND supply-demand conditions were expected to remain tight beyond calendar 2026. Micron Fiscal Second Quarter 2026 Earnings Call – Micron – March/2026 — Micron Q2 FY2026 investor disclosure.
- Microsoft disclosed USD 31.9 billion of quarterly capital expenditure in fiscal Q3 2026, with approximately two-thirds directed to shorter-lived assets, primarily GPUs and CPUs. Microsoft Fiscal Year 2026 Third Quarter Earnings Conference Call – Microsoft – April/2026 — Microsoft FY2026 Q3 investor disclosure.
- Alphabet disclosed USD 44.9 billion of capital expenditure in Q2 2026, with the vast majority directed to technical infrastructure supporting AI; approximately 60% of technical-infrastructure investment was in servers and 40% in data centres and networking. Alphabet 2026 Q2 Earnings Call – Alphabet – July/2026 — Alphabet Q2 2026 investor disclosure.
9.2 Forecast variables and measurement units
| Variable | Forecast unit | Base-year interpretation |
|---|---|---|
| Consumer and professional GPU price | Nominal price index, 2026=100 | Comparable AI-capable hardware basket |
| Quality-adjusted accelerator price | Cost per successful workload index | Controls for memory, throughput and quality |
| HBM price | Price-per-GB index | Blended high-bandwidth-memory generation |
| DRAM price | Price-per-GB index | Blended server and workstation DRAM |
| Advanced packaging | Capacity index | Quality-adjusted CoWoS-equivalent capacity |
| Data-centre CAPEX | Nominal global investment index | AI and cloud technical infrastructure |
| Electricity demand | US data-centre TWh plus global index | Physical US anchor |
| AI inference demand | Quality-adjusted successful-task index | Not raw generated tokens alone |
| API price | Cost per successful workload index | Model-, tool- and quality-adjusted |
| Enterprise AI expenditure | Real expenditure index | Licences, compute, integration and governance |
| Local-AI affordability | Cost-to-income index | Lower values indicate greater affordability |
| Global access inequality | Pilot GAI Gini | Country-average quality-adjusted access |
| Sectoral adoption cost | Real total-cost index | Full implementation and governance burden |
9.3 Forecasting methods
9.3.1 Trend and CAGR model
For a variable X:
CAGR = (XT ÷ X0)1/T − 1
and:
X̂t = X0(1 + CAGR)t
CAGR is used only as a descriptive baseline. It is unreliable when supply constraints, product cycles or regulatory shocks change the growth regime.
9.3.2 Exponential smoothing
Simple exponential smoothing is:
Lt = αXt + (1−α)Lt−1
A damped-trend extension is:
X̂t+h = Lt + (φ + φ2 + … + φh)Bt
where φ below one prevents implausibly persistent exponential growth.
9.3.3 ARIMA
An ARIMA(p,d,q) model is:
φ(B)(1−B)dXt = c + θ(B)εt
ARIMA should be used only for sufficiently long and consistent monthly or quarterly series, such as memory spot prices or electricity demand. It is unsuitable for a five-point annual series of non-comparable accelerator products.
9.3.4 Dynamic regression
For memory or accelerator prices:
ΔlnPt = α + β1ΔlnDemandt − β2ΔlnCapacityt + β3Energyt + β4FXt + β5Restrictiont + εt
This allows supply, demand, energy, currency and geopolitical variables to affect prices.
9.3.5 Panel-data model
For country i and year t:
GAIi,t = α + β1Incomei,t + β2Broadbandi,t + β3Poweri,t + β4Cloudi,t + β5Skillsi,t + μi + τt + εi,t
Country fixed effects μi control for time-invariant structural characteristics; period effects τt control for global changes.
9.3.6 Stock-flow model
Installed AI capacity evolves as:
Kt+1 = Kt + It − δKt − Rt
where:
- Kt is productive capacity;
- It is new installation;
- δ is physical retirement;
- Rt is economic obsolescence or inaccessible capacity.
Effective capacity is:
Keffective,t = Kt × Ut × At × Et
where U is utilisation, A is availability and E is software efficiency.
9.4 Learning curves
The requested learning relationship is:
Ct = C0(Qt ÷ Q0)−b
The implied learning rate is:
LR = 1 − 2−b
If cumulative inference output increases tenfold and quality-adjusted cost falls from 100 to 32:
0.32 = 10−b
b = −ln(0.32) ÷ ln(10)
b ≈ 0.495
Therefore:
LR = 1 − 2−0.495
LR ≈ 29.0%
Under this central case, each doubling of cumulative output reduces quality-adjusted cost by approximately 29%.
9.4.1 Learning-rate sensitivity
| Learning exponent b | Implied learning rate | Cost after four doublings |
|---|---|---|
| 0.10 | 6.7% | 75.8 |
| 0.20 | 12.9% | 57.4 |
| 0.30 | 18.8% | 43.5 |
| 0.40 | 24.2% | 33.0 |
| 0.50 | 29.3% | 25.0 |
| 0.60 | 34.0% | 18.9 |
Learning effects can be offset by capability escalation. If users migrate from inexpensive basic models to more compute-intensive reasoning and agentic systems, the cost of a fixed 2026 task may decline while average expenditure per user rises.
9.5 Central annual forecast, 2026–2031
9.5.1 Technology and infrastructure
All indices use 2026=100 unless otherwise specified.
| Variable | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 | 2026–2031 CAGR |
|---|---|---|---|---|---|---|---|
| Nominal GPU/accelerator price | 100 | 101 | 103 | 104 | 105 | 105 | 1.0% |
| Quality-adjusted accelerator price | 100 | 82 | 68 | 57 | 48 | 41 | −16.4% |
| HBM price per GB | 100 | 94 | 88 | 83 | 78 | 73 | −6.1% |
| Conventional DRAM price per GB | 100 | 95 | 90 | 86 | 82 | 78 | −4.9% |
| Advanced-packaging capacity | 100 | 130 | 165 | 205 | 245 | 285 | 23.3% |
| Data-centre CAPEX | 100 | 122 | 145 | 166 | 184 | 199 | 14.8% |
| AI inference demand | 100 | 180 | 310 | 500 | 760 | 1,100 | 61.6% |
| Quality-adjusted API price | 100 | 75 | 58 | 46 | 38 | 32 | −20.4% |
| Enterprise AI expenditure | 100 | 130 | 165 | 205 | 250 | 300 | 24.6% |
| Local-AI cost-to-income ratio | 100 | 95 | 89 | 84 | 80 | 76 | −5.3% |
The central result is not contradictory: API prices and quality-adjusted hardware costs can fall while total enterprise expenditure rises. Demand, integration depth, model capability, multimodality, agentic tools and governance can expand faster than unit costs decline.
9.5.2 US data-centre electricity
| Year | Central forecast | Low trajectory | High trajectory |
|---|---|---|---|
| 2026 | 330 TWh | 280 TWh | 380 TWh |
| 2027 | 390 TWh | 310 TWh | 480 TWh |
| 2028 | 450 TWh | 325 TWh | 580 TWh |
| 2029 | 515 TWh | 350 TWh | 690 TWh |
| 2030 | 585 TWh | 380 TWh | 810 TWh |
| 2031 | 660 TWh | 410 TWh | 930 TWh |
The official DOE range applies to 2028; values beyond 2028 are scenario extensions, not official federal projections. The central 2028 value of 450 TWh lies within the official 325–580 TWh range.
9.6 Price forecasts
9.6.1 GPU and accelerator prices
Nominal accelerator prices are expected to be substantially more rigid than quality-adjusted prices. Successive products may provide more memory and throughput while preserving or increasing headline prices.
| Forecast | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Nominal low | 100 | 95 | 91 | 87 | 84 | 80 |
| Nominal central | 100 | 101 | 103 | 104 | 105 | 105 |
| Nominal high | 100 | 110 | 121 | 133 | 146 | 161 |
| Quality-adjusted central | 100 | 82 | 68 | 57 | 48 | 41 |
The high nominal path corresponds to persistent scarcity, premium product migration, memory bottlenecks, tariffs and limited competition. The low path assumes strong competition, inventory normalisation and improved alternatives.
9.6.2 HBM
HBM’s price trajectory depends on two opposing forces:
- generational improvements and capacity expansion reduce price per bit;
- rapid demand, complex stacking, yield and capacity trade-offs sustain premiums.
Micron stated that the expansion of HBM creates a significant trade ratio against conventional DRAM capacity and continued to describe supply-demand conditions as tight. Micron Fiscal First Quarter 2026 Earnings Call – Micron – December/2025 — Micron HBM and DRAM supply disclosure.
| HBM price-per-GB index | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Low | 100 | 82 | 69 | 58 | 50 | 43 |
| Central | 100 | 94 | 88 | 83 | 78 | 73 |
| High | 100 | 112 | 120 | 125 | 128 | 130 |
9.6.3 DRAM
| DRAM price-per-GB index | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Low | 100 | 85 | 74 | 66 | 59 | 53 |
| Central | 100 | 95 | 90 | 86 | 82 | 78 |
| High | 100 | 108 | 113 | 116 | 118 | 120 |
The high case reflects wafer allocation toward HBM and advanced products, constraining conventional memory supply despite technological progress.
9.7 Advanced-packaging forecast
TSMC’s 2025 annual report described continuing development of CoWoS, InFO, SoIC and related three-dimensional integration technologies in response to AI demand. TSMC 2025 Annual Report – TSMC – 2026 — TSMC 2025 Annual Report.
| Advanced-packaging capacity index | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Low | 100 | 115 | 135 | 155 | 175 | 195 |
| Central | 100 | 130 | 165 | 205 | 245 | 285 |
| High | 100 | 150 | 215 | 300 | 400 | 515 |
Capacity growth does not guarantee equivalent growth in completed accelerator systems. HBM, substrates, interposers, optical networking, power equipment and final-server integration must expand simultaneously.
9.8 Data-centre CAPEX
9.8.1 Central trajectory
| Data-centre CAPEX index | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Low | 100 | 108 | 116 | 122 | 126 | 128 |
| Central | 100 | 122 | 145 | 166 | 184 | 199 |
| High | 100 | 145 | 195 | 255 | 325 | 400 |
The central path assumes continued expansion but declining annual growth as grid, construction, permitting and financing constraints become binding. The high path assumes a sustained infrastructure arms race, sovereign capacity programmes and rapid agentic demand.
9.8.2 CAPEX stock-flow relation
KDC,t+1 = KDC,t + CAPEXt − Depreciationt − Impairmentt
A rise in capital expenditure does not translate immediately into productive capacity because of:
- construction lead time;
- grid connection;
- accelerator delivery;
- cooling commissioning;
- software integration;
- customer ramp;
- stranded or underutilised assets.
9.9 AI inference demand
Inference demand should be measured in successful quality-adjusted tasks rather than raw tokens. Let:
Dinference,t = Userst × TasksPerUsert × ComputePerTaskt ÷ Efficiencyt
| Inference-demand index | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Low | 100 | 140 | 200 | 280 | 380 | 500 |
| Central | 100 | 180 | 310 | 500 | 760 | 1,100 |
| High | 100 | 250 | 550 | 1,100 | 2,000 | 3,400 |
Efficiency can increase demand through a rebound mechanism:
Total Compute = Cost per Task × Number of Tasks
If cost per task falls by 70% while task volume increases tenfold, total expenditure and physical compute can still rise substantially.
9.10 API price forecast
9.10.1 Quality-adjusted cost per successful task
| API price index | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Low-price path | 100 | 65 | 43 | 30 | 22 | 17 |
| Central | 100 | 75 | 58 | 46 | 38 | 32 |
| High-price path | 100 | 100 | 103 | 108 | 115 | 125 |
The central path assumes learning, routing, caching and competition. The high path reflects premium frontier migration, reasoning-token billing, priority charges, long-context multipliers and tight capacity.
9.10.2 Tariff decomposition
PAI,t = Pbase,t × Mmodel,t × Mlatency,t × Mcontext,t × Mregion,t + Ctools,t
A falling base price can coexist with a rising completed-task charge if the multipliers grow.
9.11 Enterprise expenditure
| Enterprise AI expenditure index | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Low | 100 | 115 | 132 | 150 | 168 | 185 |
| Central | 100 | 130 | 165 | 205 | 250 | 300 |
| High | 100 | 155 | 225 | 320 | 440 | 590 |
The central forecast implies 24.6% annual growth. Expenditure includes:
- licences;
- API usage;
- local and cloud infrastructure;
- data preparation;
- integration;
- cybersecurity;
- governance;
- compliance;
- human supervision;
- evaluation;
- switching and redundancy.
Unit-price deflation therefore does not imply budget deflation.
9.12 Sectoral adoption-cost forecasts
The table reports real total adoption-cost indices.
| Sector | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Microenterprises | 100 | 117 | 135 | 153 | 170 | 185 |
| SMEs | 100 | 122 | 148 | 176 | 205 | 235 |
| Large corporations | 100 | 130 | 164 | 202 | 243 | 285 |
| Banking | 100 | 135 | 170 | 207 | 245 | 285 |
| Insurance | 100 | 130 | 160 | 195 | 230 | 265 |
| Healthcare | 100 | 138 | 180 | 225 | 272 | 320 |
| Legal/professional | 100 | 120 | 140 | 160 | 178 | 195 |
| Manufacturing | 100 | 128 | 158 | 190 | 222 | 255 |
| Media | 100 | 125 | 150 | 176 | 200 | 225 |
| Telecommunications | 100 | 135 | 172 | 212 | 255 | 300 |
| Defence/critical infrastructure | 100 | 142 | 188 | 238 | 290 | 345 |
| Public administration | 100 | 126 | 155 | 188 | 220 | 250 |
| Education/research | 100 | 125 | 150 | 178 | 205 | 230 |
High-assurance sectors rise fastest because deployment depth, security, validation and redundancy expand faster than model-unit prices decline.
9.13 Local-AI affordability
The local-AI affordability ratio is:
LAAc,t = Annualised Local AI TCOc,t ÷ Median Disposable Incomec,t
| Local-AI affordability index | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Improving path | 100 | 88 | 77 | 67 | 58 | 50 |
| Central | 100 | 95 | 89 | 84 | 80 | 76 |
| Deteriorating path | 100 | 108 | 116 | 125 | 133 | 140 |
The central case improves slowly because quality-adjusted hardware costs fall but adequate-model requirements, memory, electricity and capability expectations rise.
9.14 Global-access inequality forecast
Using the Chapter 8 provisional GAI Gini of 0.193 as the 2026 analytical baseline:
| GAI Gini | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Convergence path | 0.193 | 0.185 | 0.178 | 0.171 | 0.165 | 0.160 |
| Central tiered-access path | 0.193 | 0.198 | 0.203 | 0.208 | 0.212 | 0.215 |
| Divergence path | 0.193 | 0.210 | 0.228 | 0.247 | 0.268 | 0.290 |
The central case assumes basic access diffuses but frontier infrastructure remains concentrated.
9.15 Monte Carlo framework
For each simulation b:
- Draw inference-demand growth.
- Draw capacity growth.
- Draw learning exponent.
- Draw energy price and grid-connection delay.
- Draw memory and packaging constraints.
- Draw geopolitical regime.
- Calculate capacity, price, expenditure and access.
- Repeat at least 50,000 times.
Illustrative parameter distributions:
| Parameter | Distribution | Central parameter |
|---|---|---|
| Annual inference growth | Lognormal | 55% |
| Advanced-packaging growth | Triangular | 15%, 25%, 45% |
| API learning exponent | Normal, truncated positive | 0.50 |
| Accelerator economic life | Triangular | 3, 4, 6 years |
| Data-centre power delay | Discrete | 0–4 years |
| Export-control shock | Bernoulli | Scenario-dependent |
| HBM supply shock | Bernoulli | Scenario-dependent |
| Electricity-price growth | Lognormal | 3% |
| Enterprise adoption | Logistic diffusion | Sector-dependent |
9.16 2031 forecast intervals
These intervals are model-based simulation ranges, not classical confidence intervals derived from long stationary histories.
| Variable | 2031 central | 80% interval | 95% interval |
|---|---|---|---|
| Nominal GPU price index | 105 | 82–145 | 65–190 |
| Quality-adjusted accelerator price | 41 | 25–65 | 15–100 |
| HBM price-per-GB index | 73 | 50–110 | 35–150 |
| DRAM price-per-GB index | 78 | 55–110 | 40–145 |
| Advanced-packaging capacity | 285 | 220–360 | 170–470 |
| Data-centre CAPEX | 199 | 155–270 | 120–380 |
| US data-centre electricity | 660 TWh | 480–800 TWh | 390–930 TWh |
| AI inference demand | 1,100 | 600–1,900 | 350–3,200 |
| Quality-adjusted API price | 32 | 20–50 | 12–80 |
| Enterprise AI expenditure | 300 | 220–410 | 160–600 |
| Local-AI affordability index | 76 | 55–105 | 40–140 |
| GAI Gini | 0.215 | 0.185–0.250 | 0.160–0.300 |
9.17 Scenario probabilities
The scenario probabilities are structured judgments derived from:
- current investment commitments;
- concentration documented in previous chapters;
- official packaging and memory statements;
- power forecasts;
- export-control developments;
- observed pricing segmentation;
- model-efficiency trends;
- historical semiconductor cyclicality.
They are not objective frequencies. The scenarios are partly overlapping: financialised pricing can coexist with oligopoly, fragmentation or scarcity.
| Scenario | Central judgement | Plausible range |
|---|---|---|
| A — Abundant Compute | 20% | 15–25% |
| B — Managed Oligopoly | 35% | 30–40% |
| C — Persistent Scarcity | 18% | 15–25% |
| D — Fragmented AI World | 17% | 15–25% |
| E — AI Utility and Financialised Access | 10% as dominant regime | 10–20% |
The central point estimates sum to 100% only to construct a decision-weighted baseline. The plausible ranges do not need to sum to 100%.
9.18 Scenario A — Abundant Compute
9.18.1 Assumptions
- advanced-packaging capacity expands rapidly;
- HBM4 and later generations achieve strong yields;
- foundry and memory competition increases;
- grid and generation projects arrive on schedule;
- open-weight models remain competitive;
- inference software produces strong learning effects;
- export restrictions remain limited;
- cloud providers compete aggressively;
- small and specialised models absorb routine demand.
9.18.2 Annual trajectory
| Indicator | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Quality-adjusted compute price | 100 | 75 | 58 | 45 | 35 | 28 |
| HBM price-per-GB | 100 | 82 | 69 | 58 | 50 | 43 |
| Packaging capacity | 100 | 150 | 215 | 300 | 400 | 515 |
| API successful-task price | 100 | 65 | 43 | 30 | 22 | 17 |
| Enterprise expenditure | 100 | 125 | 155 | 185 | 215 | 245 |
| Global GAI mean | 56.4 | 60 | 64 | 68 | 72 | 76 |
| GAI Gini | 0.193 | 0.185 | 0.178 | 0.171 | 0.165 | 0.160 |
9.18.3 Consequences
- Companies obtain positive ROI from more use cases.
- SMEs gain access to advanced functions without large CAPEX.
- Household and student affordability improves.
- Low- and middle-income countries benefit from mobile and small-model diffusion.
- Frontier providers face margin pressure.
- Total electricity demand can still rise because lower prices generate higher usage.
9.18.4 Early-warning indicators
- packaging lead times fall below normal procurement cycles;
- HBM contract premiums decline;
- accelerator inventories rise;
- API prices fall without tighter limits;
- local models close capability gaps;
- grid-connection queues shorten.
9.18.5 Invalidation events
- major semiconductor disruption;
- sustained HBM shortages;
- severe power constraints;
- widening export controls;
- frontier workloads become much more compute-intensive than efficiency gains.
9.19 Scenario B — Managed Oligopoly
9.19.1 Assumptions
- physical supply expands;
- advanced foundry, HBM, packaging and cloud control remains concentrated;
- providers maintain differentiated model tiers;
- competition reduces some prices but preserves rents;
- enterprise contracts and ecosystem lock-in increase;
- interoperability remains incomplete;
- basic models become cheap while frontier access remains premium.
9.19.2 Annual trajectory
| Indicator | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Quality-adjusted compute price | 100 | 88 | 78 | 70 | 64 | 59 |
| HBM price-per-GB | 100 | 98 | 94 | 90 | 86 | 82 |
| Packaging capacity | 100 | 128 | 160 | 195 | 230 | 265 |
| API successful-task price | 100 | 82 | 70 | 61 | 55 | 50 |
| Enterprise expenditure | 100 | 132 | 170 | 215 | 265 | 320 |
| Global GAI mean | 56.4 | 58 | 60 | 62 | 64 | 66 |
| GAI Gini | 0.193 | 0.197 | 0.202 | 0.207 | 0.211 | 0.215 |
9.19.3 Consequences
- Large companies obtain volume discounts and reserved capacity.
- SMEs remain dependent on bundled subscriptions.
- Consumer free tiers improve but remain below frontier capability.
- Platform vendors capture a growing share of enterprise technology budgets.
- Countries without cloud regions or bargaining scale remain price takers.
9.19.4 Early-warning indicators
- gross margins remain elevated despite capacity expansion;
- multi-year supply agreements dominate;
- enterprise prices remain opaque;
- provider-specific tools deepen switching costs;
- priority and long-context premiums expand.
9.19.5 Invalidation events
- successful open standards materially reduce switching;
- new foundry or accelerator competitors capture substantial share;
- regulators impose structural interoperability;
- oversupply causes severe price competition.
9.20 Scenario C — Persistent Scarcity
9.20.1 Assumptions
- inference demand exceeds forecasts;
- HBM, packaging, substrates and power remain constrained;
- data-centre grid connections are delayed;
- frontier-model compute per task rises;
- equipment prices and financing remain elevated;
- supply expansion arrives more slowly than demand.
9.20.2 Annual trajectory
| Indicator | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Nominal accelerator price | 100 | 110 | 121 | 133 | 146 | 161 |
| HBM price-per-GB | 100 | 112 | 120 | 125 | 128 | 130 |
| Packaging capacity | 100 | 115 | 135 | 155 | 175 | 195 |
| API successful-task price | 100 | 105 | 110 | 115 | 120 | 125 |
| Enterprise expenditure | 100 | 150 | 215 | 300 | 410 | 540 |
| Global GAI mean | 56.4 | 55.5 | 54.5 | 53.5 | 52.5 | 51 |
| GAI Gini | 0.193 | 0.205 | 0.218 | 0.230 | 0.240 | 0.250 |
9.20.3 Consequences
- Companies ration AI use and prioritise highest-margin tasks.
- Heavy agentic workflows become prohibitively expensive for SMEs.
- Students and researchers rely increasingly on inferior tiers.
- Wealthy countries secure capacity through long-term contracts.
- Used local hardware retains high resale values.
- Grid and electricity access become competitive differentiators.
9.20.4 Early-warning indicators
- HBM remains sold out more than twelve months forward;
- packaging lead times stop improving;
- data-centre projects are delayed by power;
- provider rate limits tighten;
- reserved-capacity premiums increase;
- used accelerator prices remain above reference values.
9.20.5 Invalidation events
- major inference-efficiency breakthrough;
- demand growth slows materially;
- multiple new capacity sources achieve high yields;
- small models replace frontier systems across most workloads.
9.21 Scenario D — Fragmented AI World
9.21.1 Assumptions
- export controls broaden;
- countries impose data-localisation and sovereignty rules;
- technology blocs adopt incompatible accelerators and software;
- cross-border cloud service becomes politically contingent;
- sanctions and foreign-currency restrictions intensify;
- national subsidies duplicate infrastructure.
9.21.2 Annual trajectory
| Indicator | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Global average compute price | 100 | 105 | 112 | 118 | 125 | 132 |
| Cross-bloc price dispersion | 100 | 130 | 165 | 205 | 250 | 300 |
| Duplicated infrastructure CAPEX | 100 | 140 | 190 | 245 | 305 | 370 |
| Global API interoperability | 100 | 90 | 78 | 67 | 58 | 50 |
| Global GAI mean | 56.4 | 55.8 | 55.2 | 54.8 | 54.3 | 54 |
| GAI Gini | 0.193 | 0.210 | 0.230 | 0.250 | 0.270 | 0.290 |
9.21.3 Consequences
- Allied countries receive different hardware and cloud access.
- Global companies operate duplicate AI stacks.
- Compliance and switching costs rise.
- Countries with domestic ecosystems gain sovereignty but sacrifice some scale efficiency.
- Scientific collaboration and model reproducibility decline.
- Smaller countries face pressure to align with a technology bloc.
9.21.4 Early-warning indicators
- new destination-based semiconductor restrictions;
- regional model-licensing rules;
- sovereign-cloud mandates;
- incompatible national AI standards;
- cross-border model withdrawal;
- rising duplication of national compute clusters.
9.21.5 Invalidation events
- multilateral technology agreements;
- effective interoperable standards;
- relaxation of export controls;
- broad availability of competitive hardware from multiple jurisdictions.
9.22 Scenario E — AI Utility and Financialised Access
9.22.1 Assumptions
- compute capacity is treated as a reservable strategic service;
- prices vary by latency, model, context, geography and scarcity;
- reserved, priority, batch and flex markets deepen;
- providers sell forward capacity;
- large users hedge future compute costs;
- spot capacity becomes dynamically priced;
- completed-task and outcome pricing expand.
9.22.2 Annual trajectory
| Indicator | 2026 | 2027 | 2028 | 2029 | 2030 | 2031 |
|---|---|---|---|---|---|---|
| Standard compute-price index | 100 | 92 | 88 | 85 | 83 | 82 |
| Priority-price multiplier | 1.5× | 1.7× | 2.0× | 2.3× | 2.6× | 3.0× |
| Batch discount | 50% | 52% | 55% | 58% | 60% | 62% |
| Reserved-capacity share | 100 | 135 | 180 | 230 | 285 | 350 |
| Enterprise expenditure | 100 | 140 | 185 | 235 | 290 | 350 |
| Global GAI mean | 56.4 | 57 | 58 | 59 | 59.5 | 60 |
| GAI Gini | 0.193 | 0.200 | 0.208 | 0.218 | 0.227 | 0.235 |
9.22.3 Consequences
- Large firms hedge capacity and obtain predictable service.
- SMEs pay volatile on-demand prices.
- Latency-sensitive sectors pay substantial premiums.
- Batch research and back-office tasks become cheaper.
- Households receive low-cost standard service but expensive frontier priority.
- Compute-finance products create counterparty and concentration risks.
9.22.4 Early-warning indicators
- expansion of reserved and priority tiers;
- compute-capacity marketplaces;
- forward contracts and minimum-spend commitments;
- dynamic congestion pricing;
- task and outcome billing;
- financial institutions offering compute hedging.
9.22.5 Invalidation events
- abundant capacity eliminates scarcity premiums;
- regulation prohibits dynamic discrimination;
- model efficiency makes reservation unnecessary;
- compute becomes too heterogeneous for standard contracts.
9.23 Company-level consequences by scenario
| Company type | A | B | C | D | E |
|---|---|---|---|---|---|
| Microenterprise | Strong access improvement | Subscription dependence | Exclusion risk | Regional service gaps | Volatile usage bills |
| SME | Lower entry cost | Vendor lock-in | Margin pressure | Duplicate compliance | Need for capacity budgeting |
| Large corporation | Broad deployment | Bargaining advantage | Reserved-capacity race | Multiple regional stacks | Hedging and portfolio optimisation |
| Bank/insurer | Lower analytical cost | Concentrated supplier risk | High priority premiums | Data-localisation burden | Financialised capacity contracts |
| Healthcare | Wider validated use | Enterprise-provider dependence | High-assurance scarcity | National-system divergence | Priority service becomes essential |
| Manufacturer | Cheap edge AI | Platform integration | Hardware rationing | Supply-chain fragmentation | Long-term compute procurement |
| Research institution | Wider experimentation | Tiered frontier access | Publication inequality | Restricted collaboration | Batch discounts but priority disadvantage |
| Government | Broader public services | Sovereignty concern | Capacity rationing | National stacks | Strategic compute reserve |
9.24 Household and distributional consequences
| Scenario | Household effect | Student/researcher effect | Inequality effect |
|---|---|---|---|
| A | Falling effective prices | Broad advanced access | Convergence |
| B | Cheap basic access, premium frontier | Persistent tiering | Gradual divergence |
| C | Higher prices and tighter limits | Severe exclusion | Strong divergence |
| D | Access depends on nationality and bloc | International collaboration declines | Geographic divergence |
| E | Basic access cheap; priority expensive | Time-sensitive users disadvantaged | Price discrimination increases |
9.25 Early-warning dashboard
| Indicator | Abundance signal | Scarcity/fragmentation signal |
|---|---|---|
| HBM lead time | Falling below six months | Above twelve months |
| Packaging utilisation | Normalising | Persistently near effective maximum |
| GPU street premium | Below 5% | Above 25% |
| Used-market premium | Negative/normal depreciation | Persistent positive premium |
| API list prices | Broad decline | Frontier and priority increase |
| Batch discount | Expands | Contracts |
| Data-centre power queue | Shortens | Multi-year delay |
| Memory producer margins | Normalise | Remain exceptionally high |
| Cloud-region expansion | Broad geographic diffusion | Bloc-specific concentration |
| Export controls | Stable/narrow | Broaden by destination and capability |
| Enterprise RPO | Stable | Rapid long-term reservation |
| Local-model quality gap | Narrows | Widens |
| GAI Gini | Declines | Rises |
| Research-compute distribution | Broadens | Concentrates |
9.26 Model validation and annual updating
The forecast should be updated every quarter for market variables and annually for country and social variables.
Forecast error is:
et = Xt − X̂t
Recommended metrics are:
MAE = (1 ÷ n)Σ|et|
RMSE = √[(1 ÷ n)Σet2]
MAPE = (100 ÷ n)Σ|et ÷ Xt
Interval coverage is:
Coverage = Number of Observations Inside Interval ÷ Total Observations
If an 80% interval contains materially fewer than 80% of subsequent observations, the model is overconfident and its variance assumptions must be revised.
9.27 Final forecast conclusions
- Nominal accelerator prices are likely to fall much more slowly than quality-adjusted prices.
- HBM and advanced packaging remain the most important semiconductor bottlenecks through at least the early forecast period.
- Advanced-packaging capacity could nearly triple by 2031 in the central case, yet demand can still exceed supply.
- US data-centre electricity demand could reach approximately 660 TWh by 2031 in the central scenario, with a very wide model-based range.
- Inference demand is likely to grow substantially faster than physical capacity because lower prices generate new uses.
- Quality-adjusted API costs can fall by approximately two-thirds while total enterprise AI expenditure triples.
- Enterprise expenditure will increasingly shift from experimental licences to integration, governance, security, agents and reserved capacity.
- Local-AI affordability improves only slowly in the central case because the threshold of adequate capability continues to rise.
- Global access inequality is more likely to increase modestly than collapse, even if basic AI becomes widely available.
- Scenario A produces broad convergence but requires simultaneous success in capacity, competition, energy and model efficiency.
- Scenario B is the central structural case: supply expands while bargaining power and frontier access remain concentrated.
- Scenario C remains plausible because demand can expand faster than memory, packaging, power and construction.
- Scenario D would make geography and geopolitical alignment direct determinants of AI capability.
- Scenario E is already partially visible in reserved, priority, batch and flexible service tiers, although a fully financialised market is not inevitable.
- Prediction intervals are necessarily wide because technology, demand and policy are endogenous. Narrow intervals would be statistically misleading.
- The five scenarios are not mutually exclusive in practice. The world can exhibit abundance in small models, oligopoly in frontier systems, scarcity in HBM, fragmentation across geopolitical blocs and financialised pricing for priority capacity at the same time.
- The decisive monitoring variable is not the sticker price of one GPU or one million tokens. It is the quality-adjusted cost of a successfully completed workload, available at the required latency, privacy, reliability and geographic location.
- The dominant five-year risk is tiered abundance: large quantities of inexpensive basic intelligence coexist with scarce, costly and concentrated frontier capability.
CHAPTER 10 — Policy Responses, Strategic Options and Final Conclusions
10.1 Final policy thesis
The evidence developed throughout this report supports neither technological fatalism nor the stronger claim that all AI hardware has become uniformly more expensive after adjusting for performance. The defensible conclusion is more specific. Raw computational capability continues to improve rapidly, and quality-adjusted prices can decline when workloads fit the memory, software and reliability constraints of newer systems. At the same time, economically useful access to AI depends on more than arithmetic throughput. Accelerator memory, memory bandwidth, advanced packaging, power availability, data-centre capacity, model quality, context length, software compatibility, privacy, concurrency and access to complementary expertise form a joint production system. A fall in price per theoretical operation therefore does not guarantee a fall in the cost of completing a real workload. Users who require large memory capacity, sustained inference, high reliability, private deployment or frontier-level model quality can face rising absolute expenditure even while price per TFLOP declines. This distinction explains why households, students and small organisations can experience genuine exclusion without contradicting semiconductor learning curves. The central policy problem is not a universal shortage of every chip. It is the concentration of strategically complementary capabilities in a limited number of firms, jurisdictions, fabrication processes, packaging facilities, memory suppliers, cloud regions and software ecosystems.
The appropriate policy objective is consequently not to make every organisation own frontier-scale infrastructure. That would waste capital, electricity and specialised labour. The objective should be to ensure contestable, resilient and affordable access to a workload-appropriate capability floor. Public intervention is justified when a combination of indivisibility, high fixed costs, knowledge spillovers, market concentration, switching costs, public-interest privacy requirements and unequal purchasing power causes socially valuable users to receive less compute than produces the highest social return. Intervention must nevertheless avoid converting scarcity into permanent subsidy dependence or protecting inefficient national suppliers from competition. The recommended architecture combines shared infrastructure, targeted demand subsidies, interoperability, transparent metering, competitive procurement, repair and refurbishment, open-weight research, grid investment and narrowly conditioned industrial support. Supply subsidies should be milestone-based, technology-neutral where possible, competitively awarded and subject to clawbacks. Access subsidies should target students, researchers, SMEs and public-interest institutions rather than reduce prices indiscriminately for the largest buyers. Competition enforcement should be evidence-driven: high concentration or margins establish a reason to investigate but do not independently prove exclusionary conduct, unlawful coordination or excessive pricing.
European policy already provides a foundation. The Commission reports an ambition to mobilise EUR 200 billion for AI, including EUR 20 billion for as many as five AI gigafactories, while an expanding network of AI Factories is intended to serve research, industry, start-ups and SMEs. AI Continent – European Commission – 2026 — official programme page. EuroHPC has separately described investment approaching EUR 2 billion in AI Factories; individual projects demonstrate the scale involved, including the EUR 290 million IT4LIA system, financed equally by EuroHPC and Italy. EuroHPC JU signs contract to boost AI capabilities at the IT4LIA AI Factory – EuroHPC Joint Undertaking – April 2026 — verified project announcement. These programmes should be treated as infrastructure platforms rather than one-time machine purchases. Their economic performance must be measured through utilisation, queue time, workload completion, geographic distribution, research output, SME survival and the share of capacity reaching users who could not otherwise purchase equivalent service.
10.2 Governing principles
A defensible access policy should satisfy seven tests:
- Additionality: public expenditure must add capacity, competition, resilience or access that the market would not supply on comparable terms.
- Capability targeting: eligibility must depend on workload requirements rather than parameter count or marketing labels.
- Contestability: funded infrastructure must support multiple frameworks, models and providers.
- Portability: users must be able to export data, prompts, embeddings, evaluation records, fine-tuning artefacts and operational metadata in documented formats.
- Cost transparency: metering must separate model inference, storage, retrieval, networking, tool use, reserved capacity and priority premiums.
- Distributional accountability: programmes must report who receives compute, not merely aggregate utilisation.
- Sunset and review: subsidies require expiry dates, counterfactual evaluation and recovery provisions when recipients fail investment, access or pricing obligations.
The Data Act supplies an existing European foundation for cloud switching and interoperability and has applied since 12 September 2025. Data Act Explained – European Commission – December 2025 — official explanation. The Digital Markets Act complements competition law by imposing obligations and prohibitions on designated gatekeepers, but its application to AI-related self-preferencing must follow the Act’s defined scope and legal tests rather than assume that every vertically integrated AI supplier is automatically covered. The Digital Markets Act – European Commission – 2026 — official DMA portal.
10.3 Policy-design matrix: mandate, recipients, cost and timing
All amounts below are planning estimates, not enacted appropriations, unless expressly identified as an official programme figure.
| Policy instrument | Problem addressed | Legal or institutional basis | Implementing authority | Principal target | Estimated public cost | Implementation period |
|---|---|---|---|---|---|---|
| Public compute infrastructure | Indivisible capital costs; private under-provision for research and public-interest workloads | National research, industrial and digital-infrastructure statutes; in the EU, EuroHPC, Digital Europe and research competences | Digital ministry, science ministry, national HPC body, EuroHPC | Universities, public research, SMEs, hospitals, public administration | EUR 100m–1bn per national node; EUR 2bn–10bn for a multi-node system | 2–5 years |
| European AI cloud and gigafactories | Dependence on non-European frontier capacity; insufficient scale | EuroHPC Regulation, Digital Europe, Horizon Europe, InvestAI-compatible finance, applicable State-aid rules | Commission, EuroHPC JU, EIB, Member States | European research, industry, start-ups and sovereign workloads | EUR 10bn–25bn additional public and risk-sharing capital over five years; official wider ambition includes EUR 20bn for up to five gigafactories | 3–6 years |
| Shared university compute | Fragmented procurement and low utilisation of departmental systems | University and research-funding mandates | Research councils and university consortia | Students, doctoral candidates and researchers | EUR 25m–300m per consortium plus annual operating expenditure of approximately 10–20% of CAPEX | 1–4 years |
| Student compute vouchers | Income-based exclusion from capable AI | Education and equal-opportunity legislation | Education ministries, universities, grant agencies | Low-income students; disability and language-access users | EUR 250–1,000 per eligible student/year | 6–18 months |
| Research compute grants | Early-career and independent researchers lack purchasing power | Public research-grant authority | Research councils and foundations | Researchers without institutional clusters | EUR 2,000–20,000 per researcher/year; larger competitive awards for high-compute projects | 6–18 months |
| Investment tax credits | Socially useful capacity may be delayed by financing constraints | Tax code; EU State-aid compatibility where applicable | Treasury and tax authority | Data centres, packaging, memory, energy-efficient compute | Credit of 10–25% of qualifying incremental investment, with a jurisdictional fiscal cap | 1–3 years |
| Cooperative purchasing | SMEs and universities face weak bargaining power and duplicative procurement | Procurement and cooperative statutes | Central purchasing bodies or sector consortia | SMEs, schools, municipalities, universities | EUR 1m–10m national administration; EUR 20m–50m for an EU-scale platform | 6–24 months |
| Open-weight model programme | Dependence on closed services; weak local-language coverage | Research, cultural and digital-innovation mandates | Research agencies, public laboratories, universities | Researchers, SMEs, public services, minority-language communities | EUR 25m–250m nationally or EUR 500m–2bn at EU scale over five years | 2–5 years |
| Interoperability and portability rules | Cloud, data and workflow lock-in | EU Data Act; national data and competition law | Commission, national competent authorities, standards bodies | All enterprise and public-sector customers | Public administration EUR 20m–100m at EU scale; private compliance costs additional | 1–3 years |
| Standardised AI price disclosure | Complex tariffs obscure effective prices and switching costs | Consumer, contract, procurement and sectoral transparency law | Consumer authority, digital regulator, procurement agencies | Consumers, SMEs and public buyers | EUR 5m–20m per major jurisdiction plus provider compliance cost | 12–24 months |
| Competition enforcement unit | Information asymmetry and technically complex vertical conduct | TFEU Articles 101 and 102; national equivalents | DG Competition and national competition authorities | Markets for accelerators, cloud, models and distribution | EUR 50m–200m over five years across a major regional authority network | Immediate; permanent |
| Enhanced merger scrutiny | Acquisitions can remove nascent competitors or close complementary inputs | EU Merger Regulation and national merger law | Commission and national authorities | Semiconductor, cloud, AI and data transactions | Included in enforcement staffing; specialist analysis EUR 5m–30m/year | Immediate |
| Restriction of proven self-preferencing | Vertically integrated platforms may disadvantage rivals | DMA where legally applicable; competition law elsewhere | Commission and digital regulators | Business users and competing AI services | Enforcement cost included above | 1–3 years per proceeding |
| AI-aware public procurement | Public contracts may create lock-in or reward opaque pricing | Procurement directives and national procurement law | Central purchasing bodies and contracting authorities | Public administrations, schools, hospitals | Administrative reform EUR 10m–100m; underlying purchases remain programme expenditure | 1–3 years |
| Grid and low-carbon generation investment | Power and interconnection constrain data-centre expansion | Energy, network and planning law | Energy ministry, regulator and grid operators | Economy-wide users, including data centres | EUR 5bn–50bn nationally or EUR 50bn–250bn regionally over five years; only a fraction attributable to AI | 3–10 years |
| Semiconductor and packaging support | Advanced fabrication and packaging are geographically concentrated | European Chips Act; national industrial legislation | Commission, Member States, development banks | Foundries, packaging, equipment, substrates | EUR 5bn–40bn per major jurisdictional programme | 4–10 years |
| HBM diversification | Limited qualified supply and long capacity cycles | Industrial, R&D and State-aid frameworks | Industry ministry, research agencies, development banks | Memory fabrication, packaging and materials | EUR 2bn–15bn per programme | 4–8 years |
| Right to repair and longer support | Short hardware life raises effective access costs and e-waste | Consumer, ecodesign and repair legislation | Consumer authorities and product regulators | Households, schools, SMEs and refurbishers | EUR 20m–100m administration; manufacturer transition costs additional | 2–5 years |
| Verified refurbished market | Quality uncertainty, fraud and absent warranties weaken secondary supply | Consumer safety, warranty and certification law | Standards body, consumer authority, accredited certifiers | Low-income users, schools, SMEs | EUR 25m–150m setup, testing network and optional guarantee reserve | 1–3 years |
| Educational and research AI entitlement | Adequate AI may become a necessary educational input | Education and research law; annual budget legislation | Education ministries and research councils | Students, teachers and accredited researchers | EUR 0.5bn–3bn/year for a large jurisdiction, depending on coverage | 1–3 years |
| International Compute Development Fund | Countries with weak grids, foreign-exchange constraints or absent cloud regions cannot finance access | IBRD, IDA and regional-development-bank mandates; donor agreements | World Bank, regional development banks, ITU-linked partners and national governments | Low- and lower-middle-income economies | USD 10bn–30bn over five years | 2–7 years |
Table metadata: Unit: programme expenditure or benefit level. Currency: 2026 EUR unless USD is expressly stated. Price basis: real planning cost, excluding private co-finance unless noted. Reference year: 2026. Geographic scope: EU and representative national or international programmes. Sources: official programme records cited in this chapter plus author estimates. Method: benchmark scaling from official compute projects, programme design and beneficiary counts. Uncertainty: generally ±50–100% because facility configuration, energy works, financing structure and beneficiary coverage remain unspecified.
10.4 Policy performance, safeguards and feasibility
| Instrument | Expected benefit | Mandatory KPI | Principal unintended consequence | Distortion-control mechanism | Feasibility |
|---|---|---|---|---|---|
| Public compute | Access to workloads too large or sensitive for ordinary users | Successful jobs; queue time; utilisation; cost per completed workload; share allocated to underserved users | Political allocation, low utilisation, rapid obsolescence | Independent capacity auction; access quotas; external evaluation; depreciation reserve | High, if attached to existing HPC institutions |
| EU AI cloud | Strategic capacity and cross-border scale | Delivered accelerator-hours; workload portability; regional latency; SME/research share | Duplication, national capture, subsidy race | Competitive site selection; cross-border access; open interfaces; clawbacks | Medium–high |
| University sharing | Higher utilisation and lower unit cost | Active institutions; job success; publications; cost per research output | Dominance by large universities | Ring-fenced capacity and transparent queue rules | High |
| Vouchers | Immediate affordability improvement | Redemption; user income distribution; learning or research outcomes | Providers capture subsidy through higher prices | Multiple-provider eligibility; price ceilings; randomised or phased evaluation | High |
| Tax credits | Accelerated investment | Incremental capacity and verified investment additionality | Windfall for already-planned projects | Baseline tests, caps, expiry and recapture | Medium |
| Cooperative purchasing | Lower prices and contract complexity | Discount against comparable standalone contracts; switching rate | Excess standardisation around one vendor | Multi-lot procurement and multi-provider frameworks | High |
| Open-weight support | Local control, research spillovers and language inclusion | Reproducible benchmarks; adoption; language coverage; documented safety testing | Funding low-use models; release-related misuse | Milestones, model cards, security review and usage evaluation | High |
| Portability | Lower switching costs and stronger competition | Migration time, egress cost and successful export rate | Superficial formal compliance | Conformance testing and machine-readable export standards | High |
| Price disclosure | Comparable effective costs | Percentage of offers reporting standard workload prices | Gaming benchmark workloads | Multiple representative workloads and audit rights | High |
| Competition unit | Better ability to identify foreclosure and exclusion | Investigations completed; remedies; measured customer savings | Over-deterrence of integration and investment | Effects-based assessment and judicial review | High |
| Merger scrutiny | Preservation of future competition and input access | Cases reviewed; remedies monitored; post-merger price and innovation outcomes | Delayed efficient transactions | Deadlines, transparent theories of harm and targeted remedies | High |
| Self-preferencing controls | Neutral access where discrimination is proven | Ranking/access parity; rejection and latency differentials | Forced equivalence between non-equivalent services | Technical equivalence tests and security exceptions | Medium–high |
| Procurement reform | Lower lock-in and life-cycle cost | Open-interface compliance; total switching cost; supplier diversity | Complex tenders disadvantage SMEs | Standard clauses, lotting and procurement assistance | High |
| Grid investment | Removes physical constraint and limits price shocks | Connection lead time; firm capacity; emissions and reliability | Costs socialised while benefits concentrate | Beneficiary connection charges and location signals | Medium, because of long lead times |
| Semiconductor support | Greater resilience and technological capacity | Qualified output, yield, private co-investment and customer diversity | International subsidy race and stranded plants | Milestone disbursement, open access and clawbacks | Medium |
| HBM diversification | Reduced single-point dependency | Qualified HBM capacity, yield and supplier count | Technology becomes obsolete before scale | Staged funding linked to customer qualification | Medium–low |
| Repair rights | Lower lifetime cost and e-waste | Repairability score; parts availability; device life | Safety or cybersecurity risks from unsupported components | Qualified repair rules and firmware-security obligations | High |
| Refurbished certification | Trustworthy lower-cost supply | Failure rate, warranty claims, resale price and device life | Certification burdens small refurbishers | Tiered fees and public testing support | High |
| AI entitlement | Minimum educational and research capability | Eligible-user access, task success and attainment gap | Vendor dependence; displacement of teaching | Provider plurality, privacy requirements and pedagogy evaluation | Medium–high |
| International fund | Infrastructure and capability in underserved countries | Additional reliable compute, trained users, local-language coverage and public-service outcomes | Debt burden, imported systems without local capacity | Grant-heavy finance, maintenance reserves and local training | Medium |
Table metadata: Unit: qualitative feasibility plus programme-specific KPIs. Currency and price basis: not applicable; costs are in the preceding table. Reference year: assessment as of August 2026. Scope: international, EU and national implementation. Source: official legal and programme sources cited herein; analytical assessment by this report. Method: institutional-feasibility and market-failure screening. Uncertainty: qualitative, approximately one feasibility category in either direction.
Public compute should operate as a capacity wholesaler with public-interest obligations, not as a permanently free retail cloud. Universities, accredited researchers, start-ups and privacy-sensitive public institutions should receive an initial allocation based on project merit and inability to acquire equivalent resources. Larger commercial users should normally pay a transparent cost-recovery tariff. Allocation should combine reserved public-interest capacity, competitive grants and spot access for otherwise idle resources. The accounting system must publish the fully loaded cost per accelerator-hour and per completed reference workload, including energy, maintenance, staff and depreciation. An apparently inexpensive public system can otherwise conceal low utilisation or unfunded replacement liabilities.
The EuroHPC AI Factory call provides a useful empirical scale. The Union made up to EUR 400 million available for new or upgraded AI-optimised supercomputers, supporting total acquisition costs of as much as EUR 800 million, with further Horizon Europe support for AI Factory services. EuroHPC Joint Undertaking launches AI Factories calls – EuroHPC Joint Undertaking – September 2024 — official call announcement. These amounts show why shared procurement is economically preferable to duplicating underutilised departmental clusters. They do not, however, prove that every funded installation generates net social benefit. Publication of queue distributions, job-failure rates, workload outcomes and capacity allocation is essential.
10.4.2 Vouchers and access entitlements
Compute vouchers should be denominated in quality-adjusted service units rather than tied to a named supplier. A student entitlement could comprise a protected baseline of inference, document analysis, coding and accessibility services, with higher allocations for technical courses or disability accommodation. Research grants should support verified workloads, including local deployment where privacy, intellectual property or research reproducibility justifies it. Voucher design should avoid a simple reimbursement percentage, which favours users who can first finance the remaining expenditure. Prepaid credits, institutional brokerage and income-sensitive co-payments are more inclusive.
A voucher programme should be suspended or redesigned when more than 25% of its value is captured by price increases relative to untreated customers, when fewer than three substitutable providers satisfy the standard, or when measured educational or research outcomes do not improve after two evaluation cycles. AI access should become an entitlement only to a minimum adequate capability, not an unlimited right to the most expensive frontier model. The entitlement must remain technologically neutral and capable of being fulfilled by shared public systems, competitive cloud services, approved open-weight deployments or hybrid delivery.
10.4.3 Competition, portability and pricing transparency
Competition enforcement must distinguish bottleneck control from competitive success. Authorities should collect transaction-level evidence on discounts, capacity reservations, supply refusals, cloud credits, egress charges, interoperability failures, ranking, bundling and discrimination between affiliated and independent downstream services. A high HHI or margin may justify scrutiny but cannot establish an infringement without a defined market, a theory of harm and evidence of effects.
Merger review should examine complementary assets rather than only horizontal overlap. A small acquisition may materially affect competition when the target controls a compiler, model-optimisation layer, dataset, interconnect technology, capacity broker or distribution channel required by competing providers. The Commission is reviewing its merger guidelines to maintain an evidence-based framework adapted to changing competitive conditions. Review of the Merger Guidelines – European Commission, DG Competition – 2026 — official review page.
Standardised price disclosure should require providers to report:
- effective price per million input and output tokens;
- cached-input treatment;
- context-length and reasoning premiums;
- storage, retrieval and tool-use charges;
- reserved-capacity commitments;
- service-level guarantees;
- egress and migration costs;
- prices for a set of audited reference workloads;
- model-version retention rules;
- unilateral price-change notice periods.
The disclosure standard should not impose a regulated price. Its purpose is to expose the effective total price and reduce information-driven lock-in.
10.4.4 Repair, refurbishment and equipment life
Repair and verified refurbishment are distribution policies as well as environmental measures. They expand the secondary supply of capable hardware, reduce the annualised ownership cost and allow schools or small organisations to obtain higher-memory systems than new-device budgets permit. Certification should test memory, thermal stability, storage health, power delivery, firmware integrity and sustained AI workload performance. Certified devices should carry a minimum warranty and disclose prior enterprise or high-duty-cycle use.
The EU Directive promoting repair was adopted in June 2024 and applies through national implementation from 31 July 2026. Directive on Repair of Goods – European Commission – June 2024 — official directive portal. AI-relevant computers and components should be incorporated through product-specific repairability, parts and firmware-support rules where the legal scope permits. Requirements must not oblige manufacturers to support unsafe configurations indefinitely, but firmware withdrawal should not become an artificial method of shortening useful life.
10.4.5 Semiconductor, packaging, memory and energy policy
Industrial subsidies should concentrate on bottlenecks that plausibly remain binding after demand and technology uncertainty are considered. Advanced packaging, HBM qualification, substrates, interposers, power electronics and test capacity may generate greater near-term resilience per public euro than attempting to duplicate every layer of leading-edge fabrication. The European Chips Act entered into force in September 2023 and expressly targets design, manufacturing and advanced-packaging capacity, resilience and reduced dependencies. European Chips Act – European Commission – September 2023 — official policy page.
Subsidies must include:
- measurable capacity and yield milestones;
- private co-investment;
- non-discriminatory customer access;
- supply-continuity obligations during shortages;
- restrictions on relocating subsidised assets;
- repayment for non-delivery;
- workforce-development commitments;
- disclosure of other public assistance;
- environmental and grid-connection conditions.
Grid investment is a prerequisite rather than an AI subsidy. Data centres should pay cost-reflective connection charges and face locational signals where capacity is scarce. Public authorities should not reserve subsidised electricity for low-value or highly relocatable workloads while households and productive industry bear congestion costs. New capacity should therefore be evaluated by social value per constrained megawatt, reliability requirements, flexibility and recoverable heat or ancillary-grid services.
10.4.6 Lower-income countries and international development
The recommended international fund should finance more than imported accelerators. A viable package includes reliable electricity, cooling, connectivity, cybersecurity, local technical staff, foreign-exchange risk protection, educational access and long-term maintenance. Countries without sufficient demand to utilise a national cluster should receive federated access to regional systems, negotiated cloud capacity and latency-sensitive edge infrastructure. Grant finance is preferable for education, research and basic public services; revenue-generating commercial facilities can use concessional debt or guarantees.
The World Bank’s development framework treats compute, connectivity, data and skills as complementary foundations rather than interchangeable inputs. Digital Progress and Trends Report 2025: Strengthening AI Foundations – World Bank – 2025 — official publication. The ITU likewise documents persistent differences in connectivity and digital adoption. Facts and Figures 2025 – International Telecommunication Union – 2025 — official statistics portal. Funding accelerators without those complements can create expensive, underutilised installations and renewed dependence on foreign maintenance and software.
10.5 Policy portfolio by stakeholder
| Stakeholder | First-priority action | Second-priority action | Action to avoid |
|---|---|---|---|
| National governments | Shared public-interest compute plus grid planning | Targeted vouchers and repair/refurbishment standards | Untargeted hardware subsidies without access obligations |
| European institutions | Interoperable EuroHPC/AI Factory network | Cross-border capacity allocation and bottleneck investment | Fragmented national clouds unable to exchange workloads |
| Competition authorities | Specialist technical and economic investigation capacity | Complementary-asset merger review and switching-cost evidence | Treating concentration alone as proof of illegality |
| Universities | Federated procurement, workload scheduling and cost accounting | Minimum student/research access allocation | Numerous low-utilisation departmental clusters |
| Research organisations | Reproducible benchmarks and open research artefacts | Privacy-aware hybrid infrastructure | Selecting models solely by parameter count |
| Large companies | Workload-level ROI gates and multi-provider architecture | Portability tests and internal metering | Blanket AI deployment based on strategic fashion |
| SMEs | Cooperative purchasing and small-model/RAG specialisation | Contractual spending ceilings and exit plans | Frontier models for tasks that bounded systems complete adequately |
| Schools | Teacher-governed access entitlement | Privacy, accessibility and learning-outcome evaluation | Replacing instruction with uncontrolled chatbot use |
| Civil society | Independent access audits and price monitoring | Support for vulnerable groups and minority languages | Assuming free tiers provide capability equivalence |
| Lower-income countries | Regional infrastructure and negotiated capacity pools | Local skills, language resources and maintenance reserves | Debt-financed prestige clusters without utilisation demand |
| Development institutions | Grant-heavy foundational infrastructure | Open procurement and outcome-based disbursement | Vendor-tied finance that reproduces lock-in |
Table metadata: Unit and currency: not applicable. Reference year: 2026. Scope: institutional strategy. Source: synthesis of Chapters 1–9 and official sources cited in Chapter 10. Method: prioritisation by expected access benefit, implementation speed and distortion risk. Uncertainty: qualitative.
10.6 Quantitative decision rules
Policy should be triggered by observable conditions rather than general concern. Let:
AAIg,t = Annual cost of adequate AI access for group g at time t / Median disposable resources of group g at time t
CapabilityGapg,t = SuccessRatefrontier,t − SuccessRateaccessible,g,t
QueueStresst = Median waiting time for eligible public-interest workload / Maximum operationally acceptable waiting time
PortabilityFailuret = Failed or materially incomplete migrations / Attempted audited migrations
Intervention thresholds should include:
- student or low-income AAI above 5%: targeted voucher review;
- AAI above 10%: presumptive access-support eligibility;
- audited capability gap above 20 percentage points on essential educational or research tasks: capability-floor intervention;
- public research queue stress above 1.5 for two consecutive quarters: capacity procurement or allocation reform;
- portability failure above 10%: compliance investigation;
- provider switching cost above 20% of annual AI expenditure: contract and interoperability review;
- a verified effective-price increase exceeding consumer inflation by 15 percentage points without documented model or service improvement: enhanced disclosure and market inquiry;
- critical input lead time above 12 months combined with capacity utilisation above 90%: resilience investment assessment.
These thresholds are decision rules, not claims that intervention will automatically generate net benefit. Each programme still requires a counterfactual and cost-benefit test.
10.7 Explicit answers to the report’s central questions
1. Has AI-relevant hardware genuinely become less affordable?
For important user groups and configurations, yes; for hardware as a homogeneous category, the evidence does not support a universal answer. Nominal entry prices, high-memory requirements, complementary system costs and income differences can reduce affordability even when capability per unit of arithmetic performance improves. Frontier-equivalent local deployment is substantially less affordable than merely running a small quantised model.
2. Is the increase nominal, quality-adjusted or both?
The increase is clearly nominal in selected high-demand categories and periods. It is also quality-adjusted where memory, bandwidth, reliability or completed-workload costs deteriorate relative to the requirement. It is not generally quality-adjusted when the denominator is theoretical throughput alone. The conclusion changes with the quality measure, which is why price per successfully completed workload is preferable to sticker price or TFLOP alone.
3. Is used-market inflation systematic?
The evidence supports episodic, product-specific scarcity, especially for high-memory, power-efficient or unusually capable legacy devices. It does not establish that used AI hardware generally trades at two or three times original MSRP across countries and years. Such ratios should be reported as verified observations, not as the median market condition, until transaction-level data establish representativeness.
4. Which bottlenecks are most serious?
The highest systemic significance lies in leading-edge fabrication, advanced lithography, HBM, advanced packaging, substrates and interposers, high-speed interconnects, power equipment, grid connections and specialised software ecosystems. Their criticality reflects both concentration and limited short-run substitution.
5. Does concentration produce demonstrable pricing power?
It produces the capacity and incentive for pricing power, but concentration and profitability alone do not prove its exercise or illegality. Demonstration requires evidence of persistent margins above a defensible counterfactual, customer discrimination, capacity withholding, foreclosure, non-cost-based switching barriers or price responses inconsistent with innovation and risk.
6. Are current AI-service prices economically sustainable?
Basic and highly utilised inference services may be sustainable. Frontier subscriptions, long-context reasoning and aggressively priced promotional services remain uncertain because public filings generally do not isolate complete model-level economics. Provider-wide investment and depreciation imply that current posted token prices cannot automatically be treated as long-run equilibrium prices.
7. How much could enterprise AI expenditure rise by 2031?
Using Chapter 9’s indexed scenarios, aggregate enterprise AI expenditure rises from 100 in 2026 to approximately:
- 185 in the low case: +85%;
- 300 in the central case: +200%;
- 590 in the high case: +490%.
These are expenditure forecasts, not pure unit-price forecasts. They combine adoption, usage intensity, model mix, governance, infrastructure and labour costs.
8. Can local AI remain realistic?
Yes, for bounded, privacy-sensitive and repeatable workloads. It is less realistic as a universal substitute for continuously updated frontier services, especially where many concurrent users, very long contexts or near-frontier multimodality are required.
9. What constitutes genuinely usable local AI?
A system is usable only when it meets a workload-specific threshold for:
- task success and factual reliability;
- sufficient context and usable memory;
- acceptable tokens per second;
- bounded time to first token;
- required concurrency;
- operational uptime;
- privacy and auditability;
- manageable energy cost;
- maintainable software and model updates.
Loading weights into memory is necessary but not sufficient.
10. Are students and researchers without frontier access disadvantaged?
The evidence supports a credible and testable disadvantage, particularly in coding, multilingual work, document synthesis, advanced reasoning and rapid iteration. The magnitude is not yet universally established because longitudinal causal studies remain limited. The appropriate conclusion is measurable task-level disadvantage with incomplete evidence about lifetime educational effects.
11. Which countries and groups face the greatest exposure?
Those with low disposable income, weak currencies, import duties, unreliable electricity, limited broadband, long latency, no nearby cloud region, export restrictions and weak local-language support. Within countries, exposure is highest for low-income students, independent researchers, rural users, small organisations, minority-language communities and privacy-sensitive institutions without infrastructure budgets.
12. Could AI deepen global inequality?
Yes. AI complements skilled labour, data, capital and organisational capacity, allowing initially advantaged users to compound gains. Chapter 9’s central scenario increased the illustrative AI-access Gini from 0.193 in 2026 to 0.215 in 2031; the divergence case reached approximately 0.290. Conversely, low-cost models and shared access can reduce inequalities if capability floors improve faster than frontier costs rise.
13. Is the “AI priced like Bitcoin” analogy defensible?
Only narrowly. Compute can be scarce, metered, reserved and exposed to demand-sensitive pricing. Unlike Bitcoin, however, compute is heterogeneous, depreciating, location-dependent, perishable when idle, continuously produced and inseparable from memory, energy, networking and software. It is therefore better understood as a differentiated capacity service or electricity-like industrial input than as a standard store of value.
14. Which interventions expand access without suppressing innovation?
The strongest portfolio combines:
- shared public-interest compute;
- targeted student and research vouchers;
- cooperative purchasing;
- portability and open interfaces;
- standard price disclosure;
- open-weight and local-language research;
- repair and verified refurbishment;
- evidence-based competition enforcement;
- grid and bottleneck investment;
- competitively awarded, milestone-based industrial support.
These measures enlarge demand and contestability without imposing general price controls.
15. Under what conditions would the report’s warning be proven wrong?
The central warning would be materially falsified if, by 2031:
- quality-adjusted cost per completed representative workload fell by at least 70% from 2026;
- the Global AI Access Gini fell to 0.15 or below;
- adequate-access cost fell below 3% of median disposable resources in lower-income countries and student populations;
- the frontier-versus-basic capability gap fell below 15 percentage points on essential tasks;
- median lead time for critical accelerators, HBM and packaging fell below six months;
- audited provider migration succeeded in at least 95% of cases without material functional loss;
- no strategically essential layer maintained both HHI above 2,500 and persistent supranormal margins unsupported by innovation, risk or capital recovery;
- at least 90% of accredited students and researchers obtained adequate, reliable AI access.
10.8 Final risk classification framework
| Classification | AI-access Gini | Adequate-access affordability burden for exposed groups | Essential-task capability gap | Structural supply condition |
|---|---|---|---|---|
| Limited | Below 0.10 | Below 3% | Below 10 percentage points | Critical lead times below 6 months; broad substitution |
| Moderate | 0.10–0.15 | 3–5% | 10–20 points | Occasional bottlenecks; no persistent multi-layer scarcity |
| High | Above 0.15–0.22 | Above 5–10% | Above 20–35 points | Persistent concentration or 9–12 month lead times in critical layers |
| Severe | Above 0.22–0.30 | Above 10–20% | Above 35–50 points | CR4 above 80% or HHI above 2,500 across at least three complementary layers, with lead times above 12 months |
| Systemic | Above 0.30 | Above 20% | Above 50 points | Multi-region access failure, sustained rationing or exclusion of more than 25% of qualified users from essential workloads |
Table metadata: Unit: ratios, percentage of disposable resources, percentage-point benchmark gap and market-structure indicators. Currency: not applicable. Price basis: adequate annual access cost at 2026 real purchasing power. Reference year: baseline 2026, assessment horizon 2031. Scope: global and exposed-user populations. Source: thresholds defined by this report using results from Chapters 3, 7, 8 and 9. Method: multidimensional risk-classification rule; the highest persistent dimension determines the provisional classification, subject to corroboration by at least one additional dimension. Uncertainty: classification uncertainty of approximately one category because access and benchmark datasets remain incomplete.
Final classification: HIGH, with a material upward risk toward SEVERE
The current evidence exceeds the high-risk thresholds through the estimated access Gini, affordability burdens in exposed populations, substantial capability differences between unrestricted frontier access and constrained local or free-tier access, and concentration across several complementary supply-chain layers. It does not yet justify a global severe or systemic classification because falling quality-adjusted costs, expanding capacity, open-weight innovation and public infrastructure remain credible counterforces; universal exclusion has not been demonstrated; and several critical indicators rely on estimates rather than harmonised transaction data. Under the managed-oligopoly central scenario, the risk remains high through 2031. Under persistent scarcity or geopolitical fragmentation, the access Gini and capability gap can enter the severe range. The policy implication is preventive rather than punitive: intervention should create capacity, contestability and a minimum capability floor before deprivation becomes structurally self-reinforcing.
10.9 Required-table completion and audit register
| Required table | Primary chapter | Publication requirement |
|---|---|---|
| Global AI value-chain map | 1 and 3 | Consolidate firms, layers, dependencies and geography |
| Historical GPU price database | 2 | Retain product-level observations and exact dates |
| Historical memory-price database | 2 | Separate HBM, DRAM and NAND |
| Official, street and used prices | 2 | Preserve observation date and condition |
| Quality-adjusted hardware comparison | 2 | Publish every denominator and benchmark version |
| Supply-chain market shares | 3 | State market definition and share year |
| CR4 and HHI by segment | 3 | Preserve authority-specific interpretation |
| Five-year local-system TCO | 4 | Include discount rate and residual value |
| Model-size and memory matrix | 4 | Separate weights, KV cache, activations and overhead |
| Tokens per second and cost per million tokens | 4 and 5 | Report framework, quantisation and workload |
| Cloud-provider pricing | 5 | Timestamp every tariff |
| Provider cost stack | 5 | Separate observed values from inferred allocations |
| Enterprise cost by firm size | 6 | Publish archetype assumptions |
| Sectoral exposure matrix | 6 | Include regulation, privacy and failure cost |
| Student and researcher affordability | 7 | State income and resource denominator |
| Wage months for local hardware | 7 | Use net or disposable wage consistently |
| Country affordability table | 8 | Apply PPP and market exchange rates separately |
| Global AI Access Index | 8 | Publish raw indicators, weights and normalisation |
| Scenario assumptions and outputs | 9 | Separate scenario logic from forecasts |
| Annual forecasts, 2026–2031 | 9 | Include low, central, high and intervals |
| Sensitivity analysis | 8 and 9 | Report weighting and model alternatives |
| Early-warning dashboard | 9 and 10 | Assign trigger, frequency and responsible authority |
| Policy cost-benefit matrix | 10 | Included above; requires jurisdiction-specific refinement |
| Facts, estimates and projections | All chapters | Final classification table below |
| Unresolved data gaps | All chapters | Maintain as a living audit register |
Mandatory metadata for every final-publication table: unit; currency; nominal or real price basis; reference year; observation date; geographical scope; primary source; calculation method; sample size where relevant; uncertainty interval or confidence classification. Analytical tables appearing in individual chapters should be reissued in a consolidated statistical appendix so that metadata are not lost during publication formatting.
10.10 Evidence-status register
| Finding | Status | Confidence | Principal unresolved requirement |
|---|---|---|---|
| AI capacity requires complementary hardware, energy, software and skills | Confirmed structural fact | High | Improved cross-layer capacity data |
| Selected high-memory and frontier systems show large nominal costs | Confirmed for cited products | High | Harmonised worldwide street-price history |
| All AI hardware became quality-adjusted more expensive | Not confirmed | Low | Constant-workload hedonic panel |
| Used devices generally sell for two or three times MSRP | Not confirmed as representative | Low | Transaction-level marketplace data |
| Supply is highly concentrated in several critical layers | Confirmed | High | Consistent global market definitions |
| Concentration alone proves abusive pricing | Rejected inference | High | Conduct and counterfactual price evidence |
| Local AI is viable for bounded workloads | Confirmed conditionally | Medium–high | Standardised task and reliability tests |
| Local AI universally substitutes for frontier cloud services | Not confirmed | High confidence in rejection | Rapid change in model efficiency could alter result |
| Current frontier-service prices are fully cost-reflective | Unresolved | Low | Provider model-level cost and revenue accounts |
| Enterprise AI expenditure could triple by 2031 | Central projection | Medium–low | Usage, price and governance-cost observations |
| Differential access can compound educational advantage | Supported mechanism | Medium | Longitudinal causal studies |
| Global access inequality rises under the central scenario | Projection | Medium–low | Annual Global AI Access Index observations |
| Public/shared compute can improve access | Supported policy estimate | Medium | Programme-level counterfactual evaluation |
| Targeted vouchers outperform universal subsidies | Policy inference | Medium | Randomised or phased programme evidence |
| Compute will become Bitcoin-like money | Economically unsupported | High | No plausible fungibility or store-of-value mechanism identified |
Table metadata: Unit: evidence classification. Currency and price basis: not applicable. Reference date: August 2026. Geographic scope: global. Source: Chapters 1–10 and live-verified primary sources. Method: separation of directly observed facts, model estimates, forecasts and unresolved propositions. Uncertainty: categorical confidence stated in the table.
Final judgment
The emerging AI divide is not primarily a shortage of intelligence expressed through a single price. It is a layered inequality in the ability to finance, locate, power, operate and productively apply systems of different quality. Hardware learning can continue while access becomes more unequal; abundant aggregate compute can coexist with scarcity for specific users, regions, memory configurations and privacy requirements. On the evidence and thresholds established in this report, the present risk is HIGH. Without faster diffusion of shared capacity, effective portability, expanded energy and packaging supply, trustworthy secondary hardware markets and targeted educational access, the risk approaches SEVERE under persistent-scarcity or fragmented-world conditions. It becomes SYSTEMIC only if access inequality exceeds the defined thresholds and essential AI capability ceases to be realistically obtainable by a substantial share of qualified students, researchers, firms or countries.
Copyright of debugliesintel.com
Even partial reproduction of the contents is not permitted without prior authorization – Reproduction reserved
