Artificial IntelligenceInnovation

The Physics and Economics of Training a Frontier Model

How the world’s most powerful artificial intelligence quietly became a heavy industry — governed less by code than by electricity, memory and heat

Executive Summary

For most of its short history, artificial intelligence was imagined as software — weightless, infinitely copyable, bound only by the cleverness of its authors. This investigation argues that the frontier of AI has crossed a threshold and become something else entirely: a heavy industry, closer in character to steelmaking or aluminium smelting than to writing code. Training the most capable models now consumes gigawatt-scale electricity, mobilises tens of billions of dollars in fixed capital, and depends on a supply chain of chips, memory and lithography machines that a handful of firms and governments control.

We trace that transformation through its physical foundations — the thermodynamics of computation, the counter-intuitive dominance of memory over arithmetic, the tyranny of waste heat — and through its economics: the exponential rise in training costs, the scaling laws that govern them, and the quiet migration of expense from a one-off training run to the perpetual cost of serving a model to the world. Readers will finish with a materially different picture of modern AI: not a mind in the cloud, but a new kind of factory, and a contest over who can afford to build one.

Key Takeaways

  • Frontier AI is now capital-intensive heavy industry. In 2025 the combined data-centre capital spending of the five largest technology firms exceeded USD 400 billion — more than the world invested in oil and gas production.
  • The dominant physical cost is not calculation but data movement. On modern accelerators, an arithmetic operation costs a fraction of a picojoule, while fetching its operands from memory can cost a hundred times more.
  • Memory, not logic, is the binding constraint. High-bandwidth memory now accounts for more than half the bill of materials of a top-tier AI chip and is sold out years in advance.
  • Electricity has become the true bottleneck. Global data-centre demand is on course to roughly double to around 945 terawatt-hours by 2030, forcing operators toward gas turbines and a nuclear revival.
  • Training costs have grown roughly 2.4 times per year since 2016; the first billion-dollar training runs are expected before 2027, concentrating frontier capability in very few hands.
  • The economic centre of gravity is shifting from training to inference — the cost of running a model, which can represent 80 to 90 per cent of its lifetime compute bill.
  • Compute has become an instrument of statecraft, with export controls on chips, memory and lithography turning the supply chain into contested strategic terrain.

I.  The Illusion of Software

Before any explanation, stand for a moment inside the machine.

The first thing is the sound. Not the whine most people expect from computers, but a deep, oceanic roar — the massed breathing of coolant pumps and heat exchangers moving liquid through thousands of processors. The air in the cold aisle is engineered, dry and faintly cool; step through a door into the mechanical yard behind the hall and the temperature jumps, because everything the machines consume comes back out as heat. Overhead, busbars thick as a forearm carry current at scales that belong to substations, not offices. Along each aisle, black cabinets stand shoulder to shoulder, veined with transparent tubing through which water flows to within millimetres of the silicon. A single one of these cabinets can draw more electrical power than a small apartment block.

Outside, the illusion of software finally collapses. Beside some of the newest sites sit rows of gas turbines, installed because the electricity grid could not deliver power fast enough. Elsewhere, transmission towers march toward a dedicated substation, and paperwork is under way to restart a nuclear reactor whose entire output has been promised to a single technology company. This is not the popular image of artificial intelligence — a disembodied intelligence humming somewhere in “the cloud.” It is a construction site, a power station and a factory floor, fused into one.

For most of the field’s history, the metaphor of weightless software was accurate enough. A program could be written once and copied a billion times at almost no cost; the marginal price of another instance rounded to zero. But the models now defining the frontier of AI — the systems behind the most capable chatbots, reasoning engines and scientific tools — are no longer meaningfully described that way. Their creation has become a physical undertaking measured in megawatts, tonnes of copper and litres per second of coolant. Somewhere between the research paper and the industrial reality, artificial intelligence acquired mass.

DATA BOX The scale of the build-out In 2025, the capital expenditure of just five technology companies on data centres and computing infrastructure surpassed USD 400 billion, and is expected to rise by a further 75 per cent in 2026 — a figure the International Energy Agency notes is now larger than global investment in oil and gas production. The OpenAI-led “Stargate” programme, announced in January 2025, targets USD 500 billion and 10 gigawatts of computing capacity over four years. Its flagship site near Abilene, Texas, is designed to reach roughly 1.2 GW and to house up to 400,000 next-generation GPUs. Satellite tracking by the IEA finds that purpose-built “AI factories” more than tripled in capacity in the eighteen months to early 2026.

The people building these sites reach instinctively for the vocabulary of the industrial age. The chief executive of Crusoe, one of Stargate’s construction partners, described the effort on a site tour as “the largest capital investment in infrastructure in human history.” Commentators compared it to the Manhattan Project. Whatever one makes of the hyperbole, the underlying claim is sober enough: the constraint on advanced AI is migrating out of the realm of ideas and into the realm of concrete, copper and current. To understand why, we have to descend below the software, past the mathematics, to the physics of what a computation actually is.

WHY IT MATTERS If intelligence at the frontier is now bounded by physical inputs — power, chips, cooling, capital — then the questions that decide its future are no longer only about algorithms. They are questions of energy policy, industrial strategy and geopolitics. That reframing changes who gets to shape the technology, and how quickly it can advance.

II.  Measuring Intelligence in Joules

In 1961, a physicist at IBM named Rolf Landauer proved something startling: that logic itself has a thermodynamic price. Erasing a single bit of information — forcing a system from two possible states into one — must dissipate at least a minimum quantity of energy as heat, set by the temperature of the surroundings. At room temperature that floor, known ever since as the Landauer limit, is astonishingly small: about three zeptojoules, or three thousand-billionths of a billionth of a joule. It is the closest thing computation has to a speed of light — an absolute boundary imposed not by engineering but by the laws of physics.

The Landauer limit matters here for a reason that is easy to miss. Today’s best processors operate roughly a billion times above that floor. Every real operation in a frontier data centre spends vastly more energy than thermodynamics strictly demands, lost to the electrical resistance of wires, the switching of transistors and, above all, the shuttling of data. The gap is the point. It tells us that the energy hunger of artificial intelligence is not a fundamental law we have run up against; it is an engineering problem, and therefore an economic one. We are nowhere near the physical minimum. What we are near is the limit of what our present machines, grids and budgets can bear.

EXPERT INSIGHT Computation is physical The pioneers of computing understood their machines as physical systems. John von Neumann’s 1945 blueprint for the stored-program computer — the architecture nearly every processor still follows — separated memory from the unit that does the arithmetic. That separation was a brilliant simplification. Eighty years later it is also the source of AI’s deepest physical constraint, because it means that before any sum can be computed, its numbers must first be moved. A useful way to read the whole industry: intelligence, rendered in silicon, is a very large number of very small physical events — charges pushed across gates, bits flipped, values fetched — each with an energy cost that no cleverness in software can wish away.

How many such events does a frontier model require? The unit of account is the floating-point operation, or FLOP — a single multiplication or addition on the decimal-like numbers that neural networks manipulate. Training GPT-4 is estimated to have consumed around two hundred septillion FLOP; that is a 2 followed by twenty-five zeros. The models arriving now are trained on ten to fifty times more. These are quantities with no intuitive human analogue: if every person on Earth performed one calculation per second, without rest, they would need far longer than the age of the universe to match a single modern training run. That torrent of arithmetic is what the power stations, the racks and the coolant exist to sustain.

The energy of a single operation

Action inside the chipApproximate energyRelative cost
Landauer limit (erase one bit, room temperature)≈ 0.000003 pJ1 (physical floor)
32-bit integer addition≈ 0.1 pJ≈ 30,000×
32-bit floating-point multiply–accumulate≈ 1.5–4.6 pJ≈ 0.5–1.5 million×
Read 32 bits from on-chip cache≈ 5 pJ≈ 1.5 million×
Fetch 32 bits from off-chip HBM memory≈ 200 pJ≈ 65 million×
Fetch 64 bits from off-chip DRAM≈ 1,300–2,600 pJ≈ 0.4–0.9 billion×

Indicative figures after Horowitz (2014), normalised to the Landauer limit. The lesson is in the last two rows: moving data dwarfs the cost of the arithmetic it feeds.

That table contains the hidden thesis of this entire investigation, and it is worth pausing on. The arithmetic — the actual “thinking” — is cheap. What is expensive is remembering: reaching out to fetch the numbers the arithmetic needs. Hold that fact. It will reappear, in one guise after another, as the master key to why frontier AI costs what it costs.

III.  The Hidden Economics of Frontier Models

A decade ago, a state-of-the-art model could be trained for a sum a well-funded university laboratory might muster. That era is over. Analysts at Epoch AI, who maintain the most careful public accounting of the field, estimate that the amortised hardware-and-energy cost of the final training run for a frontier model has grown by a factor of about 2.4 every year since 2016. Exponential growth of that kind does not stay abstract for long. GPT-4, trained in 2022, is estimated to have cost on the order of USD 80 to 100 million for its final run. On the present trajectory, the largest training runs are expected to pass one billion dollars before 2027.

The head of one leading laboratory, Anthropic’s Dario Amodei, has said publicly that frontier developers are likely to spend close to a billion dollars on a single training run, and to contemplate ten-billion-dollar runs within a couple of years. Whether or not those precise figures hold, the direction is unmistakable, and it has a consequence that ought to trouble anyone who cares about who controls this technology: at these prices, the number of organisations able to train a frontier model from scratch shrinks toward a handful. Cost is quietly performing an act of political selection.

ECONOMIC PERSPECTIVE Where the money actually goes Contrary to intuition, energy is a minor line in the training budget. Epoch AI’s decomposition of frontier development cost finds that hardware accounts for 47–67 per cent, research and engineering salaries for 29–49 per cent, and energy for only 2–6 per cent of the total. The dominant cost, in other words, is the depreciation of extraordinarily expensive machines. A single top-tier server rack can cost USD 2–3 million; a data centre with one gigawatt of computing power represents on the order of USD 38 billion in up-front capital. This is the signature of heavy industry: enormous fixed costs, rapid obsolescence, and a relentless pressure to keep the machines fully utilised, because an idle GPU is a depreciating asset earning nothing.

The economics resemble those of an airline or a semiconductor fab more than those of a traditional software company. The asset is the hardware, and the hardware decays. A leading-edge AI accelerator may hold its commercial value for only a few years before newer silicon renders it uncompetitive for frontier work. That short useful life turns the whole enterprise into a race: capital must be deployed, run at full tilt, and earn its return before the technology beneath it moves on. It is why operators speak less about elegance of code than about “utilisation,” “depreciation schedules” and “time to first token” — the vocabulary of factory managers.

And yet the training run, for all its drama, is only the visible half of the ledger. The figure that makes headlines — the cost of teaching the model — is a capital expense, incurred once. The cost that will ultimately dominate is the expense of running the finished model for hundreds of millions of users, day after day, for years. We will return to that inversion, because it is where the economics and the physics finally converge. First, though, we must confront the single most surprising fact about these machines — the one already hinted at by the energy table above.

IV.  Why Memory Costs More Than Computation

Here is the counter-intuitive heart of the matter. In a modern AI machine, thinking is nearly free. It is remembering that is expensive.

The popular image of an AI “training” is of prodigious calculation — trillions upon trillions of sums, a furnace of arithmetic. That image is not wrong, but it is misleading about where the difficulty lies. As the earlier table showed, the arithmetic itself is astonishingly cheap in energy terms. A floating-point multiply-accumulate on a current accelerator costs perhaps one and a half picojoules. Fetching the two numbers it needs to multiply, if they sit in the chip’s external high-bandwidth memory, costs on the order of two hundred picojoules — well over a hundred times more. The processor could do the sum a hundred times over in the energy it takes merely to deliver the ingredients.

This is the “memory wall,” a phrase coined by the computer scientists William Wulf and Sally McKee in 1995, when they warned that processor speed was racing ahead of memory speed and that the gap would eventually strangle performance. Their prophecy has come true with a vengeance in the age of AI. A modern neural network is, at bottom, a colossal collection of numbers — its “weights” — that must be read from memory, multiplied against incoming data, and written back, over and over. The compute cores can perform those multiplications faster than memory can feed them. For long stretches, the most expensive silicon on Earth sits waiting for data to arrive.

ENGINEERING PERSPECTIVE The memory that feeds the maths The solution the industry has converged on is high-bandwidth memory (HBM): stacks of DRAM chips bonded vertically with thousands of microscopic vertical wires (through-silicon vias) and placed as close to the processor as physics allows. It is a manufacturing triumph and an economic chokepoint. HBM is difficult to make, and only three firms — SK Hynix, Samsung and Micron — can produce it at scale. It commands a premium of roughly five to eight times the price of ordinary memory per gigabyte, and by 2026 it had become, by analysts’ estimates, more than half the bill of materials of a top-tier AI chip. The physics of the memory wall has become the economics of the memory wall: the component that feeds the arithmetic is now worth more than the arithmetic engine itself.

Nothing captures the imbalance more sharply than the trajectory of a single chip family. Over roughly eleven years of successive generations, the raw arithmetic throughput of leading AI processors rose by a factor of around five hundred. The memory capacity attached to each processor, over the same span, rose by a factor of only about eighteen. Compute has sprinted; memory has walked. Each new generation is therefore more starved for data than the last, more likely to leave its arithmetic units idle, more tightly bound by the cost and scarcity of the memory that surrounds them. The frontier of AI is, in a precise engineering sense, bandwidth-bound.

Intelligence at industrial scale is limited not by how fast the machine can think, but by how fast it can remember.

This single fact reorganises everything. It explains why memory, not logic, is the sold-out component of the AI boom. It explains why a chip’s value is increasingly set by the memory strapped to it. It explains, as we will see, why waste heat is so ferocious (moving data dissipates energy), why inference is so costly (serving a model is one long act of memory-reading), and why the supply chain’s most contested chokepoint is not the logic fab but the memory stack. When people speak of the “intelligence” of these systems, they picture calculation. The machines themselves are built, priced and bottlenecked around remembering.

WHY IT MATTERS Because the constraint is data movement, the winners of the next phase will be those who solve memory and bandwidth, not merely those with the fastest arithmetic. That reframes chip design, data-centre architecture and even which nations hold strategic leverage — as the fight over high-bandwidth memory in the following chapters makes plain.

V.  Heat: The Invisible Enemy

Return to the physics for a moment, because it exacts a second tax that is even harder to escape than the first. Almost every joule of electricity a processor consumes is converted, in the end, into heat. This is not a flaw to be engineered away; it is thermodynamics. A data centre that draws a gigawatt of electrical power is, to a very good approximation, a gigawatt heater. The entire elaborate apparatus of cooling exists to move that heat somewhere else before it destroys the machines producing it.

The numbers have crossed a threshold that has quietly ended an era of data-centre design. For decades, a well-filled server rack drew something like five to ten kilowatts and could be cooled by blowing chilled air across it. The Uptime Institute pegs the industry-average rack at around 7.6 kilowatts. A single rack of the latest AI hardware — seventy-two processors wired together to behave as one giant chip — draws roughly 120 kilowatts, and often more under full load. That is a fifteen- to twenty-fold jump in a single generation. Air can no longer carry the heat away; at the surface of these chips, heat pours off at more than 500 watts per square centimetre, a flux comparable to the nozzle of a rocket engine.

DATA BOX When air cooling fails Air cooling tops out at roughly 8–25 kW per rack. The newest AI racks draw ≈ 120 kW — five to fifteen times beyond that ceiling — and weigh around 1.36 tonnes. The only viable answer is direct-to-chip liquid cooling: water (or a water–glycol mix) piped through cold plates pressed against the silicon, roughly two litres per second circulating through a single rack. A representative flagship rack integrates 72 GPUs and 36 CPUs, carries around 13.5 terabytes of high-bandwidth memory, and costs on the order of USD 2–3 million — sold not as a server but as a piece of factory infrastructure.

The consequences ripple outward through the building. Most of the world’s existing data centres — halls designed for air-cooled racks a fraction as dense — simply cannot host this hardware. Their floors are not rated for the weight; their electrical distribution cannot deliver the current; they have no plumbing for coolant. Retrofitting an air-cooled hall for liquid cooling is a six-to-twelve-month project, if it is possible at all. The rise of frontier AI is, among other things, a vast construction programme to build the only kind of building that can contain it. The efficiency of that building is measured by a ratio called Power Usage Effectiveness — total energy drawn divided by the energy that actually reaches the computers — and shaving it toward its ideal value of 1.0 has become a discipline unto itself, pursued through liquid cooling, waste-heat reuse and water-free designs.

HISTORICAL PERSPECTIVE From office appliance to industrial furnace For half a century the story of computing was miniaturisation: more capability in less space, at lower power. The personal computer, the laptop and the smartphone were all children of that trend, in which each generation of transistors ran cooler and cheaper. AI has reversed the arrow at the top end. To gather enough computation in one place, the industry has pushed power density to levels last seen in industrial process equipment. The machine that writes poetry now shares a cooling philosophy with a smelter. The return of heat as the central design problem is, in a sense, computing rejoining the physical economy it once seemed to transcend.

VI.  Scaling Laws and Diminishing Returns

Why build machines this monstrous at all? The answer lies in one of the most consequential empirical discoveries of the past decade: that the capability of these models improves in a smooth, predictable way as you pour in more of three ingredients — computation, data and parameters. These “scaling laws,” first mapped in detail by researchers at OpenAI in 2020, behave like power laws over many orders of magnitude. They turned model-building from an art into something closer to a forecastable engineering process: spend ten times the compute, and performance improves by a reliable, if modest, amount. That predictability is precisely what justified the willingness to spend billions.

In 2022, a team at Google DeepMind refined the recipe in a paper the field now treats as canonical, built around a model named Chinchilla. Training more than four hundred models, they found that most large systems had been built the wrong way: too big, and trained on too little data. The compute-optimal balance, they showed, is roughly twenty tokens of training text for every parameter in the model. Their 70-billion-parameter Chinchilla, trained on 1.4 trillion tokens, outperformed models several times its size that had been fed far less. The result rewrote industry practice overnight and put a number on the trade-off between a model’s size and the quantity of text it must read.

Two scaling recipes

 Kaplan school (2020)Chinchilla (2022)
Guiding ideaGrow the model faster than the dataBalance model size and data
Tokens per parameter≈ 1.7≈ 20
Illustrative modelGPT-3: 175B params, 300B tokensChinchilla: 70B params, 1.4T tokens
ConsequenceLarge models, undertrainedSmaller models, far more data
LegacyJustified the race to scaleRedefined “compute-optimal”

The Chinchilla correction showed that many earlier models were undertrained — and reframed the economics of every run that followed.

But the story has a twist that carries us straight back to the economics. Chinchilla optimises for the cost of training. Yet a model, once trained, is then run — served to users — billions of times. If a slightly smaller model can match a larger one’s quality by being trained on far more data than the compute-optimal recipe suggests, it is cheaper to run for its entire operational life. So developers now deliberately “overtrain” smaller models, pushing well past twenty tokens per parameter — often to a hundred or a thousand — accepting a higher training bill in exchange for a lower lifetime cost. The scaling laws, in other words, are no longer read as instructions for cheap training. They are read as instructions for cheap serving. The industry is already optimising for the inference era before it has fully arrived.

SCIENTIFIC PERSPECTIVE The limits of the curve Scaling laws describe smooth improvement, but they are curves of diminishing returns: each further gain demands exponentially more compute. Doubling capability may require ten or a hundred times the resources — which is exactly why costs and power draw climb so steeply. There is also a looming ceiling on the highest-quality human text available to train on — the so-called “data wall.” The field’s response has been to seek new axes of scaling: better data curation, synthetic data, and — most importantly — spending more computation at the moment a model answers, rather than only when it learns. That last idea reshapes the economics entirely, as the penultimate chapter shows.

VII.  Electricity Becomes the Bottleneck

Follow the constraint far enough and it stops being about chips and becomes about power. You can order the world’s finest processors and still be unable to use them, because the one input no cheque can conjure on demand is a gigawatt of electricity delivered to a specific patch of ground. Across the United States, the queues to connect large new loads to the grid are now measured in years — sometimes five to seven — as transmission systems strain to keep pace. Compute is abundant on paper and scarce in practice, and the scarcity is electrical.

The aggregate numbers are sobering. The International Energy Agency estimates that data centres consumed around 415 terawatt-hours of electricity in 2024 — roughly 1.5 per cent of the world’s total — and projects that this will roughly double to around 945 terawatt-hours by 2030, close to 3 per cent of global demand and comparable to the entire electricity consumption of Japan today. The AI-specific portion is the fastest-growing slice, projected to triple or quadruple over the same window. Data-centre electricity is rising some four times faster than demand from every other sector combined.

DATA BOX Powering the machines Global data-centre electricity: ≈ 415 TWh in 2024 (≈ 1.5% of world demand) rising to ≈ 945 TWh by 2030 (≈ 3%), and toward 1,300 TWh by 2035 in the IEA base case. The power needed for a single frontier training run has grown more than twice per year, with the largest now exceeding 100 megawatts — and is forecast to reach 4–16 gigawatts by 2030, enough to power millions of homes. Today, coal still supplies about 30 per cent of the electricity feeding data centres worldwide, and renewables about 27 per cent — a reminder that the cleanliness of AI depends entirely on the grid it plugs into.

When the grid cannot deliver fast enough, operators build their own supply. The most visible expedient has been the gas turbine, wheeled onto site to burn natural gas beside the servers. One of the most powerful clusters in the world, xAI’s Colossus in Memphis, went from empty warehouse to 100,000 processors in a reported 122 days — in part by installing dozens of gas turbines, a decision that drew objections from residents and regulators over permits and local air quality. The frontier’s appetite for power is running ahead of the institutions meant to govern it, and the friction is spilling into courtrooms and city halls.

The more strategic bet is nuclear. In September 2024, Microsoft signed a twenty-year agreement to buy the entire output of a restarted reactor at Three Mile Island — the site of America’s most notorious nuclear accident — rebranded the Crane Clean Energy Center, delivering 835 megawatts wholly dedicated to its data centres. Within weeks, Google contracted with a start-up, Kairos Power, for small modular reactors, and Amazon backed several such projects and a large nuclear plant in Pennsylvania. Meta issued a request for proposals for gigawatts of new nuclear generation. The pipeline of small modular reactors tied to data-centre operators grew from roughly 25 gigawatts at the end of 2024 to around 45 gigawatts a year later. An industry born of software has become one of the most consequential patrons of the atom.

GEOPOLITICAL PERSPECTIVE Power as sovereignty Electricity is not evenly distributed, and neither is the ability to build it quickly. Nations with abundant, dispatchable power — and the political capacity to permit new generation and transmission at speed — gain a structural advantage in the AI era. This is why energy policy has quietly become AI policy. The Gulf states, courting hyperscalers with cheap power and capital; the United States, fast-tracking approvals and reviving nuclear; China, marshalling coal and hydro at scale — each is competing on electrons as much as on algorithms. The map of future AI capability may be drawn, in part, by the map of the grid.

VIII.  The Geopolitics of Compute

No modern industry has a supply chain as narrow, as specialised, or as geographically concentrated as advanced computing. Nearly all of the world’s leading-edge AI chips are manufactured by a single company, Taiwan Semiconductor Manufacturing Company, on the island of Taiwan. The machines that etch the finest features onto those chips — extreme-ultraviolet lithography systems — are built by a single Dutch firm, ASML, using light-generating technology of almost science-fictional complexity. The high-bandwidth memory that feeds the chips comes from just three suppliers in South Korea and the United States. At every layer, the chain runs through a chokepoint that only a handful of actors can hold.

A supply chain this concentrated is, inevitably, an instrument of statecraft. Since October 2022, the United States has imposed escalating controls restricting the export of the most advanced AI chips and the equipment to make them, aimed principally at slowing China’s progress. The rules have tightened in stages: bans on top-tier processors in 2023, restrictions on high-bandwidth memory and an order halting certain chip shipments to China in late 2024, and a running cat-and-mouse over intermediate chips designed specifically to fall just under the legal thresholds. Extreme-ultraviolet lithography has been withheld from China since 2019, and proposals have since reached toward the less advanced tools as well.

DATA BOX The chokepoints Logic: the great majority of the world’s most advanced AI processors are fabricated by one firm, on one island. Lithography: a single company supplies the extreme-ultraviolet machines required to make them; these have been export-controlled to China since 2019. Memory: three firms control roughly 96 per cent of the DRAM market and effectively all high-bandwidth memory — capacity that is sold out years ahead. Concentration of use: by one estimate, the United States hosts around three-quarters of global GPU-cluster computing performance.

The policy has proved anything but stable. Controls on one intermediate chip, the H20, were tightened, loosened, and re-negotiated repeatedly across 2025 — at one point reportedly tied to a revenue-sharing arrangement with the US government — while Chinese authorities, citing their own security concerns, discouraged domestic firms from buying it. Late in 2025, a more capable chip was cleared for sale to China after all. Industry has learned to live in a state of permanent regulatory uncertainty, and each turn of the screw accelerates the very thing it seeks to prevent: China’s drive for self-reliance, embodied in domestic processors, a growing memory sector, and a national campaign to indigenise the whole stack.

EXPERT INSIGHT A contest, not a verdict Reasonable people disagree sharply about export controls. Proponents argue they meaningfully slow a strategic rival’s access to frontier compute and buy time. Critics counter that they fragment global markets, hand domestic Chinese champions a captive market, and spur an indigenisation that will erode Western leverage over the longer run. What is not in dispute is that compute has joined oil, sea lanes and semiconductors on the short list of things nations now treat as matters of national security. The frontier model, once a research artefact, has become a strategic asset — and its inputs, levers of power.

IX.  From Training to Inference

The headline number — the cost of training — is the one everyone quotes. It is also, increasingly, the wrong one.

Training a model is a capital expense, paid once. Running it — the industry calls this inference, the act of generating an answer to each query — is an operating expense, paid every time anyone uses it, for as long as the model lives. Multiply a modest per-query cost by hundreds of millions of users making billions of requests, sustained over years, and the arithmetic inverts. By most industry estimates, inference now represents 80 to 90 per cent of the total compute a model consumes over its lifetime; training is the smaller share. Google has reported that some 60 per cent of its machine-learning energy went to inference; Meta has cited a figure of around 65 per cent for one of its large models.

This inversion is being amplified by a new paradigm. The latest “reasoning” models — systems that visibly deliberate before answering, generating long internal chains of thought — spend far more computation at the moment of response than their predecessors. Researchers describe a rough equivalence: capability can be bought either by training a larger model or by letting a smaller one think longer at inference time, and for many hard problems the two are interchangeable. The result is that intelligence is migrating out of the one-off training run and into the perpetual, per-query act of thinking. And crucially, that per-query thinking is itself an exercise in reading memory — which returns us, once more, to the memory wall of Chapter IV.

ECONOMIC PERSPECTIVE Cheaper per token, dearer overall The price of a unit of inference has collapsed. Stanford’s AI Index documents the cost of running a model at a given capability falling from about USD 20 per million tokens in late 2022 to roughly USD 0.07 by late 2024 — a reduction of more than 280-fold in two years, driven by better hardware and software. Yet total inference bills are rising, because demand is growing even faster than price is falling — a modern instance of the Jevons paradox, where efficiency breeds consumption. The newest driver is “agentic” AI, where a single task fans out into many reasoning steps and tool calls, consuming by some estimates five to thirty times the tokens of a simple query.

The shift reshapes the physical footprint of AI as well as its accounting. Training can be concentrated in one enormous, remote campus and amortised over years. Inference must happen close to the user, quickly, every single time — which pushes computation outward, toward regional data centres and the network’s edge, and makes the geography of inference as strategic as its cost. The demonstration that a capable reasoning model could be built and served far more cheaply than assumed — as one much-discussed 2025 release showed — only sharpened the point: the frontier is no longer only about who can afford the largest training run, but about who can serve intelligence to billions at a sustainable price.

X.  The Future of Industrial Intelligence

Assemble the pieces and a single shape emerges. At the bottom sits physics: the thermodynamic cost of computation, and the far larger cost of moving data to and from memory. That physics expresses itself as engineering: bandwidth-bound chips, ferocious heat, mandatory liquid cooling. The engineering expresses itself as economics: exponential training costs, depreciating hardware, and the dominance of memory in the bill of materials. The economics expresses itself as energy: gigawatt appetites, grid queues, gas turbines and a nuclear revival. And the energy expresses itself, finally, as geopolitics: a contest over chips, memory, lithography and power that now runs through the heart of relations between the great powers. Each layer is a translation of the one beneath it. The whole stack is the anatomy of a new heavy industry.

What follows from this? First, that the pace of AI progress is now coupled to the pace at which the physical world can be built — substations, transmission lines, fabs, reactors, cooling plants. Software iterates in weeks; infrastructure iterates in years. That mismatch will increasingly set the tempo. Second, that capability will concentrate wherever capital, power and supply-chain access concentrate, raising uncomfortable questions about who is left out. Third, that efficiency — in memory, in energy, in the number of joules per unit of intelligence — becomes the master variable, because it determines how much intelligence a given quantity of the physical world can yield.

WHY IT MATTERS For a century, the frontier of technology seemed to be dematerialising — moving from steel and oil toward information and code. Frontier AI has reversed that arrow. It has taken the most abstract of human pursuits, reasoning, and rebuilt it as a physical process with a supply chain, a carbon footprint and a geopolitics. Understanding AI now requires understanding electricity, memory and heat as surely as it requires understanding mathematics.

None of this diminishes the achievement. That a machine can be taught to reason at all, by pushing charge through silicon a billion times a second, remains among the most remarkable feats of the age. But the romance of weightless intelligence should give way to a clearer, and in its own way grander, picture: of vast industrial cathedrals humming with liquid-cooled processors, drinking gigawatts, tended by engineers, straining against the limits of memory and heat. The mind in the cloud has a body after all — and the body is a factory.

Conclusion

We began with an illusion — that artificial intelligence is software, weightless and free to copy — and we end with its correction. The frontier of AI is a physical enterprise, and its constraints are physical: the thermodynamics of computation, the primacy of memory over arithmetic, the tyranny of waste heat, the ceiling of the electrical grid, the narrowness of the supply chain. The dollars follow the physics. The geopolitics follows the dollars. To ask what it takes to train the world’s most powerful models is, in the end, to ask what it takes to build a new kind of factory — and to reckon with who can afford to build one, who controls its inputs, and what it will cost the planet to run.

The most durable insight is also the most counter-intuitive. In these machines, thinking is nearly free; it is remembering, moving, powering and cooling that cost. The intelligence we are building is bounded less by the speed of thought than by the price of everything that surrounds it. That is not a temporary awkwardness to be engineered away next year. It is the defining condition of the industry — the reason frontier AI has become, quietly and irreversibly, one of the great physical undertakings of our time.

Final Reflection

For seventy years, we imagined that computation was escaping the physical world — growing smaller, cooler, freer, until it dissolved into pure information. The frontier model is the moment that story reversed. To make our machines think, we have had to give them a body of copper and coolant, feed them the output of power stations, and place their fate at the mercy of the grid, the fab and the atom.

We set out to build a mind and discovered we were building a furnace. Artificial intelligence is the point at which computation returned, at last, to the physical world — and brought the physical world’s oldest constraints back with it.

Timeline: The Road to Industrial Intelligence

YearMilestone
1945Von Neumann’s EDVAC report defines the stored-program architecture — and the separation of memory from processing that underlies the memory wall.
1961Rolf Landauer proves that erasing information has a minimum thermodynamic energy cost — computation’s absolute floor.
1995Wulf and McKee name the “memory wall”: processors are outrunning the memory that feeds them.
2012Deep learning’s breakthrough on image recognition ignites the modern era of GPU-driven AI.
2020Scaling laws are formalised, making model capability a predictable function of compute, data and size.
2022The Chinchilla study redefines “compute-optimal” training at ≈ 20 tokens per parameter.
2024Microsoft contracts to restart Three Mile Island; Google and Amazon back nuclear — AI becomes a patron of the atom.
2025The USD 500 billion, 10-gigawatt Stargate programme is announced; frontier training runs pass 100 MW of power.
2026Data-centre capital spending by five firms exceeds global oil-and-gas investment; memory shortages tighten across the industry.
2030 (proj.)Data centres approach ≈ 945 TWh of annual electricity (≈ 3% of global demand); inference dominates AI compute.

Glossary

  • FLOP (floating-point operation) — a single arithmetic operation on decimal-like numbers; the standard unit for measuring how much computation a model requires.
  • Parameter — one of the adjustable numbers (“weights”) a model learns during training; frontier models have hundreds of billions to trillions.
  • Token — a chunk of text (a word or word-piece) that a language model reads and generates; training data and running costs are counted in tokens.
  • HBM (high-bandwidth memory) — stacked DRAM placed beside the processor to feed it data quickly; the scarce, costly component at the heart of the memory wall.
  • Memory wall — the widening gap between how fast processors can compute and how fast memory can supply them with data.
  • Inference — running a trained model to produce an answer; an operating cost incurred on every query, as opposed to the one-off cost of training.
  • Scaling laws — empirical relationships showing how model performance improves predictably with more compute, data and parameters.
  • PUE (Power Usage Effectiveness) — total facility energy divided by the energy reaching the computers; a measure of data-centre efficiency, ideal value 1.0.
  • Test-time (inference-time) compute — computation spent while a model answers, letting it “think longer”; a way to trade training cost for inference cost.
  • Landauer limit — the minimum energy that must be dissipated to erase one bit of information; computation’s fundamental thermodynamic floor.

Ishraqa7 Editorial Team

The ISHRAQA7 Editorial Team produces premium documentary-style journalism covering history, science, geopolitics, exploration, engineering and innovation. Every article is carefully researched, fact-checked and written to provide readers with reliable, evidence-based analysis.
Back to top button