For fifty years, the semiconductor industry operated on a simple premise: pack more transistors onto a chip, and computers get faster. This principle—Moore’s Law—served as the prima facie blueprint for technological progress. But something fundamental has changed in the chips that power artificial intelligence. The new constraint is not how fast processors can calculate. It is how fast they can be fed information. Understanding this shift—and its implications for infrastructure investment, supply chain strategy, and competitive positioning—is now essential knowledge for any leader navigating the AI transformation.
The Essay: Why Moving Data Now Matters More Than Computing It
Imagine a highly advanced manufacturing facility. Inside are thousands of specialised machines capable of extraordinary precision and speed. These machines can complete millions of operations per second. On paper, it is the most productive factory ever built.
Now imagine that this factory sits at the end of a narrow country lane. Lorries carrying raw materials queue for hours to reach the loading bays. Finished goods pile up inside because there is no capacity to transport them out. The machines, however fast, spend most of their time waiting. The bottleneck is not the factory floor; it is the road.
This is the situation facing modern AI chips. The processors inside—NVIDIA calls them ‘tensor cores’—are astonishingly capable. They can perform trillions of mathematical operations per second. But those operations require data: the numerical representations of language, images, and knowledge that constitute an AI model. When a system like ChatGPT generates a response, it must read hundreds of billions of numbers from memory, perform calculations, and write results back. The speed of this data movement, not the speed of calculation, determines how quickly the system responds.
For executives, the strategic implication is significant: the chips powering AI are no longer defined primarily by their processing power. They are defined by their ability to move data. This reframing changes how we should evaluate infrastructure investments, understand supply chain constraints, and anticipate competitive dynamics.
Why AI Has a Unique Appetite for Data Movement
To understand why this matters specifically for AI, consider what happens when a large language model generates text. The model itself is essentially a vast collection of numbers—’weights’—that encode patterns learned from training data. A model like GPT-4 reportedly contains over a trillion such numbers. When the model produces each word of output, it must consult these weights, perform calculations, and generate the next word.
Here is the challenge: the model generates words one at a time, and each word requires reading a substantial portion of those weights from the chip’s memory. Unlike traditional computing tasks—where the same data might be reused many times before new data is needed—AI inference demands a constant stream of fresh information flowing from memory to processor.
The mathematics are instructive. A large AI model might occupy 350 gigabytes of memory (roughly the storage of 70 feature films). If the chip’s memory system can deliver 2 terabytes per second—an impressive figure by historical standards—reading the entire model takes approximately 175 milliseconds. That yields perhaps six words of output per second, regardless of how fast the processors themselves can calculate. Double the processing power and you achieve nothing; the memory system remains the sine qua non. This fundamental constraint has reshaped how AI chips are designed.
The Revolution in Memory Technology
The semiconductor industry’s response to this challenge has been ingenious: rather than moving data faster across long distances, move it shorter distances altogether.
Traditional computer memory sits on chips separate from the processor, connected by electrical traces on a circuit board. These traces act like motorways—adequate for conventional traffic but insufficient for AI’s demands. The solution, called High Bandwidth Memory (HBM), stacks multiple layers of memory chips vertically, like floors in a high-rise building, and places this memory tower directly adjacent to the processor on a shared foundation.
The analogy might be relocating a warehouse from across town to next door to the factory, and then adding multiple levels to that warehouse—each level connected by high-speed lifts rather than congested roads. The result is a dramatic increase in throughput. NVIDIA’s latest chips, built on the Blackwell architecture, incorporate memory systems capable of delivering 8 terabytes per second—roughly equivalent to streaming 1,600 feature films simultaneously. This bandwidth enables AI systems to generate responses far faster than previous generations.
The Hidden Manufacturing Challenge
Here is where the strategic picture becomes more complex. Building these advanced chips requires manufacturing techniques that are extraordinarily difficult and capacity-constrained.
The shared foundation that holds the processor and memory together—called an ‘interposer’ in industry parlance—must be manufactured with semiconductor-grade precision. The connections between components are microscopic, fabricated using the same advanced lithography techniques that produce the processors themselves. This assembly process, dominated by Taiwan Semiconductor Manufacturing Company (TSMC), represents a genuine production bottleneck.
Consider the implications. A single advanced AI chip like NVIDIA’s B200 requires not only the processor itself but also multiple memory stacks, an interposer to connect them, and specialised packaging to integrate everything. Each component has its own supply chain, its own capacity constraints, and its own yield challenges. The packaging process alone can cost thousands of dollars per chip and requires factory capacity that takes years to expand.
This reality inverts traditional semiconductor economics. For decades, advances in chip-making simultaneously improved performance and reduced costs—the famous virtuous cycle. Today, the most advanced AI chips are becoming more expensive and more supply-constrained even as demand accelerates. For enterprise leaders, this means that AI infrastructure planning increasingly resembles capacity planning for manufacturing: long lead times, constrained supply, and strategic supplier relationships matter as much as technical requirements.
The Art of Doing More with Less
If moving data is the constraint, engineers have also pursued a complementary strategy: reduce the amount of data that needs moving.
This is achieved through ‘quantisation’—representing numbers with fewer digits of precision. Consider the difference between measuring a room to the nearest centimetre versus the nearest metre. The latter requires less information to communicate, though it sacrifices precision. AI researchers have discovered that models can often tolerate substantial reductions in numerical precision without meaningful degradation in output quality.
The latest NVIDIA chips support formats that use as few as four bits per number—down from sixteen bits in previous generations. This four-fold reduction in data volume translates directly into faster inference, as the memory system can now deliver four times as many ‘useful’ numbers per second. Combined with the hardware improvements in memory bandwidth, the effective data delivery has improved by an order of magnitude in just a few years.
For business leaders, this illustrates a broader principle: performance in AI systems emerges from the interplay of hardware, software, and algorithmic innovation. Investments in any one dimension without attention to the others will yield suboptimal returns. The organisations achieving the best AI economics are those treating the entire stack as an integrated system.
Scaling Beyond a Single Chip
The largest AI models cannot fit on any single chip, however capable. Training frontier models like GPT-4 reportedly required tens of thousands of chips working in concert. Even running these models for inference often requires distributing the work across multiple processors.
This introduces another constraint: the speed of communication between chips. Just as memory bandwidth limits a single chip’s performance, the connections between chips limit a cluster’s collective capability. NVIDIA has invested heavily in proprietary interconnect technologies—branded NVLink and NVSwitch—that enable chips to communicate at speeds approaching their internal memory bandwidth.
The result is that a group of eight interconnected chips can, for many purposes, operate as though they were a single chip with eight times the memory. This architectural choice has profound implications for how AI workloads can be structured and, ultimately, for what becomes computationally tractable. It also means that AI infrastructure is best understood not as a collection of components but as an integrated system where the connections are as important as the components themselves.
The Strategic Picture for Enterprise Leaders
What does this mean for business leaders evaluating AI investments?
First, specifications can be misleading. Chip vendors advertise peak performance figures that assume ideal conditions rarely achieved in practice. The meaningful questions concern system-level performance: how quickly can the infrastructure serve your actual workloads? This requires understanding not just processing power but memory bandwidth, interconnect speed, and software optimisation.
Second, supply chain considerations are now strategic, not merely operational. Advanced packaging capacity is genuinely constrained, and the hyperscalers—Amazon, Google, Microsoft, Meta—have secured multi-year capacity agreements with NVIDIA and TSMC. Enterprises without similar foresight may find themselves waiting in queue, unable to execute AI initiatives at the pace their strategies require.
Third, the pace of architectural change is accelerating. HBM4, expected by 2026, will deliver further bandwidth improvements. Chiplet architectures—disaggregating processors into smaller, modular components—promise new flexibility in system design. The organisations that understand these trajectories can make infrastructure investments that remain relevant as the technology evolves.
Architecture as the New Competitive Frontier
Perhaps the most apt historical parallel is not Moore’s Law but the railway mania of the nineteenth century. Then, as now, new infrastructure technologies created possibilities previously inconceivable. Then, as now, the organisations that understood the systemic nature of the opportunity—not merely the locomotives but the tracks, the signalling, the stations, the supply chains—were those that captured enduring advantage.
We have entered an era where integrated design—across compute, memory, interconnect, and packaging—defines real-world capability. The question for today’s leaders is not whether to embrace AI infrastructure but whether they comprehend its architecture deeply enough to make wise decisions about their place within it.
The data must move. How we ensure it moves swiftly, efficiently, and at scale will shape the competitive landscape of the coming decade.
Follow me
That’s all for this month’s Essay. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.

