Morten Falch Sortland/Getty Images
Investor’s Guild
Investor’s Guild

AI re-routing

AI re-routing

Tuesday, July 28, 2026 by Stephanie Guild, CFA and Maddie MahoneySteph is Chief Investment Officer. Maddie is an investment strategist. Both are Wall Street alums.
Morten Falch Sortland/Getty Images
Morten Falch Sortland/Getty Images

As kids, many of us heard adults claim they “walked miles to school, in the rain and snow.” It was one way that generation tried to stress how good we had it. So I suppose I sound just as old when I talk about printing directions for a roadtrip. In those days, re-routing was not something we could just do. But today, it happens with a few taps on a phone, or for intelligence building, a few engineering tests. 

In the last month, there’s been no shortage of papers, articles, and research pieces published by various tech firms discussing the trials and tribulations of AI. Taken with the appropriate amount of skepticism, they’re chock full of pointers on the future of the AI buildout. These included Meta’s July 1 blog, Samsung’s blog from July 6, Micron’s joint blog with Meta from July 20, a ZDNet Korea article about memory in light of China’s CXMT’s niche strategy from July 24, and Nomura’s very long Advanced Testing piece, also from July 24.  

Like an old-school college kid, we went through all of it and even added in a Stanford course syllabus. Despite coming from different angles, they generally arrived at the same conclusion: the AI bottleneck keeps migrating around from compute, to memory, to storage, to the physical connections between all of it, along with funding. Here’s a breakdown:

Where everyone agrees:

  1. Memory, not just raw compute, is a durable chokepoint. GPUs have gotten fast enough that they now spend meaningful, expensive, time sitting idle, waiting on memory and storage to feed them data. So companies are willing to spend on faster storage and memory to keep chips fed, such as flash. 

  2. And because of the bottleneck, a diverse memory hierarchy has become the solution. While high-bandwidth memory feeds the GPU directly, a step below that sits a newer, more power-efficient memory standard, called Low-Power Double Data Rate (LPDDR), that's cheaper to deploy at scale. Below that, some cloud players are experimenting with pooling memory together so it can be shared more efficiently. And at the bottom of the stack, flash storage is now a place to park data the GPU isn't using yet. All of them are getting built out right now, simultaneously.

What reinforced our view (or skepticism?) of the memory shortage (and any shortage): human ingenuity. 

Some of the memory shortage fix is anti-shortage technology. letting existing memory get shared or substituted cleverly:

  • Compute Express Link (CXL) pooling: CXL is an industry-standard connection technology that lets a computer's processor talk to memory (and other devices) that aren't physically right next to it, at speeds fast enough that it can still be used like regular memory. Instead of every GPU server having its own dedicated pile of DRAM, multiple servers draw from one shared pool, which is more useful work out of it.

  • LPDDR offload: it's a category of DRAM (the memory chips that hold data your processor is actively working with) designed specifically to use less power than standard DDR memory. Instead of paying HBM prices for data that doesn't need HBM's speed, you place it onto cheaper, more power-efficient memory.

Both are engineered solutions and proof that part of the "AI needs more memory" story is being answered by needing less of the specific memory that's short. History says a shortage can usually be routed around it, because the profit motive to engineer around them drives it. 

For example, in their paper, Micron and Meta took memory technology (LPDDR5X) and ran it two different ways against real workloads, then compared each way to old DDR5 as the baseline. The chart below shows the trade-off: make LPDDR5X faster and you get more AI-pipeline throughput and lower tail latency; keep it slower but double the capacity, and “Spark”-style workloads more than double or triple their throughput.

This data shows memory is becoming a genuinely flexible tool that hyperscalers can tune to whichever problem they have.

Even so, total AI usage is growing, so total memory demand keeps rising (more queries, more models, more context length). But the amount of the specific, expensive memory type needed per unit of work is shrinking, because of the aforementioned engineering shifts.

That's the Jevons paradox shape: making something more efficient to use usually doesn't reduce total consumption, it expands the market. Cheaper memory-per-token means more AI gets built, which uses more memory in aggregate, even though each individual AI task may be lighter on memory than before.

What this means for our current preferred areas of investment or ones we have been considering:

  • A focus on all forms of memory may be a more diverse way to play demand. The newest, most exotic memory designs (the ones that stack logic and memory together), still look like custom, bespoke businesses, which opens a door for smaller or overseas competitors. 

  • Memory or chips still need the same specialized deposition, etching, and packaging equipment to get built, regardless of which specific technology wins.

  • As chips get more complex, testing eats up more of the total cost before they ship—though most of the specialized testing companies trade overseas. 

  • There's another gap: fast storage. If the GPU-idle-time math is right, storage is becoming part of the same bottleneck story as memory, just one tier further out. 

None of this breaks the broader thesis that AI infrastructure spending has a long runway. If anything, every source we reviewed reinforced it from a different angle. The winners will be whoever owns the narrowest pipe at any given moment or whoever sells the tools that build all the pipes. That's a more useful lens for the next leg of this trade than the one most of us started with.

More from Investor's Guild
The information provided here is for general informational purposes only and is not an individualized recommendation of any security, digital asset, or investment strategy. Expressions of opinion are as of this date and are subject to change without notice. There is no guarantee that these statements or opinions provided herein will prove to be correct. Past performance is no guarantee of future results. Investing involves risk including loss of principal. Diversification does not ensure a profit or guarantee against a loss. Information shown is as of a certain date and represents a point in time. Data will generally not be updated after publishing. Data is obtained from what are considered reliable sources. However, its accuracy, completeness, or reliability cannot be guaranteed. Supporting documentation for any claims or statistical information is available upon request. Keep in mind that individuals cannot invest directly in any index, and index performance does not include transaction costs or other fees, which will affect actual investment performance. 5793128