AI infrastructure 3 min read

The Heretics Who Say the GPU Shortage Is a Mirage

There’s a phrase you can’t escape in tech right now: “We don’t have enough GPUs.” Nvidia’s stock has gone vertical, and Big Tech is busy pouring tens of billions into new data centers. But in the middle of all this noise, a few people are saying something quieter and more pointed: what if this is a bubble?

A False Note in the GPU-Shortage Chorus

For two years, the AI infrastructure story has run in exactly one direction. Bigger models, more compute, more GPUs. The formula became almost religious doctrine: more parameters means smarter models, so whoever hoards the most silicon wins.

But that logic smuggles in an assumption. It assumes every task needs a giant model. The small-model camp goes straight for that weak spot. The job you keep summoning a supercomputer for, they ask, couldn’t a laptop handle it just fine?

The poster child here is something like moondream, a tiny vision model. Not billions of parameters but hundreds of millions, a fraction of the size of the frontier models, and it still reads and describes images. It runs on a phone or an ordinary PC. As models like this multiply, the question sharpens: do we really need all those GPUs?

The Small-Model Argument, in Three Parts

The case is surprisingly simple, and surprisingly hard to wave away.

First, most real-world work doesn’t need a frontier model. Document classification, image tagging, quick summaries, the everyday stuff runs fine on small models. There’s no reason to fire up the most expensive chip on the planet for a routine task.

Second, small models are brutally efficient. They process more requests per watt. They respond faster. And they don’t need to ship your data off to the cloud, which means you get lower cost and better privacy in the same package.

Third, model efficiency keeps improving. What took a frontier model a year ago now runs on something far smaller. So the obvious question: will the GPU demand everyone is stacking up like mad today still hold a few years from now? That question is the whole bubble thesis in one sentence.

Bubble, or Genuine Shortage

The other side has its own strong rebuttals. Training still demands staggering amounts of compute. Even if inference gets dramatically more efficient, the demand to build and fine-tune new models isn’t going anywhere.

There’s also the counterintuitive twist: efficiency can detonate usage rather than shrink it. Economists call it the Jevons paradox. When something gets cheaper and more efficient, we don’t conserve it, we use far more of it. AI may follow the same path. The cheaper it gets, the deeper it embeds into everything, and total demand climbs higher than before.

So this isn’t really a fight about whether GPUs are scarce or abundant. It points to a more fundamental question: what size of model makes sense for what kind of task.

The Takeaway

To be honest, this isn’t a debate that’s been raging across the communities over the past month. The bubble thesis is still more of a fringe provocation than a mainstream view. But we’ve watched fringe provocations flip the board plenty of times before.

The point is this: when everyone sprints in the same direction, someone has to ask whether it’s actually the right one. Right now, the small-model camp is playing that role. So ask yourself: are you reaching for the biggest, priciest model every single time, or would something small and fast do the job? The real answer probably lives somewhere in between.

AI infrastructure GPU small models moondream tech trends

Comments

    Loading comments...