On Tuesday afternoon at Stanford, Norm Jouppi, an engineering fellow at Google, put two unnamed chips on a screen at the hottest “chip” event in the world.
Which one, he asked, is the inference chip?

The one on the left was smaller, with six stacks of memory. The one on the right was bigger, with eight.
Most of the room picked the big one as the training chip.
They were wrong.
The big one is the inference chip.
That little pop quiz tells you more about where AI hardware is heading than any spec sheet from three days of Hot Chips.
Hot Chips is a massive event where the world’s leading semiconductor makers explain the technical details of their pipeline of cutting-edge chips.
And this year, it was all about inference and speed.
Inference, the part where a trained model actually answers you, now needs more memory per unit of compute than training does.
For years, the story was that training was the hard bit and inference was the cheap afterthought. That has completely changed in the last few months.
Everyone, and I really mean everyone, is chasing speed.
Google (Nasdaq: GOOGL) has now split its eighth-generation TPU in two in the pursuit of speed – the TPU 8t trains while the TPU 8i thinks.
And Google wasn’t the only one at Hot Chips showing just how fast things are getting… …
Forget AI, America’s No.1 forecaster says a bigger boom is coming:
“I’ve invested $1 million of my own money to prepare for this…”
He predicted the Financial Crash, both Trump victories and 2025’s record rare metals surge that saw stocks soar as much as 645%
Now discover the move he is making as America seeks to unlock a home grown fortune potentially worth trillions on Friday, May 15th
Find out what that move is right here >>
Capital at risk
Everyone built for the answer, not the lesson
Nvidia (Nasdaq: NVDA) used Hot Chips to announce that its Groq 3 LPX rack is in full production.
I’ve covered Groq and the Nvidia deal a couple of times since January. But now we’re looking at real products shipping out, with Nebius as the first customer.
Nvidia’s numbers say the Rubin system with LPX tops out at 3,400 tokens per second on the Gemma 4 31B AI model, with a 100,000-token context.
The idea is simple. Rubin GPUs chew through the prompt, and the Groq chips spit out the answer. This is actually relatively similar to the process Amazon is taking with its Trainium chips alongside Cerebras’ inference chips.
Two chips for two halves of one job.
Nvidia calls it “extreme codesign.” I call it admitting that the GPU can’t win the inference race on its own.
Speaking of Cerebras (Nasdaq: CBRS), it showed a roadmap of where it’s heading only a week after debuting its CS-4 rack design, complete with Wafer Scale Engine (WSE) “backpacks.”
Cerebras said the next generation, CS-5, is coming in 2027, and it’s targeting up to 10,000 tokens per second per user on the smaller open models, and up to 5,000 on frontier ones. It’s also targeting a three million TPS per megawatt ratio.
That’s fast!
Then there’s the CS-6, which Cerbras says goes three-dimensional, stacking memory on top of the wafer. No specs on the TPS yet, but the expectation is that it’ll blow away everything again.
Cerebras also had a go at Nvidia’s rack design.
It was happy to point out that Rubin NVL72 needs about 5,000 cables. A CS-4 wafer has the fabric baked into the silicon and needs none.
They’re all chasing speed, and they’re all looking at side-by-side training and inference setups. And all the likely operators were in full force: Nvidia, Google, Cerebras…
Oh, and OpenAI.
The chip that built itself
Did you think OpenAI was just an AI frontier model company?
Think again.
It’s now in the chip business.
Jalapeño is OpenAI’s first custom chip, built with Broadcom (Nasdaq: AVGO). Tuesday was the first time anyone outside the company saw benchmarks from working silicon.
Impressive numbers, and also OpenAI’s numbers.
Semianalysis ran the comparison against GB200 and GB300, not Rubin. It also ran without the speculative decoding tricks Nvidia’s customers use in production.
Richard Ho, OpenAI’s hardware boss, said the picture could shift by the time Jalapeño ships at scale in 2027.
What makes Jalapeno so remarkable, though, is how it even came into existence,
AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip’s arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule.
Nine months, from a blank page to a finished design at the foundry.
Is that good?
Hell yeah, it’s good!
The process usually takes 18 to 24 months from concept to production, and that’s considered quick.
OpenAI is claiming it did the lot in less than half the usual time, on its first-ever chip.
And the AI efficiency gains don’t stop with the design process.
OpenAI also says AI-written code for parts of its GPT-OSS model ran 1.5 to 1.8 times faster on Jalapeño than the versions its own engineers wrote.
And they’re already near the end of the design and completion of generation two. Generation three won’t be far behind it.
Speed was on display everywhere. Fast chips, fast chipmaking, fast AI models, everything getting better, faster, and… accelerating.
Now, not everyone reads these incredible announcements at Hot Chips the way I do.
My colleague Jim Rickards has spent months arguing the other side of this.
The capital pouring into AI, the concentration of the market in a handful of names like Google and OpenAI, and the financial engineering underneath it all look to him like every great bubble he’s studied.
He’d point to the draft US Treasury report that compared AI to the dot-com bubble and warned that a downturn would hit chipmakers, cloud providers, and private credit at once.
Nvidia’s move on 10 August to help mobilise more than US$500 billion of outside capital for AI compute wouldn’t reassure him either.
I don’t agree with Jim’s conclusion, but you should hear his case from him rather than from me.
Jim’s full AI Meltdown briefing is here. If you invest, or plan to invest, in any of the names I’ve already mentioned today, Jim’s briefing is worth an hour or so of your time.
Why I say meltup
Jim’s argument is about money.
Mine is about time.
A bubble bursts when the promised productivity never shows up and the borrowed capital has to be repaid.
Fibre in the ground in 1999. Nobody to use it. Fifteen years to recover.
What Hot Chips showed is productivity smacking everyone square in the face.
The tools that many argue won’t pay back, are designing their own next generation, and each generation ships faster than the last.
Maybe this is the singularity, or at the very least a self-improvement loop that lifts productivity at a scale no one can quite comprehend.
If you want to compare this to 1999, ask yourself: did Cisco’s routers ever design Cisco’s next-generation routers?
While I’m asking questions, here’s another one…
Which past technology boom offers the most interesting parallel to AI designing its own next generation of hardware?
Find out what your friends and family think as well. Share this email with them so they can answer today’s poll and subscribe as well. |
Chip design is one of the hardest, most expensive, most physically constrained design problems humans undertake.
If AI can cut that process in half, think about what happens to car development and design, or a next-generation jet engine, or a battery cell, or as we’re already starting to see… drug development and discovery?
Everything is up for grabs.
Everything that takes time to do, design, or make should be using AI to improve productivity. In time, everything will.
Humans and AI converging can lift productivity to astronomical heights.
And the payback on that can come fast.
As this happens, the demand for AI-related components and hardware accelerates too.
Google’s inference chip needs eight memory stacks instead of six. Micron (Nasdaq: MU) says HBM eats about three times the wafer area of ordinary DDR5 for the same capacity.
SK Hynix, Samsung, and Micron have already sold their HBM capacity through 2027. If OpenAI scales Jalapeño across its 10-gigawatt Broadcom deal, it joins the queue for the same memory.
Faster design plus a genuine multi-year backlog in manufacturing equals a shortage that gets worse before it gets better. We’re moving ahead faster than current production capacity can keep up.
And when it does catch up, the baseline demand will be many times higher anyway.
So, waiting for some fallback on the demand curve…? Well, you might be waiting forever.
Every chipmaker at Hot Chips this week is spending record sums to make AI faster and cheaper, and using AI to do it.
Bubbles end when the spending stops producing things. There is no slowdown in production here, just a constant ramp-up. And from my view, any pullback or fear that runs through the market now is just a good excuse to buy quality, long term pillars of the AI revolution.
I’d rather own the shortage than bet against it.
Until next time,

Sam Volkering
Investment Director, Southbank Investment Research
PS Nvidia reports its second quarter after the US close today. It forward guided $91 billion of revenue in May and Wall Street wants about $92 billion.
I haven’t seen the numbers yet, even though they might have been released before you read this. Still, I’m going to say it now: it’s going to blow those away. I would think $95 billion-plus. Maybe even $100 billion could be on the cards.
In particular, pay attention to the data centre segment results and anything Jensen says about Rubin, because after Hot Chips the market will see anything Rubin related as inference not training, which will give you a good idea of the importance they’re putting now on inference development amid a lot of fast-approaching competition.