For the last three years, I’ve been planning, developing, and building an AI-powered engine to help me find the best, and potentially most explosive, stocks in the entire market.
An AI co-pilot, custom-built by me, that watches all 50,000 (or so) listed stocks across the US and UK. It whittles all of that down to just five that show every sign of being primed for a big leg higher.
I call it Hyperion .
You’ve probably heard me bang on about it for the past week. Good news: the demo is LIVE, and you can watch it right now, right here.
And for a limited time, we’re opening the doors to anyone who wants to use it and see just how powerful it is.
But while I was working on it the other day, I had 11 AI agents running at once. They were refining how it looked, checking that the backtesting was thorough and accurate, and making tweaks so the engine was bulletproof.
And while those 11 agents chipped away at their tasks, I just sat there waiting… and waiting… and waiting.
I had more questions lined up, more tasks to do. Things to add, to tweak, to make it all just better. But I couldn’t interrupt them, and I was pinned to the speed of the AI model doing the work.
I wished for more speed.
Twice the speed would have been a dream. Ten times, an epiphany. Twenty times would have been like I’d landed on an alien planet with a technologically advanced civilisation.
But 30 times… that would just be silly, right?
Right?
No.
Turns out that asking for 30x more speed isn’t silly. In fact, it’s available… as of this week.
Three chips the size of dinner plates
On Tuesday, Cerebras (Nasdaq: CBRS) launched the CS-4, calling it “the fastest AI accelerator in the industry, and a new foundation for frontier AI.”
The claim is up to 30 times faster AI inference than GPU systems (you can read that as Nvidia, but they don’t like calling it out directly).
Inference, as I explained in January when Nvidia bought inference chipmaker, Groq, is the phase where a trained model thinks and answers you.
Training teaches the model, and inference is the thought-out response it gives. It’s measured in tokens per second, and on that measure, the CS-4 is a beast.
More than 4,400 tokens per second per user on a 120-billion-parameter model. More than 1,000 tokens per second on models above 10 trillion parameters.
In short, that’s insanely fast.
It can do this because Cerebras has their wafer-scale engine (WSE). Most chipmakers carve a silicon wafer into hundreds of small chips, then wire them back together with the memory somewhere else entirely.
Cerebras keeps the whole wafer as one giant processor, with the memory sitting right on the silicon. The CS-4 loads what it calls “backpacks” into the rack. Each backpack has its own WSE chip. And each CS-4 can take three of these backpacks. They pass data between each other in as little as two microseconds.

A GPU spends most of its life waiting for data to arrive from memory. The CS-4 barely waits at all.
What’s one of the biggest things slowing an AI chip down?
|
Cerebras boss Andrew Feldman said at their Supernova event this week, “In AI, speed is productivity.”
None of this should be news to regular readers here at Investor’s Daily.
I first wrote to you about Cerebras back in February, when its wafer-scale systems were pumping out 2,000 tokens per second, compared with just 30 to 100 from GPU clouds. Back then, Cerebras wasn’t even a public company.
I personally had been watching the company for over a year, and liked it enough to buy in on private markets, before it listed.
On 14 January, this year, OpenAI signed a multi-year deal for 750 megawatts of Cerebras systems. That’s the largest high-speed AI inference deployment in the world, and it will roll out in stages starting this year
The thinking is that Cerebras chips are now a critical part of OpenAI’s rollout plans, and that it’s that speed advantage they believe will see the rise to the top of the frontier model food chain and stay there.
Forget AI, America’s No.1 forecaster says a bigger boom is coming:
“I’ve invested $1 million of my own money to prepare for this…”
He predicted the Financial Crash, both Trump victories and 2025’s record rare metals surge that saw stocks soar as much as 645%
Now discover the move he is making as America seeks to unlock a home grown fortune potentially worth trillions on Friday, May 15th
Find out what that move is right here >>
Capital at risk
Jane Street doesn’t buy toys
Cerebras isn’t alone in this race though. There is plenty of competition building in the inference chip world.
Also this week, an inference chip startup called Etched raised US$700 million at a US$21 billion valuation. That’s roughly double what its valuation was just a month earlier.
The round wasn’t led by a venture fund. It was led by Jane Street, arguably the best quant trading firm on the planet. It quite literally pays mega-bucks for microseconds of latency advantage, where speed is often the difference between profit and loss.
Jane Street took delivery of Etched’s first-ever production rack. It installed the rack in its own data centre and ran live workloads through it.
Then it led the new round.
So after testing the technology, it decided to pump US$700 million into Etched. That’s saying something!
And then there’s Groq…
Nvidia paid about US$20 billion on Christmas Eve for Groq’s inference technology and most of its team. It’s the biggest deal in Nvidia’s history. I wrote in January that it was proof AI is going exponential.
Eight months on, it’s clearly Nvidia’s play for more speed too.
Now, I don’t know which of these architectures wins.
WSE, transformer-specific silicon, Groq’s LPUs, or perhaps all of them carve out a big chunk of the market, such is the demand and opportunity in AI.
But whichever way you look at it, in AI, both now and in the future, speed wins.
A fast token is worth more than a slow one. It burns less power, and every agent, robot, and chatbot on Earth is about to demand billions of them every second. The potential growth in that demand is enormous.
The AI factory of the next five to 10 years will be built on inference chips. It will also get built on the optics and photonics that carry the tokens between machines because, again, when speed wins, what’s faster than light?
The companies that make tokens faster, and the ones that move those tokens around faster, are all part of a very exciting corner of the market.
But it’s not the only one.
Today, using Hyperion, I’ve pinpointed our first new stock for readers.
Watch the demo and join us at Hyperion to find out which one it is.
Until next time,

Sam Volkering
Investment Director, Southbank Investment Research