So What About AI Agents
So What About AI Agents
Podcast Description
π What About AI Agents is your go-to podcast for exploring the rapidly evolving world of AI agents. From automating workflows to revolutionizing industries, we break down the latest advancements, real-world applications, and emerging trends in AI.
Join us weekly as we uncover how AI agents are shaping our future, featuring expert interviews, thought-provoking insights, and stories that bridge the gap between humans and intelligent systems. Whether you're an AI enthusiast, industry professional, or simply curious about the tech shaping tomorrow, What About AI Agents has something for you.
Podcast Insights
Content Themes
The podcast covers a range of topics related to AI agents, such as automation workflows, industry revolution, and emerging trends. For example, recent episodes discussed the impact of AI in customer experience, the role of AI in risk assessment with Alec Crawford from Artificial Intelligence Risk, and marketing strategies involving AI with Chris Hood from PolyAPI, showcasing how these developments affect real-world applications.

π What About AI Agents is your go-to podcast for exploring the rapidly evolving world of AI agents. From automating workflows to revolutionizing industries, we break down the latest advancements, real-world applications, and emerging trends in AI.
Join us weekly as we uncover how AI agents are shaping our future, featuring expert interviews, thought-provoking insights, and stories that bridge the gap between humans and intelligent systems. Whether you’re an AI enthusiast, industry professional, or simply curious about the tech shaping tomorrow, What About AI Agents has something for you.
AI spent the last few years competing on intelligence. Now the next major AI race is about speed.
As AI shifts from simple chatbots to agents that reason, write code, call tools, search the web, and coordinate other agents, latency compounds. A task that requires dozens or hundreds of sequential model calls can quickly turn seconds into minutes β or even hours. That makes inference speed a fundamentally different problem in the agentic AI era.
In this episode of So What About AI Agents?, Philippe Trounev sits down with Vasanth Mohan of SambaNova Systems to unpack what actually makes AI faster β from model architecture and memory bandwidth to batching, specialized accelerators, and next-generation inference hardware.
Vasanth explains the two major stages of inference β prefill and decode β and why they create very different hardware bottlenecks. They discuss operator fusion, parallelization across chips, memory bandwidth, and why reducing latency often comes with a significant cost tradeoff.
They also explore why coding agents are one of the clearest use cases for premium inference, how AI providers may eventually offer commodity, premium, and ultra-fast tiers of tokens, and why businesses willing to pay for speed may gain a meaningful productivity advantage.
Vasanth also shares early performance figures for SambaNova's upcoming SN50 architecture, including benchmark results around 800 tokens per second on a large model, compared with roughly 300β400 tokens per second on GPUs in the cited comparison.
We also get into a slightly crazier question: are AI agents already helping engineers design the next generation of AI hardware? The answer is increasingly yes β although humans are still very much in the loop.
In this episode:
- Why AI agents make inference speed dramatically more important
- Why sequential agent workflows create a latency bottleneck
- Prefill vs. decode explained
- GPUs vs. specialized AI accelerators
- The relationship between speed, batching, throughput, and cost
- Why memory bandwidth matters for large AI models
- What SambaNova's SN40 and SN50 architectures are designed to solve
- 800-token-per-second AI inference
- Why coding agents benefit so much from faster models
- Commodity vs. premium vs. ultra-fast AI inference
- Whether faster AI becomes a competitive advantage
- How AI agents are already being used in hardware engineering
- Why AI infrastructure may become just as important as the models themselves
Chapters
00:00 β The AI race is shifting from intelligence to speed
00:39 β Why AI agents suddenly need faster inference
02:38 β What actually makes an AI model run faster?
05:30 β When does ultra-fast inference matter?
08:08 β Does model architecture determine inference speed?
10:05 β How do you optimize AI from model to hardware?
12:44 β How fast can AI inference actually get?
15:28 β The sequential latency problem with AI agents
16:39 β Where faster AI creates the most value
18:30 β Will premium AI inference become a competitive advantage?
20:11 β What does AI inference actually cost?
23:10 β How many tokens can AI hardware generate?
24:46 β Building private AI infrastructure
28:02 β How much faster can AI eventually become?
30:14 β Are AI agents already designing AI hardware?
32:50 β Why today's βfastβ AI will eventually feel slow
33:34 β What happens next in AI infrastructure

Disclaimer
This podcastβs information is provided for general reference and was obtained from publicly accessible sources. The Podcast Collaborative neither produces nor verifies the content, accuracy, or suitability of this podcast. Views and opinions belong solely to the podcast creators and guests.
Β
For a complete disclaimer, please see our Full Disclaimer on the archive page. The Podcast Collaborative bears no responsibility for the podcastβs themes, language, or overall content. Listener discretion is advised. Read our Terms of Use and Privacy Policy for more details.