Invest Like the Best with Patrick O'Shaughnessy

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Invest Like the Best with Patrick O'Shaughnessy·August 26, 2026

OVERVIEW

The episode features Neil Movva, founder of Sale Research, discussing his company's mission to make AI inference, specifically for long-running AI agents, dramatically cheaper. He explains Sale's approach as a "token factory" focusing on low-cost intelligence by optimizing the entire compute stack from software to hardware and power, targeting a future where AI agents run autonomously in the background for extended periods.

KEY TOPICS

  • Sale Research's "token factory" model and its focus on abundant, low-cost intelligence.
  • The shift from low-latency AI inference (chatbots) to long-horizon, proactive AI agents.
  • The importance of open-source models and owning intelligence for businesses.
  • Technical trade-offs in GPU design, particularly between throughput and latency optimization.
  • Nvidia's historical innovation in GPUs and Tensor Cores, and its current market position.
  • Neil's contrarian view on Nvidia and the potential for alternative chip architectures.
  • The role of custom silicon (e.g., Cerebras, Groq) for very low-latency applications and their memory hierarchies (SRAM versus DRAM).
  • The "original sin" of Transformers in combining memory-bound and compute-bound operations.
  • The future of data: from human-generated internet data to model self-improvement through verifiable tasks in simulation environments.
  • Neil's "scavenger strategy" for acquiring compute hardware and power.
  • The evolving data center landscape: distributed 1-megawatt data centers versus monolithic gigawatt facilities.
  • The future of software development for AI, moving towards AI-generated and optimized kernels.
  • The importance of company culture, particularly valuing performance engineering and collaboration.
  • The balance between capital-intensive vertical integration and capital-light coordination in the AI compute market.

MAIN TAKEAWAYS

  • Sale Research aims to radically reduce the cost of AI intelligence, envisioning a future where intelligence is abundant and cheap enough to be deployed broadly for long-running, proactive AI agents, not just real-time chatbots.
  • The current AI compute market is heavily optimized for low-latency, interactive chatbot inference, which may not be the optimal path for future AI applications requiring extended, background operation.
  • Open-source models are critical because they allow users to own and control their intelligence, fostering a more robust market for customized and persistent AI.
  • The fundamental trade-off between throughput and latency in GPU design leads to different optimization strategies; Sale Research is betting on throughput for long-running agents.
  • The "speed of light" ethos from Nvidia (pushing hardware to its theoretical limits) is a powerful cultural force, but Neil believes the future lies in optimizing the entire stack to find efficiencies beyond what current leading vendors prioritize.
  • The next frontier for data is not more human-generated content but AI models self-improving within verifiable, simulated environments, allowing for continuous, measurable progress.
  • The compute market for AI is shifting towards a decentralized model, leveraging disaggregated, non-premium hardware and distributed, smaller data centers to achieve cost efficiency and exploit market inefficiencies.
  • Companies should focus on understanding and optimizing for bottlenecks across the entire stack (software, hardware, power, data centers) rather than solely competing on premium hardware.
  • The long-term vision includes AI agents proactively managing digital lives and solving scientific research problems with definable answers, enabled by ultra-cheap token costs.

NOTABLE QUOTES

"The theme of our company is abundance. We want to deliver this new commodity of intelligence to as many people as possible at a cost that is sustainable for almost every industry."
"The best latency is no latency at all. When you wake up in the morning, the work's already been done overnight."
"I believe that's the most profound change we're going to see in the next year. We're going to move away from chatbots to more proactive or background agents."
"My whole goal is to so dramatically expand the supply of power across the United States that I have a home for a lot of chips that otherwise would not have earned their place in a data center."
"The one thing I cannot teach is love for performance. Love for digging into every microsecond the machine is working and understanding what's happening on the machine at that time."

Summarized with DriftNote — AI-powered podcast summaries

Try it free