
Smallest.ai, a startup founded in late 2024, raised $13 million in Series A funding to develop voice AI agents that sound indistinguishable from humans by processing speech in real time—listening, thinking, and speaking at the same moment.
Unlike most AI agents that pause noticeably while generating responses, Smallest.ai's small, specialized model handles familiar customer support topics instantly and hands off complex questions to larger models only when needed, just as a human agent would.
The company already works with RingCentral and Truecaller and competes with established players like ElevenLabs.
What happened
Smallest.ai, founded in late 2024, raised $13 million in a Series A round led by Seligman Ventures, with participation from Sierra Ventures and 3one4 Capital, bringing total funding to over $21 million. The startup builds small, specialized voice models designed to process information the way humans do—listening, thinking, and speaking simultaneously—to eliminate the lag that makes AI agents sound unnatural.
Why it matters
Most AI customer support agents still sound like machines, with noticeable pauses that feel unnatural in conversation. Smallest.ai's approach uses a small, real-time voice model for immediate responses on familiar topics, handing off complex questions to a larger model only when needed—mimicking how a human support agent would briefly research an unfamiliar issue. The startup's existing customers include RingCentral and Truecaller, and the model may appeal to any customer support company that lacks in-house voice AI expertise.
What to watch
Smallest.ai competes with ElevenLabs, Cartesia, and regional players like Sarvam. The company focuses narrowly on real-time conversational voice agents for enterprise, handling diverse accents, dozens of languages, and noisy environments—not broader voice use cases like audio dubbing or podcasting.
Smallest.ai, founded in late 2024, announced a $13 million Series A round led by Seligman Ventures, with participation from Sierra Ventures and 3one4 Capital. The funding brings the startup's total capital raised to over $21 million. The company is building small, specialized voice models designed to make conversations with AI agents feel human-like by eliminating the unnatural pauses that plague most current voice AI systems.
The core innovation is a voice model that processes speech the way humans do: listening, thinking, and speaking simultaneously. As founder and CEO Sudarshan Kamath explained to TechCrunch, "While I'm speaking to you, you're already thinking, and you might interrupt me if I talk for too long. This is exactly how the startup's model is designed to work." Standard large language models operate differently—they receive an entire prompt, then begin processing—introducing noticeable latency. In text chat, such delays are tolerable; in voice conversation, they feel jarring and unnatural. Smallest.ai's model serves as a real-time intelligence layer that enables natural customer conversations on specific topics with virtually zero response lag. When the model encounters a subject outside its limited knowledge base, it hands off the query to a large foundational model and briefly places the customer on hold to "research" the issue, mimicking how a human support agent would handle an unfamiliar problem.
Kamath believes the future of AI agents relies on two models: a small voice model for real-time interaction and an "offline" large language model called upon as needed for complex problem-solving. Unlike large foundational models, Smallest.ai focuses strictly on voice-specific nuances—handling diverse accents, supporting dozens of languages, and operating reliably in noisy environments. The startup's existing customers include RingCentral and Truecaller. Kamath noted that any customer support company, including newer entrants like Sierra and Decagon, could benefit from Smallest.ai's technology. When asked why well-funded AI customer support companies wouldn't build their own voice models, Kamath responded that becoming "extremely good at doing voice is a distraction from their core business."
Smallest.ai competes with ElevenLabs, the leading voice AI company, as well as Cartesia and regional players like Sarvam that focus on local languages. However, Smallest.ai's positioning differs: while some competitors apply voice AI to dubbing and podcasting, Smallest.ai focuses exclusively on real-time conversational voice agents for enterprise customers. Kamath stated the company's central ambition: "We want our models to break the Turing test. You should speak to our model and not know it's AI or human. That's the sole focus of the company."
The startup addresses a persistent friction point in AI customer support: the uncanny, robotic pause that occurs when large language models process input and generate output. Standard LLM-based agents operate sequentially—receive input, then think, then speak—introducing measurable latency that disrupts the flow of natural conversation. Smallest.ai's founders argue that the next frontier is not faster inference on larger models, but purpose-built, smaller models trained specifically on the mechanics of human speech and conversation.
The company's two-tier architecture—small voice model for real-time exchanges, larger foundational model for complex queries—reflects an emerging industry belief that one-size-fits-all foundation models may be inefficient for specialized tasks. By narrowing focus to voice-specific challenges (diverse accents, multilingual support, noise robustness) and real-time customer support, Smallest.ai positions itself in a narrower but less crowded niche than competitors like ElevenLabs, which address broader voice AI use cases from dubbing to podcasting. The $13 million Series A, combined with customer traction (RingCentral, Truecaller) and a late-2024 founding date, suggests venture capital confidence in the thesis that specialized, conversational voice AI can capture enterprise value even as general-purpose models grow more powerful.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Friend, the AI pendant startup, relaunched its product with a built-in speaker that can talk to users, raising…

Fish Audio, an AI startup founded by former Nvidia researcher Shijia Liao, raised $52 million in seed funding…
Google released Lyria 3.5, a music generation model that produces more natural-sounding melodies, better lyric…

OpenAI released GPT Transcribe and GPT Live Transcribe, two speech recognition models via its API

OpenAI announced GPT Transcribe, a speech-to-text model that processes completed audio files, streamed file tr…

Palo Alto-based Fish Audio, which has over 8 million users and generates $21 million in annual recurring reven…

The AI news that matters, in one minute each morning.
Sign up free