AIToday
Audio & SpeechTechCrunch AIPublished: Aug 1, 2026, 01:00 JST5 min read

Smallest.ai raises $13M for voice AI that mimics human conversation in real time

Smallest.ai raises $13M for voice AI that mimics human conversation in real time

Key takeaway

  • Smallest.ai, a startup founded in late 2024, raised $13 million in Series A funding to develop voice AI agents that sound indistinguishable from humans by processing speech in real time—listening, thinking, and speaking at the same moment.

  • Unlike most AI agents that pause noticeably while generating responses, Smallest.ai's small, specialized model handles familiar customer support topics instantly and hands off complex questions to larger models only when needed, just as a human agent would.

  • The company already works with RingCentral and Truecaller and competes with established players like ElevenLabs.

3 Key Points

  1. What happened

    Smallest.ai, founded in late 2024, raised $13 million in a Series A round led by Seligman Ventures, with participation from Sierra Ventures and 3one4 Capital, bringing total funding to over $21 million. The startup builds small, specialized voice models designed to process information the way humans do—listening, thinking, and speaking simultaneously—to eliminate the lag that makes AI agents sound unnatural.

  2. Why it matters

    Most AI customer support agents still sound like machines, with noticeable pauses that feel unnatural in conversation. Smallest.ai's approach uses a small, real-time voice model for immediate responses on familiar topics, handing off complex questions to a larger model only when needed—mimicking how a human support agent would briefly research an unfamiliar issue. The startup's existing customers include RingCentral and Truecaller, and the model may appeal to any customer support company that lacks in-house voice AI expertise.

  3. What to watch

    Smallest.ai competes with ElevenLabs, Cartesia, and regional players like Sarvam. The company focuses narrowly on real-time conversational voice agents for enterprise, handling diverse accents, dozens of languages, and noisy environments—not broader voice use cases like audio dubbing or podcasting.

In Depth

Read the full story

Smallest.ai, founded in late 2024, announced a $13 million Series A round led by Seligman Ventures, with participation from Sierra Ventures and 3one4 Capital. The funding brings the startup's total capital raised to over $21 million. The company is building small, specialized voice models designed to make conversations with AI agents feel human-like by eliminating the unnatural pauses that plague most current voice AI systems.

The core innovation is a voice model that processes speech the way humans do: listening, thinking, and speaking simultaneously. As founder and CEO Sudarshan Kamath explained to TechCrunch, "While I'm speaking to you, you're already thinking, and you might interrupt me if I talk for too long. This is exactly how the startup's model is designed to work." Standard large language models operate differently—they receive an entire prompt, then begin processing—introducing noticeable latency. In text chat, such delays are tolerable; in voice conversation, they feel jarring and unnatural. Smallest.ai's model serves as a real-time intelligence layer that enables natural customer conversations on specific topics with virtually zero response lag. When the model encounters a subject outside its limited knowledge base, it hands off the query to a large foundational model and briefly places the customer on hold to "research" the issue, mimicking how a human support agent would handle an unfamiliar problem.

Kamath believes the future of AI agents relies on two models: a small voice model for real-time interaction and an "offline" large language model called upon as needed for complex problem-solving. Unlike large foundational models, Smallest.ai focuses strictly on voice-specific nuances—handling diverse accents, supporting dozens of languages, and operating reliably in noisy environments. The startup's existing customers include RingCentral and Truecaller. Kamath noted that any customer support company, including newer entrants like Sierra and Decagon, could benefit from Smallest.ai's technology. When asked why well-funded AI customer support companies wouldn't build their own voice models, Kamath responded that becoming "extremely good at doing voice is a distraction from their core business."

Smallest.ai competes with ElevenLabs, the leading voice AI company, as well as Cartesia and regional players like Sarvam that focus on local languages. However, Smallest.ai's positioning differs: while some competitors apply voice AI to dubbing and podcasting, Smallest.ai focuses exclusively on real-time conversational voice agents for enterprise customers. Kamath stated the company's central ambition: "We want our models to break the Turing test. You should speak to our model and not know it's AI or human. That's the sole focus of the company."

Context & Analysis

The startup addresses a persistent friction point in AI customer support: the uncanny, robotic pause that occurs when large language models process input and generate output. Standard LLM-based agents operate sequentially—receive input, then think, then speak—introducing measurable latency that disrupts the flow of natural conversation. Smallest.ai's founders argue that the next frontier is not faster inference on larger models, but purpose-built, smaller models trained specifically on the mechanics of human speech and conversation.

The company's two-tier architecture—small voice model for real-time exchanges, larger foundational model for complex queries—reflects an emerging industry belief that one-size-fits-all foundation models may be inefficient for specialized tasks. By narrowing focus to voice-specific challenges (diverse accents, multilingual support, noise robustness) and real-time customer support, Smallest.ai positions itself in a narrower but less crowded niche than competitors like ElevenLabs, which address broader voice AI use cases from dubbing to podcasting. The $13 million Series A, combined with customer traction (RingCentral, Truecaller) and a late-2024 founding date, suggests venture capital confidence in the thesis that specialized, conversational voice AI can capture enterprise value even as general-purpose models grow more powerful.

FAQ

How does Smallest.ai's voice model work differently from standard AI agents?
Smallest.ai's model listens, thinks, and speaks simultaneously, mimicking human conversation flow. When a user asks about a familiar topic, the model responds in real time with virtually zero lag. If the query falls outside its knowledge base, the model hands off to a larger language model, briefly placing the customer on hold—just as a human support agent would research an unfamiliar issue.
Who are Smallest.ai's customers and competitors?
Existing customers include RingCentral and Truecaller. Smallest.ai competes with voice AI leader ElevenLabs, as well as Cartesia and regional players like Sarvam. The startup targets any customer support company, including newer ones like Sierra and Decagon, that lacks in-house voice AI expertise.
Why does Smallest.ai focus on small models instead of faster large language models?
In voice conversations, even short pauses feel unnatural, unlike text chats where slight latency is acceptable. Large language models require processing an entire prompt before responding, introducing latency unsuitable for real-time speech. Smallest.ai's small model, specialized for voice-specific nuances like accents and noisy environments, enables immediate responses.

Get the latest Audio & Speech news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Related Articles

Next articleNebius Stock Jumps 10% on $1B AI Cloud Deal Through 2029

The AI news that matters, in one minute each morning.

Sign up free