Sarvam Voice Agents are now open through a self-serve builder. In its August 13, 2026 announcement, Sarvam said the product is “available to everyone” after powering more than 350 million conversations across enterprise deployments. The company's live Voice Agents page now repeats the 350M+ figure and offers a direct Start building path.
That is a meaningful distribution change for an India-first voice platform. Until now, Sarvam's strongest public proof came from enterprise deployments, the Samvaad stack inside its broader model portfolio, and developer guides for assembling speech recognition, an LLM, and text-to-speech. The new surface packages those pieces into a builder for teams that want to describe an agent, test it, connect actions, obtain a phone number, and go live.
The headline still needs careful reading. “Available to everyone” means there is a public entry point, not that every feature is anonymous, free, instantly provisioned, or generally available in every region. And the 350M number is Sarvam's own cumulative scale claim, not an independent quality benchmark.
TL;DR: what people are asking
| Question | Direct answer |
|---|---|
| Is the builder public? | Yes. Sarvam's product page links to an authenticated Indus agents route |
| Is it usable without an account? | No. The builder redirects to login and acceptance of Sarvam's terms and privacy policy |
| What can it build? | Voice agents for sales, collections, support, and bookings, with memory, tools, testing, phone setup, and analytics |
| What does 350M+ mean? | A first-party Sarvam scale claim; timeframe, channel mix, and counting method are not disclosed |
| Is sub-500ms guaranteed? | No. It is a product-page claim that must be measured on your own end-to-end call path |
| Does it use Bulbul v4? | Not publicly confirmed. Current official IVR docs still name bulbul:v3 |
| Is pricing public on the launch page? | No pricing table appears on the Voice Agents page |
| Biggest launch risk? | Treating a fast demo as proof of reliable tools, consent, handoffs, and peak-load behavior |
What did Sarvam actually launch?
Sarvam describes the product as one platform to build, launch, and scale voice agents that carry customer context into later calls and work toward a business outcome. Its public workflow is concise:
- Describe the agent.
- Test the paths that matter.
- Get a phone number through KYC or bring an existing provider.
- Go live.
The builder is centered on Genie, a conversational interface that creates and edits the agent from a natural-language brief. The product page also lists AI-assisted test cases, custom context, cross-call memory, pause-aware behavior, voicemail detection, live API calls, transcripts, evaluations, call analytics, and production monitoring.
This is more than a newly documented TTS endpoint. It is an orchestration product: speech, dialogue, state, tools, telephony, evaluation, and operations are presented as one managed workflow. Readers comparing that approach with assembled pipelines can use explainx.ai's speech-to-speech voice-agent guide, while our types of AI agents guide explains why memory and actions change a scripted bot into an agent.
A starter brief to paste into Genie
Use a bounded workflow before asking the builder for an open-ended “customer service agent”:
Build an inbound appointment-booking voice agent for a dental clinic.
The agent may:
- answer clinic-hours and service questions from the approved knowledge base
- check available appointment slots through the scheduling API
- book, reschedule, or cancel only after reading back the date and time
- speak Hindi or English and handle code-switching
The agent must:
- identify itself as an AI voice agent at the start of the call
- ask before storing information for future calls
- never give medical advice or infer a diagnosis
- transfer to a human when confidence is low, the caller asks, or an API fails
- mask phone numbers in analytics exports
The value of this brief is not its wording. It gives tests a finite permission boundary. Every line can become a pass/fail case before the agent touches a real caller.
Is “available to everyone” the same as generally available?
No. When explainx.ai followed Start building, the official CTA resolved to indus.sarvam.ai/login?return_to=/agents. That proves a public signup and login route exists for the builder. It does not prove that an unauthenticated visitor can create an agent, that phone-number provisioning is instant in every jurisdiction, or that production traffic has no commercial approval step.
The public surface does establish four useful facts:
| What the surface shows | What it does not establish |
|---|---|
| A public product page and authenticated agents route | A free plan or included call volume |
| A conversational build-and-test workflow | Successful activation for every account or geography |
| KYC phone-number rental or bring-your-provider language | Carrier list, country coverage, and number-provisioning time |
| Cloud, private deployment, and security positioning | The exact commercial terms for each deployment mode |
This distinction matters because social replies reported that some people could not yet see Bulbul v4 in their dashboards. Those replies may accurately describe individual accounts, but they are not product documentation. They also answer a different question: a TTS model's dashboard availability is not the same as access to the end-to-end Voice Agents builder.
Does this launch include Bulbul v4?
Sarvam has not said so publicly. The Bulbul v4 reveal at Epoch emphasized richer emotion, natural expression, and vocal range, but did not publish an API model ID or migration guide. Two weeks later, Sarvam's current IVR and contact-center guide still recommends bulbul:v3 for streaming TTS.
The responsible reading is:
- Confirmed: the self-serve Voice Agents product is public.
- Confirmed: Sarvam's public API architecture uses Saaras speech recognition, a Sarvam chat model, and streaming Bulbul v3 in its documented example.
- Not confirmed: the managed product uses Bulbul v4 internally.
- Not confirmed: Bulbul v4 is enabled for every dashboard account.
- Not proven by replies: that the Voice Agents launch is blocked for everyone until v4 appears.
Sarvam may update the product independently of the API reference. Until it names the deployed voice model, builders should evaluate the output they receive rather than attaching an assumed version number to it.
What does the 350M+ conversations claim prove?
It is a strong signal that Sarvam has operated voice systems at enterprise scale. It is not proof that a new agent will achieve a particular resolution rate, latency, conversion lift, or customer preference score.
Sarvam now displays 350M+ conversations on its product page, alongside sub-500ms real-time latency and under five minutes to go live. These are first-party product claims. The page does not define:
- the date range behind the total;
- whether a conversation means an answered call, completed workflow, or any session;
- how voice, WhatsApp, and web interactions are separated;
- how many were production calls rather than tests;
- what percentile the latency number represents;
- whether go-live includes KYC, CRM integration, compliance review, and carrier setup.
That does not make the figures meaningless. It makes them vendor evidence to investigate, not a substitute for an acceptance test. The same rule applies to other managed platforms covered in explainx.ai's xAI voice-agent builder analysis and OpenAI realtime voice-model guide: the useful benchmark is the complete call path under your traffic and failure modes.
How should a team evaluate Sarvam Voice Agents?
1. Start with synthetic callers and fake records
Create a sandbox CRM and scheduling endpoint. Do not upload production customer history just to discover how memory or analytics works. Seed cases with ordinary, ambiguous, adversarial, and incomplete inputs.
2. Test language switching, not only language selection
Use names, account numbers, addresses, dates, abbreviations, and English product terms inside Hindi and regional-language sentences. Sarvam's India-first advantage should show up on code-switching and local entities, not only clean monolingual scripts.
3. Measure the entire latency budget
Capture p50 and p95 time from the caller finishing an utterance to the first audible response. Break out telephony ingress, endpoint detection, speech-to-text, tool calls, LLM generation, and first audio. A product-page number cannot account for your slow CRM.
4. Attempt unsafe actions deliberately
Ask the agent to change a phone number without verification, reveal another customer's balance, ignore a payment limit, or repeat sensitive fields. Tool permissions should enforce the rule even if the conversation model follows the malicious instruction.
5. Make human escalation a release gate
Sarvam's official IVR guide says to detect direct transfer requests, repeated failed intents, and DTMF 0, then move the caller to a human queue. Test transfers during business hours, after hours, when the tool times out, and when no human is available.
6. Audit memory as a separate data product
Cross-call memory can improve continuity, but it creates a durable profile. Verify what is stored, who can read it, whether callers can decline it, how individual memories are corrected or deleted, and whether one tenant can ever retrieve another tenant's context.
7. Score business outcomes and caller harm together
Track completion, transfer, repeat-call, complaint, opt-out, and wrong-action rates. A collections agent that improves throughput while increasing disputed actions is not a success.
What privacy and governance work remains yours?
Sarvam's Trust Center says the platform uses encryption at rest and in transit, tenant isolation, configurable retention, India-only residency for Indian deployments, PII masking, and ISO 27001:2022 and SOC 2 Type II controls. It also says its DPDP work is in progress, lists a Data Processing Addendum as available on request, and states that customer data is not used to train models for other customers.
Those controls support a deployment; they do not supply your lawful purpose or caller consent. Sarvam's own IVR documentation says call recording should happen with consent and tells teams to capture that consent in the call flow. The company's privacy policy also explicitly says enterprise content processed on behalf of business customers is governed by the customer agreement rather than the public privacy policy.
Before production, put these decisions in writing:
| Decision | Minimum question to answer |
|---|---|
| Identity | When and how does the agent disclose that it is automated? |
| Outreach | What authorizes this inbound or outbound call? |
| Recording | Is consent required, captured, and preserved for this use case? |
| Memory | Which facts persist across calls, for how long, and can the caller opt out? |
| Tools | Which actions require read-back, OTP, dual control, or human approval? |
| Retention | When are audio, transcripts, evaluations, and derived insights deleted? |
| Access | Which roles can replay calls or export transcripts? |
| Escalation | What happens when confidence is low or a caller asks for a person? |
| Incident response | What notification and deletion obligations are in the contract? |
For banking, healthcare, collections, and government use cases, review the sector-specific rules with counsel. A platform certification is not permission to call, record, profile, or take action on someone's behalf.
The explainx.ai read
The important launch is not a new voice demo. It is the shift from an enterprise-led conversational platform to a builder that anyone can enter through a public signup flow. Sarvam is turning the integrations and operating lessons behind its deployment history into a product surface.
That makes the next evaluation more demanding, not less. A voice agent earns trust through correct actions, predictable escalation, auditable memory, consent, and performance on real phone audio. Sarvam's 350M+ figure suggests the company has seen those problems at scale. The public launch gives smaller teams a chance to test whether those lessons are actually encoded in the builder.
Update — August 19, 2026: Sarvam's fourth "Build with Sarvam" showcase highlights three community builds on this stack — a burglar-deterrent device, a bargaining voice agent, and a CCTV investigation copilot.
Related on explainx.ai
- Build with Sarvam #4: 3 Voice Agents Built on India-First AI — real builder projects on this exact Voice Agents launch.
- Sarvam AI models, APIs, speech, and deployment guide
- Bulbul v4: confirmed launch details and missing API facts
- Sarvam Epoch 2026: every confirmed launch
- Build a speech-to-speech voice agent
- xAI Grok Voice Agent Builder analysis
- OpenAI GPT-Realtime 2 voice-model guide
- Types of AI agents and where voice agents fit
Primary sources
- Sarvam Voice Agents product page
- Sarvam IVR and contact-center voice-agent guide
- Sarvam Trust Center
- Sarvam Privacy Policy
- Sarvam's official X account and August 13 announcement
Product access, builder routing, public claims, model references, and security documentation were checked on August 13, 2026. Availability, pricing, phone-number coverage, model versions, and contractual controls can change; verify the live product and your agreement before deploying.
