ElevenLabs’ plan to offer voice AI services in all 22 official Indian languages over the next year is more than a product expansion. It is a test of whether India’s next digital interface can move beyond screens, keyboards and English-dominant systems into the everyday languages through which citizens access services, seek information and communicate with institutions.
The London-based voice AI company currently provides speech-to-text and text-to-speech services in 14 Indian languages, according to Business Standard. At the company’s first India Summit in Bengaluru, chief executive officer and co-founder Mati Staniszewski said the company viewed India as its biggest market outside the United States and wanted to make interaction with artificial intelligence more natural, intuitive and trusted.
That ambition places voice technology inside an existing urban challenge: digital systems may be widely deployed, but access is still shaped by language, literacy, device type and the design of public-facing institutions. A service may technically be available online while remaining difficult to use for people who are more comfortable speaking than typing, or who use a language that is not adequately supported by an interface.
The proposed expansion therefore matters not only to the technology sector. It could affect how businesses handle customer interaction, how public agencies communicate with residents and how cities design grievance and information systems. But the available evidence also shows that the transition from a working voice model to a dependable civic interface will depend on much more than the ability to generate human-like speech.
India is being treated as an interface test
Staniszewski said AI adoption would depend on how well the technology interacted with people. He argued that sounding human was only an initial step, and that systems also needed to listen, understand context, respond naturally, follow guardrails and act on user intent.
That distinction is important for urban services. A voice system that can translate text into speech may improve access to information, but a citizen-service system must also identify what a person is asking, understand the relevant administrative context and route the request to the correct department or process. In grievance redressal, for example, the value of voice AI would not be measured only by whether a complaint can be spoken in Kannada, Hindi or Tamil. It would also depend on whether the complaint is recorded accurately, assigned to the right authority and followed through.
The company’s India expansion is taking place in a country with 1.4 billion people, hundreds of millions of feature-phone users and more than 100 spoken languages in active use, the Business Standard report said. The 22 official languages are therefore a formal coverage target, but they do not represent the full linguistic complexity of the country’s urban populations.
The source report also cited Nandan Nilekani, non-executive chairman of Infosys and architect of Aadhaar, describing voice AI as a potential final frontier and a practical interface for digital equality in India. The claim points to a central design question: whether digital public infrastructure can be made more usable by adapting systems to citizens, rather than expecting citizens to adapt to the systems.
From customer conversations to government services
ElevenLabs already works with more than 250 businesses and enterprises in India, including Razorpay, TVS Motor and Urban Company. Indian businesses recorded 100 million conversations on the company’s ElevenAgents platform over the past year, with more than 70 per cent of those conversations in Hindi, Kannada, Tamil and Telugu, according to the report.
Those figures indicate that voice AI is already being tested at a substantial operational scale in commercial settings. They also reveal where demand is concentrated. The four languages identified in the report account for most of the recorded conversations, suggesting that usage is not evenly distributed across India’s linguistic landscape even as the company prepares to add more languages.
The company has crossed a global revenue run rate of $600 million and is on track to have 100 employees in India, while its valuation has reached about $22 billion, Business Standard reported. The scale of the company’s commercial growth helps explain why India is being treated as a strategic market rather than a limited localisation exercise.
Yet public-sector deployment carries different responsibilities from commercial customer service. A business can decide how to handle a failed interaction or provide an alternative channel. A government service dealing with certificates, benefits, complaints or local infrastructure has to account for accessibility, records, accountability and the consequences of misunderstanding a request. The source material does not establish how these safeguards will be designed, which means the Karnataka partnership remains an exploration of potential adoption rather than evidence of an operational public system.
The partnership was formalised through a memorandum of understanding between ElevenLabs and the Karnataka Innovation and Technology Society, with Staniszewski and Karnataka Information Technology and Biotechnology Minister Priyank Kharge participating. The areas identified include skilling, citizen services, grievance redressal, responsible AI and voice restoration.
Each of these areas involves a different institutional problem. Skilling requires systems to explain information clearly and respond to varied user needs. Citizen services require reliable access to official information. Grievance redressal requires a traceable chain from complaint to resolution. Responsible AI requires rules for how systems handle sensitive information, uncertainty and harmful or incorrect outputs. Voice restoration, meanwhile, suggests applications beyond routine service delivery, though the source report does not specify the intended use cases or implementation model.
The governance question is larger than language coverage
The proposed move from 14 to 22 official languages creates a clear product milestone, but language availability alone cannot guarantee digital equality. A voice interface has to work across accents, dialects, speech patterns, background noise and differences in how people describe the same administrative problem. The input material establishes the company’s stated ambition, but it does not provide performance data for these conditions or explain how accuracy will be assessed across languages.
This is particularly relevant in cities, where public-facing interactions often occur in noisy and crowded environments. Residents may call from streets, markets, transport hubs, homes shared with other people or workplaces where background sound affects recognition. The source report does not provide technical benchmarks on accuracy, response times or error rates, so the practical reliability of the proposed systems remains unestablished.
There is also an institutional boundary that technology cannot remove. Voice AI may help a resident express a request, but it cannot by itself resolve unclear departmental responsibilities, missing records or delayed decisions. If a grievance system is fragmented, automating the first conversation may make the front end easier to use without fixing the administrative process behind it.
The Karnataka MoU is significant because it brings a private technology company into a state-government exploration of voice-enabled public services. It also makes responsible AI part of the stated partnership, rather than treating language support as a standalone commercial feature. The next question is how that principle will be translated into operating rules, oversight and measurable service outcomes.
The company’s own description of voice interaction—listening, understanding context, following guardrails and acting on intent—implicitly sets a higher bar than speech recognition or text-to-speech. For government use, those capabilities would need to be connected to clearly defined authority, escalation and audit mechanisms. The supplied report does not state whether such mechanisms have been finalised.
What the available numbers show
The data in the report describes a rapidly expanding sector. Tracxn counted 153 companies in the global voice AI sector, including 107 funded companies that had collectively raised $3.6 billion in venture capital and private equity by August of the year covered in the report. These figures place ElevenLabs’ India strategy within a broader investment push around conversational systems.
At the company level, the reported 100 million Indian conversations on ElevenAgents over the past year provide evidence of significant commercial usage. The fact that more than 70 per cent were in Hindi, Kannada, Tamil and Telugu offers a limited but useful picture of demand. It shows that Indian-language interaction is not a marginal feature for the company’s customers, while also indicating that usage is concentrated in a small group of languages.
The company’s target of covering all 22 official languages within one year is therefore both a technical and an organisational commitment. It will require expanding the language layer of the service while maintaining performance in languages already in use. The source material does not say how the company will measure success, whether all languages will receive the same capabilities or how dialectal variation will be handled.
The institutional test will be even more demanding. Commercial adoption can demonstrate that people are willing to use voice systems, but public-sector adoption must demonstrate that the interaction produces accurate, accountable and accessible outcomes. The Karnataka partnership may provide an early setting in which those questions are examined, although the report does not include a delivery timeline, budget, pilot locations or service-level targets.
For cities, the distinction matters because citizen-facing technology is not separate from governance. A voice interface becomes part of urban infrastructure when residents depend on it to understand a public service, report a problem or access a government process. Its success must then be judged not only by convenience, but also by whether it reduces barriers without creating new forms of exclusion.
India’s language diversity makes the country a demanding test bed for that proposition. The market size, the continued presence of feature phones and the volume of multilingual conversations provide strong reasons for companies to invest. They do not, by themselves, establish that voice AI will deliver equal access. That outcome will depend on implementation, institutional accountability and evidence from real users.
ElevenLabs’ announcement confirms that voice AI is moving from a consumer and enterprise tool towards a possible layer of public-service delivery. What remains open is how the company and government partners will convert language coverage into dependable civic access. The next developments to watch are the Karnataka partnership’s concrete use cases, the company’s expansion to the remaining eight official languages and the safeguards attached to any deployment in citizen services and grievance redressal.

