Chennai, Sep 11: Bodhan AI, a Centre of Excellence in AI for Education incubated at IIT Madras, today announced a suite of open, state-of-the-art foundational AI models for Indic-languages, developed in collaboration with NVIDIA.

Working with AI4Bharat, which brings expertise in Indic-language, multilingual modelling and datasets, Bodhan AI is releasing models spanning four core capabilities: speech recognition , text-to-speech, machine translation and optical character recognition.

Bodhan AI and AI4Bharat are fostering an open AI ecosystem in India by giving developers models they can deploy, adapt, and fine-tune. The initiative aims to accelerate innovation and adoption of multilingual and multimodal AI solutions in the broader Indian language AI ecosystem, focussed for education. The goal is to make open, state-of-the-art Indian language AI more widely accessible.

To develop and deploy these models, Bodhan AI and AI4Bharat use the NVIDIA NeMo framework to train models for automatic speech recognition, machine translation and optical character recognition. This includes post-training NVIDIA Nemotron 3.5 ASR to support Indian languages, including regional dialects and accents. The models are served using NVIDIA® TensorRT™ LLM and vLLM inference microservices.

Both NVIDIA and Bodhan AI are collaborating on datasets, training recipes and evaluations to support future foundational models for Indian languages. The models are expected to benefit a wide range of stakeholders including ed-tech companies, universities, researchers, startups, technology enterprises and government agencies, by addressing a persistent gap: millions of Indians interact with technology primarily through regional languages, yet developers often lack access to reliable speech, translation and document AI models.

Bodhan AI and AI4Bharat are bridging this gap by creating foundational models that can be freely adapted and deployed across education and public-interest applications.

Elaborating on the ecosystem that is being created, Prof. Mitesh Khapra, Principal Investigator at Bodhan AI and AI4Bharat, said,

“Bodhan AI and AI4Bharat aim to give India’s developers and model builders the best open models they can deploy, adapt and fine-tune, driving fast innovation and adoption of multilingual and multimodal AI across the broader Indian language AI ecosystem. Partnering with NVIDIA lets us bring open, state-of-the-art capability to that ecosystem faster than we could alone.”

Niket Agarwal, Senior Distinguished Engineer, NVIDIA, said

“AI is becoming foundational infrastructure for every industry, and open models are essential to ensuring developers can build for the languages and communities they serve. NVIDIA accelerated computing, NVIDIA Nemotron, and NVIDIA TensorRT provide the technology foundation to train, customise and deploy these models efficiently. By using these technologies, organisations can expand access to multilingual AI capabilities across India’s developer and research ecosystem.”

Bodhan AI and AI4BHarat are making these models available through both open-weight releases and hosted APIs.  Researchers, startups and technology companies can fine-tune the models for specialised applications, while organisations seeking rapid deployment can integrate the same capabilities through low-cost APIs. Bodhan AI’s educational applications will remain free for learners, educators and partnering State Governments.

For education, the models are designed to support learners and teachers across voice, text and scanned documents, making personalised learning and high-quality educational resources more accessible in regional languages.

Sovereign Digital Public infrastructure for AI in Education is a core aspect of Bodhan AI’s approach, which it sees as critical to running AI at population scale and to customising and operating models efficiently for India’s languages and contexts. Its API infrastructure is intended to be hosted within India’s sovereign digital ecosystem, reinforcing the vision of “AI for India, governed in India.

Leave a Reply

Your email address will not be published. Required fields are marked *