Bodhan September 2026 Collection Collection of multilingual, multimodal models released as part of 5 Sep 2026 release • 5 items • Updated 10 days ago • 8
view article Article A Real-World Dataset for Noise-Robust Speech AI: 100+ Timestamped Noise Types Across 58 Indian Languages ARTPARK-IISc • 10 days ago • 3
SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages Paper • 2608.08235 • Published Aug 8 • 2
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition Paper • 2601.19451 • Published Jan 27 • 1
view article Article svara-TTS — Open Multilingual TTS for India’s Voices kenpath • Oct 27, 2025 • 26
NCERT_Dataset Collection The NCERT dataset is a collection of educational content derived from NCERT textbooks for students in standards 6 to 12. • 33 items • Updated Mar 2 • 9
Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention Paper • 2606.20945 • Published Jun 18 • 81
LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases Paper • 2602.13662 • Published Feb 14 • 1
RomanSetu Collection Romansetu is a collection of models address the challenge of extending Large Language Models (LLMs) to non-English languages using non-Latin scripts • 11 items • Updated Mar 7, 2025 • 5
IndicRagSuite Collection A comprehensive dataset collection for Indic language information retrieval. • 2 items • Updated Mar 2 • 3
IndicLLMSuite Collection Largest Collections of Pretraining and Instruction Finetuning datasets for 22 Indic languages. • 4 items • Updated Nov 5, 2024 • 20
IndicConformer Collection A collection of ASR models for 22 scheduled languages of India • 23 items • Updated Mar 2 • 42
Airavata Evaluation Suite Collection A collection of benchmarks used for evaluation of Airavata, an Hindi instruction-tuned model on top of Sarvam's OpenHathi base model. • 20 items • Updated Mar 2 • 10