Sovereign & Shared: Frugally Scalable Multilingual-Multimodal AI for Bharat
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3764194 ↗
摘要
The movement for Sovereign AI is accelerating. Meeting its promise requires vertically integrated AI stacks -spanning data, models, and reasoning systems- that remain sovereign while adhering to shared scientific principles around which global research communities can coalesce. This talk presents BharatGen as a sovereign-yet-shared effort to make AI work for all: creation of datasets, benchmarks, and models that natively support Indian languages, dialects, and code mixing across text, speech, and vision; data pipelines grounded in local realities; and frugal methods that reduce cost and lower barriers. We outline our journey to date across language infrastructure, efficient training and distillation, and early sector pilots. The R&D deep dive will draw from some of our recent work on cross-lingual knowledge distillation for low-resource languages, tokenization/phonetic design for code-mix robustness, or trustworthy document AI with visual grounding focusing on robustness under dialect/code-mix shift, and latency/cost trade-offs. We hope to inspire other Sovereign-AI efforts, especially in the low-resource ecosystems of the Global South and close by inviting international collaborations toward principled research to build people-serving AI.