India's 'Sovereign AI' Debacle: From Mistral to DeepSeek Shells
The article examines Sarvam AI’s lofty claim to build a sovereign Indian LLM, its $41 million funding, the use of Mistral and DeepSeek foundations, the government‑provided 4096 H100 GPUs, technical breakthroughs like a custom tokenizer, and the ensuing industry debate over copying versus genuine innovation.
1. The Grand Launch of "India Sovereign AI"
In December 2023 Sarvam AI, founded by IIT‑Bombay alumnus Pratyush Kumar (PhD, ETH Zurich) and IIT‑Delhi alumnus Vivek Raghavan (PhD, Carnegie Mellon), announced a $41 million seed‑plus‑Series‑A round backed by Lightspeed, Peak XV and Khosla Ventures. The founders, with deep experience in India’s Aadhaar identity system, pledged to create a large language model that understands all 22 official Indian languages.
2. First Flop: Mistral Shell
On 23 May 2025 Sarvam released Sarvam‑M, a 240‑billion‑parameter mixed model supporting ten Indian languages and fine‑tuned for math and code. The model’s base was the open‑source Mistral Small from the French company Mistral, on top of which Sarvam performed extensive post‑training with Indian language data. Despite marketing claims of “parity with leading global models,” download statistics showed only 23 Hugging Face downloads in two days, compared with ~200 000 downloads for a Korean open‑source model. Benchmark scores on IndicLLM were merely 0.02 points above the original Llama, indicating negligible improvement.
3. Government‑Backed Hardware
In April 2025 the Indian government selected Sarvam for the IndiaAI Mission, providing 4096 NVIDIA H100 GPUs for six months, hosted at the Yotta data centre. The total hardware bill was ₹2.47 billion, with a ₹986.8 million subsidy, leaving Sarvam to cover over ₹1.5 billion despite an annual revenue of only ₹291 million, creating severe financial pressure for the 114‑person team.
4. Second Flop: DeepSeek Shell
From mid‑February 2026 Sarvam launched a 14‑day product blitz modeled after OpenAI’s “12 Days of OpenAI.” On 18 Feb it announced Sarvam‑30B and Sarmam‑105B, both trained from scratch. The 30B model used ~16 trillion tokens and a Mixture‑of‑Experts architecture that activated ~1 billion parameters per inference. The 105B model featured a 128 000‑token context window, scoring 88.3 on the AIME 25 math benchmark (96.7 with tool use), 90.6 on MMLU and 98.6 on Math500. Critics quickly posted the model’s configuration file, labeling it a “shrunken DeepSeek copy.” Sarvam’s response emphasized that the architecture was inspired by DeepSeek’s Multi‑head Latent Attention and MoE design, but the underlying model was trained in‑house.
5. Industry Debate: Copying vs. Leapfrogging
Detractors argue that claiming “from‑scratch” training while reusing DeepSeek’s open‑source architecture is akin to assembling IKEA furniture and calling it handcrafted. Supporters point out that AI research inherently builds on prior work—Transformer originated at Google, MoE concepts date back decades, and DeepSeek itself aggregates earlier advances. On 6 Mar 2026 Sarvam open‑sourced the full weights of Sarvam‑30B and Sarvam‑105B under a permissive Apache 2.0 license, stating that all data pipelines, tokenizer, and reinforcement‑learning infrastructure were developed internally.
6. Reflection on Narrative vs. Reality
The core issue is India’s need for a sovereign AI capable of handling 22 official languages and thousands of dialects, where standard tokenizers are inefficient for scripts such as Sanskrit, Tamil and Bengali. Sarvam’s custom tokenizer improves token‑level efficiency for Indian scripts by three‑ to‑four‑fold, and the team built data pipelines handling ~16 trillion tokens for a 30B pre‑training run. However, the gap between the narrative of an entirely indigenous model and the reliance on external architectural blueprints fuels public backlash, echoing Aadhaar system chief Raghavan’s warning about becoming a “digital colony.”
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Black & White Path
We are the beacon of the cyber world, a stepping stone on the road to security.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
