Bharat-Tiny-LLM v3 is a major leap forward for edge-first Hindi AI. Built on the Qwen3-1.7B base and powered by Eulogik's custom-engineered Brahmi Tokenizer, it represents our most capable Indic language model — running entirely offline on consumer hardware with zero cloud dependencies.
What's New in v3: The Brahmi Tokenizer Advantage
Version 3 integrates our proprietary Brahmi Tokenizer, which eliminates byte-fragmentation on Devanagari text. While standard tokenizers require 5 to 7 tokens for typical Hindi sentences, the Brahmi Tokenizer achieves 33% to 45% token compression, preserving semantic morphology and dramatically speeding up generation.
Additionally, v3 brings several breakthrough capabilities: GGUF quantisation for llama.cpp deployment, WebGPU inference via transformers.js for in-browser applications, tool-use capability for autonomous agentic workflows, and Hindi math reasoning trained on GSM8K-Hindi.
Architecture & Training
We fine-tuned Qwen3-1.7B-Base using a curated Hindi/Hinglish SFT dataset (eulogik/bharat-v3-chat-sft-v3) combined with AI4Bharat's IndicVault and the Hindi GSM8K dataset. The model retains Qwen3's strong English performance while achieving native-quality Hindi and Hinglish conversational ability.
Deployment Options & Data Sovereignty
Bharat-Tiny-LLM v3 supports multiple deployment paths: llama.cpp via GGUF for CPU/GPU inference, transformers.js for WebGPU in-browser inference, and standard HuggingFace Transformers for Python workflows. Crucially, it requires no internet connection — ensuring complete compliance with enterprise data sovereignty mandates and the Digital Personal Data Protection (DPDP) Act.