Beyond Subwords: Rethinking Bangla Tokenization for Vernacular AI
উপ-শব্দের বাইরে: বাংলা টোকেনাইজেশনের পুনর্বিচার
Why mainstream LLM tokenizers fragment Bangla script into disproportionate bytes and how morphologically-aware chunking levels the compute ground.
Most multilingual foundation models charge a heavy tax on non-Latin scripts. When a transformer evaluates a sentence in standard Bangla, modern Byte-Pair Encoding (BPE) tokenizers split standard conjuncts (যুক্তবর্ণ) and vowel signs (কার) into dozens of fragmented UTF-8 bytes.
The Asymmetry of the Token Tax
An equivalent prompt that takes 12 tokens in English frequently balloons to 45 tokens in Bangla. This inequality creates three structural bottlenecks:
- Severe context window compression: Context lengths are effectively quartered.
- Computational cost inflation: Serving costs 3x-4x more per semantic thought.
- Representational fragility: Semantic boundaries are fractured across raw byte boundaries.
Example: "সিরাজগঞ্জ জেলা" (Sirajganj District)
Standard Llama-3 Tokenizer: [234, 182, 185, 234, 182, 176, 234, ...] (11 tokens)
Delta-Morph BPE: [সিরাজগঞ্জ, _জেলা] (2 tokens)
Morphologically-Grounded Subword Segmentation
In our research at Logicdock Studio, we trained an open vocabulary tokenizer grounded in the morphological inflection patterns of Eastern Indic dialects. By respecting root words (ধাতু) and inflectional suffixes (বিভক্তি), we retain semantic coherence while compressing byte sequences by 58.4%.
“Language models should not force vernacular communities to purchase 4x the compute for the exact same inquiry.”
Key Findings
- Conjunct Stability: Preserving grapheme clusters prevents hallucinations during text generation.
- Dialectal Resilience: Suffix normalization allows seamless comprehension across Sylheti, Chatgaya, and Varendra regional colloquial forms.
- Inference Latency Reduction: Direct 2.4x throughput boost on quantized edge devices.
Our tokenizer weights and benchmark datasets are freely available through our open research archives.