Sarvam AI launches Vision 2.1 with better document reading and Indic handwriting recognition
Sarvam AI has launched Vision 2.1, an upgraded model for reading documents, forms and tables. The update also aims to turn unstructured scans into production-ready data with better Indic OCR accuracy.
by Kazi Nasir · India TodayIn Short
- Sarvam AI launches Vision 2.1 for document intelligence
- New model can handle complex tables, forms and handwritten text
- It supports OCR across 22 official Indian languages
Homegrown AI startup Sarvam AI has rolled out Sarvam Vision 2.1, an upgraded version of its vision-language model designed to understand and process documents. Now it comes with improved capabilities for reading complex tables, extracting information from structured forms, and recognising handwritten text in Indian languages.
The company says Vision 2.1 addresses key limitations observed in its earlier version. And the focus this time is not just on reading documents, but also on turning messy, unstructured scans into clean, structured data that businesses can integrate directly into production workflows.
What can Sarvam Vision 2.1 do
The new model ships with improved tabular processing. In simpler terms, the model can work with nested, complex, and multi-page table structures that standard vision models frequently struggle to parse accurately. It can also extract key-value information from forms, such as names, dates, and financial figures, as well as recognize handwritten notes and annotations in Indic languages.
Sarvam noted that specific training interventions were made to reduce hallucinations and spatial inconsistencies, two common failure points for AI systems dealing with dense or irregular document scans.
Benchmark performance
Alongside the launch, the company introduced a new Sarvam Indic OCR benchmark covering all 22 official Indian languages. The evaluation dataset contains 6,909 samples, including 6,609 samples across the 22 regional languages and 300 English baselines. The material spans newspapers, brochures, textbooks, and historical archives dating from 1800 to the present day.
On this benchmark, Sarvam Vision 2.1 recorded an overall word accuracy of 87.39 per cent. The model also scored 87.30 on the global olmOCR-Bench, demonstrating competitive performance on international document parsing standards.
Training and deployment
Sarvam trained the new model on a mix of synthetic and real-world document scans, leaning especially on regional handwriting variations and form-extraction cases. It then fine-tuned the model with supervised learning before layering on reinforcement learning with verifiable rewards (RLVR), a step aimed at pushing the model toward exact structural and textual accuracy, rather than just plausible-looking output.
The model is accessible through Sarvam’s document intelligence APIs, which include developer tools for converting multi-page files into structured text and automatically extracting data from forms and tables.
- Ends