Logo BHASHABENCH V1

A COMPREHENSIVE BENCH-MARK FOR THE QUADRANT OF INDIC DOMAINS

ABSTRACT

The rapid advancement of large language models (LLMs) has intensified the need for domain and culture specific evaluation. Existing benchmarks are largely Anglocentric and domain-agnostic, limiting their applicability to India-centric contexts. To address this gap, we introduce BhashaBench V1, a large-scale, bilingual, exam-based benchmark covering four India-centric domains, Agriculture, Legal, Finance, and Ayurveda. BhashaBench V1 contains 74,166 meticulously curated question-answer pairs, with 52,494 in English and 21,672 in Hindi, sourced from authentic government and domain-specific exams. It is organized around a quadrant of four domains: Agriculture, Legal, Finance, and Ayurveda, comprising 90+ subdomains and covering 500+ topics, enabling fine-grained evaluation. Evaluation of 29+ LLMs reveals significant domain and language specific performance gaps, with especially large disparities in low-resource domains. For instance, GPT-4o achieves 76.49% overall accuracy in Legal but only 59.74% in Ayurveda. Models consistently perform better on English content compared to Hindi across all domains. Subdomain-level analysis shows that areas such as Cyber Law, International Finance perform relatively well, while Panchakarma, Seed Science, and Human Rights remain notably weak. BhashaBench V1 provides a large-scale, exam-sourced dataset for evaluating large language models across these four Indian knowledge domains. It enables assessment of models' ability to integrate domain-specific knowledge with bilingual understanding. All code, benchmarks, and resources are publicly available to support open research.

BHASHABENCH

The primary motivation behind BhashaBench V1 is to comprehensively assess domain-specific knowledge and reasoning capabilities of large language models within India’s diverse and culturally rich knowledge ecosystems. Unlike existing benchmarks focusing on general or Western-centric domains, our benchmark evaluates specialized Indian knowledge systems requiring deep cultural understanding and contextual awareness. BhashaBench V1 adheres to seven core design principles: (1) Critical Indian Domains: Encompasses Agriculture, Legal systems, Finance, and Ayurveda with fine-grained subfields. (2) Diverse Task Formats: Includes multiple-choice, assertion-reasoning, fill-in-blanks, and comprehension tasks. (3) India-Specific Reasoning: Evaluates domain-specific reasoning incorporating cultural contexts and regional practices. (4) Bilingual Framework: Sup- ports English and Hindi evaluation maintaining cultural authenticity. (5) Authentic Sources: Ques- tions curated from government examinations and professional certifications. (6) Difficulty Assess- ment: Categorized into Easy, Medium, Hard levels. (7) Cultural Authenticity: Prioritizes tradi- tional knowledge systems including Ayurvedic principles. 1 This framework spans 90+ subdomains covering 500+ topics, enabling comprehensive evaluation of model capabilities in India-centric con- texts.

DATA COLLECTION

The data collection process for BhashaBench V1 follows a systematic approach similar to AGIEVAL (Zhong et al., 2023), focusing on authentic examination materials from national and state-level assessments. We systematically gathered publicly available question papers from official online examination portals, which host previously released papers that are manually curated by subject matter experts, ensuring accurate topic tagging, language annotation, and validated answer keys. Our comprehensive collection encompasses over 40 different examination types across multiple categories: national competitive exams, domain-specific degree examinations, professional certification tests, and state-level civil services examinations. Regional state examinations proved particularly valuable as they incorporate state-specific topics, local knowledge systems, and cultural practices often overlooked in national assessments. These examinations are typically taken by individuals seeking higher education opportunities or career advancement, ensuring questions reflect practical, real-world knowledge requirements. The final dataset comprises 74,166 carefully curated question–answer pairs spanning four core domains, with 52,494 questions in English (70.8%) and 21,672 questions in Hindi (29.2%), reflecting practical usage patterns in Indian educational and professional contexts. This approach ensures BhashaBench V1 captures the nuanced intersection between language, culture, and domain expertise essential for effective model deployment in Indian contexts.

Data Processing and Analysis

The BhashaBench V1 dataset was built by extracting structured question-answer pairs from PDF examination papers while carefully preserving linguistic and cultural authenticity. A robust OCR pipeline powered by Surya OCR ensured high-quality text extraction across Indic languages, followed by a custom GPT-based system that structured the content into standardized JSON question-answer pairs. Extensive cleaning removed noisy and duplicate data, verified language integrity, and organized questions into six formats, with missing subdomains automatically classified. Rigorous expert validation further guaranteed accuracy, natural language flow, and cultural relevance. The final dataset contains 74,166 questions across four major domains (Agriculture, Finance, Ayurveda, and Legal) and 91 subdomains, with English (70.8%) and Hindi (29.2%) coverage. Agriculture emphasizes agronomy, Finance highlights quantitative problem solving, Ayurveda captures traditional medicine knowledge, and Legal spans both core jurisprudence and emerging areas like Cyber Law. With over 90% MCQs and a balanced distribution of difficulty levels, BhashaBench V1 stands as a high-quality, domain-rich bilingual benchmark, reflecting substantial technical and validation efforts to ensure both authenticity and reliability.

RESULTS AND DISCUSSIONS

🚨 To submit your results to the leaderboard, please send to this email with your result json files.

Logo BHASHABENCH V1

LIMITATIONS AND BIASES

BhashaBench V1 has several limitations. (1) Language Coverage: The benchmark currently covers English and Hindi, and therefore does not capture India’s full linguistic diversity, including regional variation in agricultural practices, legal terminology, and traditional medicine. Future versions will expand to additional Indian languages. (2) Domain Scope: While the benchmark covers Agriculture, Legal, Finance, and Ayurveda, other India-specific knowledge areas, including traditional crafts, regional governance, and indigenous practices, remain outside its current scope. (3) Evaluation Methodology: The benchmark uses structured questions derived from authentic government and professional examinations, which may not capture all forms of contextual reasoning required in realworld applications. Moreover, open-source and API-based models are evaluated using different scoring protocols (log-likelihood and generative prompting, respectively); cross-protocol comparisons should therefore be interpreted with caution. (4) Validation: Linguistic review was conducted by the in-house linguistics team without credentialed domain experts, independent double annotation, or inter-annotator agreement (IAA). Thus, the benchmark does not provide independent domainlevel verification of technical or factual content. Future versions should incorporate domain experts and overlapping annotation. (5) Prompt Language: Evaluation instructions and prompts were provided in English for both English and Hindi questions, which introduces a potential confound in the Hindi results. However, our localized prompt sensitivity analysis indicates that this has a negligible effect on overall performance. (6) Data Processing and Evaluation Overlap: GPT-OSS120B was used for several data construction and classification steps and is also evaluated on the resulting benchmark, creating a potential circularity. However, a paired ablation on the reconstructed subset showed that GPT-OSS-120B had the smallest reconstruction-driven gain among the three evaluated models, providing evidence against a strong self-favoring reconstruction effect. The benchmark also exhibits three broader sources of bias. Source Material Bias: Despite drawing from diverse authentic sources, some regional practices and emerging developments may be underrepresented. Language Resource Bias: The English and Hindi subsets differ in resource availability and composition; in Agriculture (BBK), for example, the Hindi subset is more concentrated in fewer subdomains and contains relatively easier questions, so the English–Hindi gap should not be attributed to language ability alone. Examination Framework Bias: Reliance on established examination systems may introduce institutional perspectives present in the original assessments. These limitations should be considered when interpreting BhashaBench V1 results.

Ethical Considerations

Societal Impact. BhashaBench V1 is intended to support the evaluation and development of LLMs for India-centric knowledge systems across agriculture, legal services, finance, and Ayurveda. Improved performance on such domains could support applications such as agricultural information access, legal information assistance, financial literacy, and access to traditional medicine knowledge. However, these applications also present risks. Models evaluated on BhashaBench V1 may produce incorrect or misleading information, and over-reliance on automated systems could be harmful in high-stakes settings. Examination-based evaluation may also favor formal educational knowledge over experiential or locally specific knowledge. We therefore recommend appropriate human oversight when models are applied to critical domains and caution against treating benchmark performance as a substitute for professional expertise.

Ethics Statement. BhashaBench V1 is constructed from publicly available government and professional examination materials. We follow the applicable terms and conditions governing the use of these sources and do not intentionally include personally identifiable information. The questionanswer pairs were reviewed against their original source materials for extraction fidelity, answer-key accuracy, and linguistic quality by our in-house linguistics team. This review did not involve domaincredentialed experts or independent double annotation, as discussed in the Limitations. We also acknowledge that the current English and Hindi coverage does not capture the full linguistic and cultural diversity of India. BhashaBench V1 is released to support academic research and educational use. Users should exercise appropriate caution when applying models evaluated with the benchmark, particularly in high-stakes domains, and should maintain human oversight and clearly communicate model limitations in downstream applications.

ACKNOWLEDGEMENTS

We gratefully acknowledge the BharatGen team for
their support, guidance, and contributions throughout this work. We also thank everyone who contributed to the development and evaluation of
BhashaBench V1.