Skip to content
Logo LogoEvalScope
Docs Blogs
⌘ K
Logo LogoEvalScope
Docs Blogs

🚀 Quick Start

  • Introduction
  • Installation
  • Quick Start
  • Visualization
  • Parameters
  • Supported Benchmarks
    • LLM Benchmarks
      • AA-LCR
      • AGIEval
      • AIME-2024
      • AIME-2025
      • AIME-2026
      • AlpacaEval2.0
      • AMC
      • AnatEM
      • ARC
      • ARC-AGI-2
      • ARC-Challenge-Indic
      • ArenaHard
      • ArXiv-Math
      • ArxivRollBench
      • ArxivRollBench-Full
      • BBH
      • BC2GM
      • BC4CHEMD
      • BC5CDR
      • BhashaBench-Multi (Ayurveda)
      • BhashaBench-Multi (Finance)
      • BhashaBench-Multi (Krishi)
      • BhashaBench-Multi (Legal)
      • BhashaBench-V1 (Ayurveda)
      • BhashaBench-V1 (Finance)
      • BhashaBench-V1 (Krishi)
      • BhashaBench-V1 (Legal)
      • BigCodeBench
      • BigCodeBench-Hard
      • BioMixQA
      • BroadTwitterCorpus
      • C-Eval
      • Chinese-SimpleQA
      • CL-bench
      • CMATH
      • C-MMLU
      • CoinFlip
      • CommonsenseQA
      • Competition-MATH
      • CoNLL2003
      • CoNLL++
      • Copious
      • CrossNER
      • Data-Collection
      • DocMath
      • DrivelologyBinaryClassification
      • DrivelologyMultilabelClassification
      • DrivelologyNarrativeSelection
      • DrivelologyNarrativeWriting
      • DROP
      • EQ-Bench
      • FinNER
      • FRAMES
      • GeneralArena
      • General-MCQ
      • General-QA
      • GeniaNER
      • GPQA-Diamond
      • GSM8K
      • GSM8K-Indic
      • HaluEval
      • HarveyNER
      • HealthBench
      • HellaSwag
      • HellaSwag-Hindi
      • Humanity’s-Last-Exam
      • HMMT25
      • HMMT26
      • HMMT-Nov-2025
      • HumanEval
      • HumanEvalPlus
      • IFBench
      • IFEval
      • IMO-AnswerBench
      • BoolQ-Indic
      • IndicParam
      • IQuiz
      • JNLPBA
      • JNLPBA-Rare
      • KINA
      • Live-Code-Bench
      • LoCoMo
      • LogiQA
      • LongBench-v2
      • LongMemEval
      • MaritimeBench
      • MATH-500
      • MathQA
      • MBPP
      • MBPP-Plus
      • Med-MCQA
      • MGSM
      • MILU
      • Minerva-Math
      • MIT-Movie-Trivia
      • MIT-Restaurant
      • MMLU
      • MMLU-Pro
      • MMLU-Redux
      • MMMLU
      • MRI-MCQA
      • MT-Bench
      • Multi-IF
      • MultiNERD
      • MultiPL-E HumanEval
      • MultiPL-E MBPP
      • MusicTrivia
      • MuSR
      • NCBI
      • Needle-in-a-Haystack
      • $OneMillion-Bench
      • OntoNotes5
      • OpenAI MRCR
      • PerspectiveGap Prompt Writing
      • PerspectiveGap Role Assignment
      • PIQA
      • PLawBench
      • PolyMath
      • PRBench
      • ProcessBench
      • PubMedQA
      • QASC
      • RACE
      • RefCOCO
      • Sanskriti
      • SciCode
      • SciQ
      • Seed-TTS-Eval
      • SimpleQA
      • SIQA
      • SuperGPQA
      • SWE-bench_Lite
      • SWE-bench_Verified
      • SWE-bench_Verified_mini
      • ToolBench-Static
      • TriviaQA
      • TriviaQA-Indic-MCQ
      • TruthfulQA
      • TweeBankNER
      • TweetNER7
      • Winogrande
      • WMT2024++
      • WNUT2017
      • ZebraLogicBench
    • VLM Benchmarks
      • A-OKVQA
      • AI2D
      • AIR-Bench-Chat
      • AIR-Bench-Foundation
      • BabyVision
      • BLINK
      • CCBench
      • CC-OCR-V2
      • ChartQA
      • CharXiv
      • CMMMU
      • CMMU
      • CommonVoice15
      • CountQA
      • DocVQA
      • EmbSpatial-Bench
      • ERQA
      • FLEURS
      • General-VMCQ
      • General-VQA
      • GSM8K-V
      • HallusionBench
      • HiPhO
      • InfoVQA
      • LibriSpeech
      • LogicVista
      • Maritime-OCR-Bench
      • MathVerse
      • MathVision
      • MathVista
      • MeasureBench
      • MedXpertQA
      • MIA-Bench
      • MicroVQA
      • MMBench
      • MMStar
      • MMAU
      • MMMU
      • MMMU-PRO
      • MSR-VTT
      • MSVD
      • MVBench
      • OCRBench
      • OCRBench-v2
      • olmOCR-Bench
      • OlympiadBench
      • OmniBench
      • OmniDocBench
      • OmniDocBench-v1.6
      • PerceptionBench
      • PhyX-MC
      • PhyX-OE
      • PMC-VQA
      • POPE
      • RealWorldQA
      • Ref-Adv-s
      • ScienceQA
      • ScreenSpot-Pro
      • SEED-Bench-2-Plus
      • SimpleVQA
      • SLAKE
      • SURDS
      • THCHS-30
      • TIR-Bench
      • TORGO
      • TVBench
      • Video-MME-v2
      • VisFactor
      • VisuLogic
      • VLMs Are Biased
      • VQAv2
      • V*Bench
      • VTCBench
      • WenetSpeech
      • WorldVQA
      • ZeroBench
    • AGENT Benchmarks
      • ACEBench
      • AutomationBench
      • BFCL-v3
      • BFCL-v4
      • BrowseComp
      • Claw-Eval
      • DeepSWE
      • DeepSearchQA
      • GAIA
      • GDPval
      • General-FunctionCalling
      • JobBench
      • K2-Vendor-Verifier
      • Kimi-Vendor-Verifier
      • MCP-Atlas
      • MiniMax-Vendor-Verifier
      • MiniWoB
      • OfficeQA
      • ResearchRubrics
      • SkillsBench
      • SWE-bench_Lite_Agentic
      • SWE-bench_Multilingual_Agentic
      • SWE-bench_Pro
      • SWE-bench_Verified_Agentic
      • SWE-bench_Verified_Mini_Agentic
      • τ²-bench
      • τ³-bench
      • τ-bench
      • Terminal-Bench-2.0
      • Terminal-Bench-2.1
      • Toolathlon Official Service Wrapper
      • WideSearch
    • AIGC Benchmarks
      • EvalMuse
      • GEdit-Bench
      • GenAI-Bench
      • General-T2I
      • HPD-v2
      • TIFA-160
    • Other Datasets
      • OpenCompass
      • VLMEvalKit Backend
      • MTEB
      • CLIP-Benchmark
  • ❓ FAQ

🔧 Tutorials

  • Evaluation Backends
    • OpenCompass
    • VLMEvalKit
    • RAGEval
      • MTEB Text Embedding Evaluation
      • CLIP Benchmark
      • RAGAS RAG Evaluation
  • Model Inference Stress Testing
    • Quick Start
    • Parameter
    • Examples
    • Multi-turn Conversation Benchmark
    • SLA Auto-Tuning
    • Speed Benchmark Testing
    • vLLM Bench vs Evalscope Perf Load Testing Comparison
    • AgentX Serving Benchmark
    • Custom Usage
  • AIGC Evaluation
    • Text-to-Image Evaluation
    • Image Editing Evaluation
  • Arena Mode
  • Sandbox Environment Usage
  • Agent Evaluation
    • Native AgentLoop Mode
    • External Agent Bridge Mode
  • EvalScope Service Deployment

🛠️ Advanced Tutorials

  • Building an Evaluation Index
    • Defining Your Schema
    • Sampling Your Index Data
    • Unified Evaluation with Your Index
  • Custom Datasets
    • Large Language Model
    • Multimodal Large Models
    • Custom Text Retrieval Evaluation Dataset
    • CLIP Model
  • Custom Model Evaluation
  • 👍 Contribute Benchmark
  • Sandbox Execution

🧰 Extended Benchmarks

  • Extended Benchmarks
    • Terminal-Bench
    • MiniWoB
    • SkillsBench
    • Toolathlon
    • GAIA
    • WideSearch
    • DeepSearchQA
    • SWE-bench
    • SWE-bench_Pro
    • τ-bench
    • τ²-bench
    • τ³-bench
    • BFCL-v3
    • BFCL-v4
    • Needle in a Haystack
    • ToolBench
    • LongBench-Write

📖 Best Practices

  • Best Practices
    • Evaluating in the Wild: How Agentic Is Your AI Model Really?
    • Benchmark Smarter: Tailor Your Model Evaluation Suite with EvalScope
    • Best Practices for Evaluating the Qwen3-Omni Model
    • Evaluating the Qwen3-VL Model
    • Evaluating the Qwen3-Next Model
    • GPT-OSS Model Evaluation
    • Evaluating Qwen3-Coder+Instruct Model
    • Evaluating Text-to-Image Models
    • Evaluating the Qwen3 Model
    • Evaluating the QwQ Model
    • How Smart is Your AI? Full Assessment of IQ and EQ!
    • Evaluating the Thinking Efficiency of Models
    • Evaluating the Inference Capability of R1 Models
    • Full-Chain LLM Training
    • ms-swift Integration

🧪 Benchmark Results

  • Benchmarking
    • MMLU
  • Speed Benchmarking
    • QwQ-32B-Preview

🌟 Blog

  • Welcome to the EvalScope Blogs!
    • RAG Evaluation Survey: Framework, Metrics, and Methods
EvalScope
/
Supported Benchmarks
/
LLM Benchmarks
/
BhashaBench-Multi (Legal)

BhashaBench-Multi (Legal)#

Overview#

BhashaBench-Multi (Legal) is a domain-specific multiple-choice benchmark evaluating LLM knowledge of Indian law across 22 Indic languages. Each question originates in English and is machine translated (with LLM-judged translation quality scores) into the target language; this adapter uses the translated question/choices.

Task Description#

  • Task Type: Domain-Specific Multiple-Choice Question Answering

  • Input: A Indian law question with 4 answer choices, in one of 22 Indic languages

  • Output: Correct answer letter

  • Languages: Assamese, Bengali, Bodo, Dogri, Gujarati, Hindi, Kannada, Kashmiri, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali, Oriya, Punjabi, Sanskrit, Santhali, Sindhi, Tamil, Telugu, Urdu

Key Features#

  • ~14,963 questions per language across 22 Indic languages per domain (~330k total per domain)

  • Machine-translated from English with LLM-judged translation quality scores

  • 22 scheduled languages of India, all in native script; no English split

  • Four domains available as separate benchmarks: Ayurveda, Finance, Krishi, Legal

Evaluation Notes#

  • Default configuration uses 0-shot evaluation (test split, the only split available)

  • Use subset_list to evaluate specific languages (e.g., ['Hindi', 'Tamil']), or limit to cap sample count — each domain is ~14,963 questions per language across 22 languages (~330k total), so evaluating every language’s full split is a large run

  • No English split exists for this dataset

Properties#

Property

Value

Benchmark Name

bhasha_bench_multi_legal

Dataset ID

bharatgenai/BhashaBench-Multi

Paper

N/A

Tags

Knowledge, MCQ, MultiLingual

Metrics

accuracy

Default Shots

0-shot

Evaluation Split

test

Data Statistics#

Metric

Value

Total Samples

536,030

Prompt Length (Mean)

490.89 chars

Prompt Length (Min/Max)

225 / 6384 chars

Per-Subset Statistics:

Subset

Samples

Prompt Mean

Prompt Min

Prompt Max

Assamese

24,365

475.08

232

2556

Bengali

24,365

482.5

235

2066

Bodo

24,365

521.23

225

4608

Dogri

24,365

487.72

225

4432

Gujarati

24,365

463.64

232

1954

Hindi

24,365

489.37

232

2202

Kannada

24,365

475.35

232

2068

Kashmiri

24,365

514.54

242

5037

Konkani

24,365

473.95

225

4000

Maithili

24,365

473.11

225

4011

Malayalam

24,365

511.29

236

2218

Manipuri

24,365

548.36

238

6384

Marathi

24,365

487.86

232

2113

Nepali

24,365

475.87

234

2058

Oriya

24,365

458.86

232

1936

Punjabi

24,365

483.03

232

2138

Sanskrit

24,365

489.0

233

1979

Santhali

24,365

549.75

233

5074

Sindhi

24,365

455.74

234

1830

Tamil

24,365

522.97

237

2479

Telugu

24,365

481.32

234

1992

Urdu

24,365

479.13

235

2120

Sample Example#

Subset: Assamese

{
  "input": [
    {
      "id": "e631cc6e",
      "content": "Answer the following multiple choice question. The entire content of your response should be of the following format: 'ANSWER: [LETTER]' (without quotes) where [LETTER] is one of A,B,C,D.\n\nকোনো আদেশ প্ৰকাশ কৰাৰ পূৰ্বতে কোনো সমস্যা সংশোধন কৰাৰ বা নতুন সমস্যা উত্থাপন কৰাৰ ক্ষমতা আদালতৰ ওচৰত থাকে, আৰু এই ক্ষমতা দিয়া হয় দেৱানী প্রক্রিয়া বিধি, ১৯০৮-ৰ কোনটো ব্যৱস্থাৰ দ্বাৰা?\n\nA) অধ্যায় ১৪, বিধি ১\nB) অধ্যায় ১৪, বিধি ৫\nC) অধ্যায় XIV, বিধি ৬\nD) ধাৰা ১৫১"
    }
  ],
  "choices": [
    "অধ্যায় ১৪, বিধি ১",
    "অধ্যায় ১৪, বিধি ৫",
    "অধ্যায় XIV, বিধি ৬",
    "ধাৰা ১৫১"
  ],
  "target": "B",
  "id": 0,
  "group_id": 0,
  "metadata": {
    "language": "Assamese",
    "topic": "Procedural Law"
  }
}

Prompt Template#

Prompt Template:

Answer the following multiple choice question. The entire content of your response should be of the following format: 'ANSWER: [LETTER]' (without quotes) where [LETTER] is one of {letters}.

{question}

{choices}

Usage#

Using CLI#

evalscope eval \
    --model YOUR_MODEL \
    --api-url OPENAI_API_COMPAT_URL \
    --api-key EMPTY_TOKEN \
    --datasets bhasha_bench_multi_legal \
    --limit 10  # Remove this line for formal evaluation

Using Python#

from evalscope import run_task
from evalscope.config import TaskConfig

task_cfg = TaskConfig(
    model='YOUR_MODEL',
    api_url='OPENAI_API_COMPAT_URL',
    api_key='EMPTY_TOKEN',
    datasets=['bhasha_bench_multi_legal'],
    dataset_args={
        'bhasha_bench_multi_legal': {
            # subset_list: ['Assamese', 'Bengali', 'Bodo']  # optional, evaluate specific subsets
        }
    },
    limit=10,  # Remove this line for formal evaluation
)

run_task(task_cfg=task_cfg)
BhashaBench-Multi (Krishi)
BhashaBench-V1 (Ayurveda)

On this page

  • Overview
  • Task Description
  • Key Features
  • Evaluation Notes
  • Properties
  • Data Statistics
  • Sample Example
  • Prompt Template
  • Usage
    • Using CLI
    • Using Python

© 2022-2024, Alibaba ModelScope Built with Sphinx 9.1.0