Skip to content
Logo LogoEvalScope
Docs Blogs
⌘ K
Logo LogoEvalScope
Docs Blogs

🚀 Quick Start

  • Introduction
  • Installation
  • Quick Start
  • Visualization
  • Parameters
  • Supported Benchmarks
    • LLM Benchmarks
      • AA-LCR
      • AGIEval
      • AIME-2024
      • AIME-2025
      • AIME-2026
      • AlpacaEval2.0
      • AMC
      • AnatEM
      • ARC
      • ARC-AGI-2
      • ARC-Challenge-Indic
      • ArenaHard
      • ArXiv-Math
      • ArxivRollBench
      • ArxivRollBench-Full
      • BBH
      • BC2GM
      • BC4CHEMD
      • BC5CDR
      • BhashaBench-Multi (Ayurveda)
      • BhashaBench-Multi (Finance)
      • BhashaBench-Multi (Krishi)
      • BhashaBench-Multi (Legal)
      • BhashaBench-V1 (Ayurveda)
      • BhashaBench-V1 (Finance)
      • BhashaBench-V1 (Krishi)
      • BhashaBench-V1 (Legal)
      • BigCodeBench
      • BigCodeBench-Hard
      • BioMixQA
      • BroadTwitterCorpus
      • C-Eval
      • Chinese-SimpleQA
      • CL-bench
      • CMATH
      • C-MMLU
      • CoinFlip
      • CommonsenseQA
      • Competition-MATH
      • CoNLL2003
      • CoNLL++
      • Copious
      • CrossNER
      • Data-Collection
      • DocMath
      • DrivelologyBinaryClassification
      • DrivelologyMultilabelClassification
      • DrivelologyNarrativeSelection
      • DrivelologyNarrativeWriting
      • DROP
      • EQ-Bench
      • FinNER
      • FRAMES
      • GeneralArena
      • General-MCQ
      • General-QA
      • GeniaNER
      • GPQA-Diamond
      • GSM8K
      • GSM8K-Indic
      • HaluEval
      • HarveyNER
      • HealthBench
      • HellaSwag
      • HellaSwag-Hindi
      • Humanity’s-Last-Exam
      • HMMT25
      • HMMT26
      • HMMT-Nov-2025
      • HumanEval
      • HumanEvalPlus
      • IFBench
      • IFEval
      • IMO-AnswerBench
      • BoolQ-Indic
      • IndicParam
      • IQuiz
      • JNLPBA
      • JNLPBA-Rare
      • KINA
      • Live-Code-Bench
      • LoCoMo
      • LogiQA
      • LongBench-v2
      • LongMemEval
      • MaritimeBench
      • MATH-500
      • MathQA
      • MBPP
      • MBPP-Plus
      • Med-MCQA
      • MGSM
      • MILU
      • Minerva-Math
      • MIT-Movie-Trivia
      • MIT-Restaurant
      • MMLU
      • MMLU-Pro
      • MMLU-Redux
      • MMMLU
      • MRI-MCQA
      • MT-Bench
      • Multi-IF
      • MultiNERD
      • MultiPL-E HumanEval
      • MultiPL-E MBPP
      • MusicTrivia
      • MuSR
      • NCBI
      • Needle-in-a-Haystack
      • $OneMillion-Bench
      • OntoNotes5
      • OpenAI MRCR
      • PerspectiveGap Prompt Writing
      • PerspectiveGap Role Assignment
      • PIQA
      • PLawBench
      • PolyMath
      • PRBench
      • ProcessBench
      • PubMedQA
      • QASC
      • RACE
      • RefCOCO
      • Sanskriti
      • SciCode
      • SciQ
      • Seed-TTS-Eval
      • SimpleQA
      • SIQA
      • SuperGPQA
      • SWE-bench_Lite
      • SWE-bench_Verified
      • SWE-bench_Verified_mini
      • ToolBench-Static
      • TriviaQA
      • TriviaQA-Indic-MCQ
      • TruthfulQA
      • TweeBankNER
      • TweetNER7
      • Winogrande
      • WMT2024++
      • WNUT2017
      • ZebraLogicBench
    • VLM Benchmarks
      • A-OKVQA
      • AI2D
      • AIR-Bench-Chat
      • AIR-Bench-Foundation
      • BabyVision
      • BLINK
      • CCBench
      • CC-OCR-V2
      • ChartQA
      • CharXiv
      • CMMMU
      • CMMU
      • CommonVoice15
      • CountQA
      • DocVQA
      • EmbSpatial-Bench
      • ERQA
      • FLEURS
      • General-VMCQ
      • General-VQA
      • GSM8K-V
      • HallusionBench
      • HiPhO
      • InfoVQA
      • LibriSpeech
      • LogicVista
      • Maritime-OCR-Bench
      • MathVerse
      • MathVision
      • MathVista
      • MeasureBench
      • MedXpertQA
      • MIA-Bench
      • MicroVQA
      • MMBench
      • MMStar
      • MMAU
      • MMMU
      • MMMU-PRO
      • MSR-VTT
      • MSVD
      • MVBench
      • OCRBench
      • OCRBench-v2
      • olmOCR-Bench
      • OlympiadBench
      • OmniBench
      • OmniDocBench
      • OmniDocBench-v1.6
      • PerceptionBench
      • PhyX-MC
      • PhyX-OE
      • PMC-VQA
      • POPE
      • RealWorldQA
      • Ref-Adv-s
      • ScienceQA
      • ScreenSpot-Pro
      • SEED-Bench-2-Plus
      • SimpleVQA
      • SLAKE
      • SURDS
      • THCHS-30
      • TIR-Bench
      • TORGO
      • TVBench
      • Video-MME-v2
      • VisFactor
      • VisuLogic
      • VLMs Are Biased
      • VQAv2
      • V*Bench
      • VTCBench
      • WenetSpeech
      • WorldVQA
      • ZeroBench
    • AGENT Benchmarks
      • ACEBench
      • AutomationBench
      • BFCL-v3
      • BFCL-v4
      • BrowseComp
      • Claw-Eval
      • DeepSWE
      • DeepSearchQA
      • GAIA
      • GDPval
      • General-FunctionCalling
      • JobBench
      • K2-Vendor-Verifier
      • Kimi-Vendor-Verifier
      • MCP-Atlas
      • MiniMax-Vendor-Verifier
      • MiniWoB
      • OfficeQA
      • ResearchRubrics
      • SkillsBench
      • SWE-bench_Lite_Agentic
      • SWE-bench_Multilingual_Agentic
      • SWE-bench_Pro
      • SWE-bench_Verified_Agentic
      • SWE-bench_Verified_Mini_Agentic
      • τ²-bench
      • τ³-bench
      • τ-bench
      • Terminal-Bench-2.0
      • Terminal-Bench-2.1
      • Toolathlon Official Service Wrapper
      • WideSearch
    • AIGC Benchmarks
      • EvalMuse
      • GEdit-Bench
      • GenAI-Bench
      • General-T2I
      • HPD-v2
      • TIFA-160
    • Other Datasets
      • OpenCompass
      • VLMEvalKit Backend
      • MTEB
      • CLIP-Benchmark
  • ❓ FAQ

🔧 Tutorials

  • Evaluation Backends
    • OpenCompass
    • VLMEvalKit
    • RAGEval
      • MTEB Text Embedding Evaluation
      • CLIP Benchmark
      • RAGAS RAG Evaluation
  • Model Inference Stress Testing
    • Quick Start
    • Parameter
    • Examples
    • Multi-turn Conversation Benchmark
    • SLA Auto-Tuning
    • Speed Benchmark Testing
    • vLLM Bench vs Evalscope Perf Load Testing Comparison
    • AgentX Serving Benchmark
    • Custom Usage
  • AIGC Evaluation
    • Text-to-Image Evaluation
    • Image Editing Evaluation
  • Arena Mode
  • Sandbox Environment Usage
  • Agent Evaluation
    • Native AgentLoop Mode
    • External Agent Bridge Mode
  • EvalScope Service Deployment

🛠️ Advanced Tutorials

  • Building an Evaluation Index
    • Defining Your Schema
    • Sampling Your Index Data
    • Unified Evaluation with Your Index
  • Custom Datasets
    • Large Language Model
    • Multimodal Large Models
    • Custom Text Retrieval Evaluation Dataset
    • CLIP Model
  • Custom Model Evaluation
  • 👍 Contribute Benchmark
  • Sandbox Execution

🧰 Extended Benchmarks

  • Extended Benchmarks
    • Terminal-Bench
    • MiniWoB
    • SkillsBench
    • Toolathlon
    • GAIA
    • WideSearch
    • DeepSearchQA
    • SWE-bench
    • SWE-bench_Pro
    • τ-bench
    • τ²-bench
    • τ³-bench
    • BFCL-v3
    • BFCL-v4
    • Needle in a Haystack
    • ToolBench
    • LongBench-Write

📖 Best Practices

  • Best Practices
    • Evaluating in the Wild: How Agentic Is Your AI Model Really?
    • Benchmark Smarter: Tailor Your Model Evaluation Suite with EvalScope
    • Best Practices for Evaluating the Qwen3-Omni Model
    • Evaluating the Qwen3-VL Model
    • Evaluating the Qwen3-Next Model
    • GPT-OSS Model Evaluation
    • Evaluating Qwen3-Coder+Instruct Model
    • Evaluating Text-to-Image Models
    • Evaluating the Qwen3 Model
    • Evaluating the QwQ Model
    • How Smart is Your AI? Full Assessment of IQ and EQ!
    • Evaluating the Thinking Efficiency of Models
    • Evaluating the Inference Capability of R1 Models
    • Full-Chain LLM Training
    • ms-swift Integration

🧪 Benchmark Results

  • Benchmarking
    • MMLU
  • Speed Benchmarking
    • QwQ-32B-Preview

🌟 Blog

  • Welcome to the EvalScope Blogs!
    • RAG Evaluation Survey: Framework, Metrics, and Methods
EvalScope
/
Supported Benchmarks
/
LLM Benchmarks
/
BhashaBench-V1 (Legal)

BhashaBench-V1 (Legal)#

Overview#

BhashaBench-Legal is the predecessor of BhashaBench-Multi’s legal domain: a domain-specific multiple-choice benchmark evaluating LLM knowledge of Indian law, covering English and Hindi.

Task Description#

  • Task Type: Domain-Specific Multiple-Choice Question Answering

  • Input: An Indian law question with 4 answer choices, in English or Hindi

  • Output: Correct answer letter

  • Languages: English, Hindi

Key Features#

  • 5,600–17,000 questions per language, covering English and Hindi only

  • Predecessor of BhashaBench-Multi: same domains, narrower language coverage

  • Each domain is a separate repository, with English and Hindi as separate configs

Evaluation Notes#

  • Default configuration uses 0-shot evaluation (test split, the only split available)

  • Use subset_list to evaluate a single language (e.g., ['Hindi'])

  • Requires access to this gated dataset - on ModelScope (the default hub), accept the terms and ensure you’re logged in; alternatively, set dataset_hub to huggingface and use HF_TOKEN after accepting the terms on huggingface.co

  • For broader language coverage of the same domain, see bhasha_bench_multi_legal (22 Indic languages, not gated)

Properties#

Property

Value

Benchmark Name

bhashabenchv1_legal

Dataset ID

bharatgenai/BhashaBench-Legal

Paper

N/A

Tags

Knowledge, MCQ, MultiLingual

Metrics

accuracy

Default Shots

0-shot

Evaluation Split

test

Data Statistics#

Metric

Value

Total Samples

24,365

Prompt Length (Mean)

513.88 chars

Prompt Length (Min/Max)

229 / 4628 chars

Per-Subset Statistics:

Subset

Samples

Prompt Mean

Prompt Min

Prompt Max

English

17,047

539.36

233

4628

Hindi

7,318

454.52

229

1748

Sample Example#

Subset: English

{
  "input": [
    {
      "id": "6e1ae42b",
      "content": "Answer the following multiple choice question. The entire content of your response should be of the following format: 'ANSWER: [LETTER]' (without quotes) where [LETTER] is one of A,B,C,D.\n\nPower to amend the issue or frame additional issues prior to passing of a decree vests in a Court by virtue of which provision of the Code of Civil Procedure, 1908?\n\nA) Order XIV Rule 1\nB) Order XIV Rule 5\nC) Order XIV Rule 6\nD) Section 151"
    }
  ],
  "choices": [
    "Order XIV Rule 1",
    "Order XIV Rule 5",
    "Order XIV Rule 6",
    "Section 151"
  ],
  "target": "B",
  "id": 0,
  "group_id": 0,
  "metadata": {
    "language": "English",
    "topic": "Procedural Law"
  }
}

Prompt Template#

Prompt Template:

Answer the following multiple choice question. The entire content of your response should be of the following format: 'ANSWER: [LETTER]' (without quotes) where [LETTER] is one of {letters}.

{question}

{choices}

Usage#

Using CLI#

evalscope eval \
    --model YOUR_MODEL \
    --api-url OPENAI_API_COMPAT_URL \
    --api-key EMPTY_TOKEN \
    --datasets bhashabenchv1_legal \
    --limit 10  # Remove this line for formal evaluation

Using Python#

from evalscope import run_task
from evalscope.config import TaskConfig

task_cfg = TaskConfig(
    model='YOUR_MODEL',
    api_url='OPENAI_API_COMPAT_URL',
    api_key='EMPTY_TOKEN',
    datasets=['bhashabenchv1_legal'],
    dataset_args={
        'bhashabenchv1_legal': {
            # subset_list: ['English', 'Hindi']  # optional, evaluate specific subsets
        }
    },
    limit=10,  # Remove this line for formal evaluation
)

run_task(task_cfg=task_cfg)
BhashaBench-V1 (Krishi)
BigCodeBench

On this page

  • Overview
  • Task Description
  • Key Features
  • Evaluation Notes
  • Properties
  • Data Statistics
  • Sample Example
  • Prompt Template
  • Usage
    • Using CLI
    • Using Python

© 2022-2024, Alibaba ModelScope Built with Sphinx 9.1.0