ROHAS

ROHAS Legal AI (India) Benchmark

10 questions test reasoning, authority accuracy, honesty about missing information, applied judgment and instruction fidelity for a total of 100 marks. Two serious failures are tracked separately as disqualifiers.

R Reasoning & Risk O Origin & Accuracy H Honesty about gaps A Applied context S Structure & Fidelity

Run configuration

Temperature0 where supported; provider default otherwise
Web & toolsDisabled
Version 3.1Provider-default reasoning · 4,000 answer tokens · no retry
Version 3.2Low reasoning · 6,000 answer tokens · one technical retry
Benchmark costAnswer calls only; judge excluded
Final gradingAdministrator reviewed

How to read these results

Performance is measured on a fixed Indian-law benchmark. Results reflect the particular model, provider, configuration, test date and administrator-reviewed grading used in that run. They do not establish a model's accuracy or suitability for every legal task. This is a research benchmark, not legal advice, and does not create a lawyer-client relationship. Law is stated as at the date disclosed in the benchmark. Model names and trademarks belong to their respective owners; ROHAS is not affiliated with or endorsed by any model provider.

ROHAS Legal AI (India) Benchmark Leaderboard

A model that triggers either explicit disqualifier cannot rank #1, however high its raw score.

Order by:
#1
Anthropic: Claude Opus 5
anthropic/claude-opus-5 · v3.2
Winner
98%
R 100%
O 100%
H 90%
A 100%
S 100%
Answer cost $0.4864 Runtime 5m 51s
v3.2 · 28 August 2026
#2
OpenAI: GPT-5.6 Sol
openai/gpt-5.6-sol · v3.1
97%
R 95%
O 100%
H 95%
A 100%
S 95%
Answer cost $0.1955 Runtime 6m 4s
v3.1 · 28 August 2026
#3
OpenAI: GPT-5.6 Terra
openai/gpt-5.6-terra · v3.1
95%
R 100%
O 90%
H 95%
A 95%
S 95%
Answer cost $0.1778 Runtime 5m 22s
v3.1 · 28 August 2026
#4
OpenAI: GPT-5.6 Luna
openai/gpt-5.6-luna · v3.1
94%
R 95%
O 95%
H 90%
A 90%
S 100%
Answer cost $0.0189 Runtime 5m 6s
v3.1 · 28 August 2026
#5
Anthropic: Claude Sonnet 5
anthropic/claude-sonnet-5 · v3.1
92%
R 95%
O 95%
H 85%
A 90%
S 95%
Answer cost $0.1638 Runtime 5m 6s
v3.1 · 28 August 2026
#6
MoonshotAI: Kimi K3
moonshotai/kimi-k3 · v3.2
91%
R 95%
O 95%
H 85%
A 85%
S 95%
Answer cost $0.1135 Runtime 2m 39s
v3.2 · 28 August 2026
#7
Qwen: Qwen3.8 Max
qwen/qwen3.8-max · v3.2
89%
R 95%
O 85%
H 85%
A 80%
S 100%
Answer cost $0.0961 Runtime 6m 19s
v3.2 · 28 August 2026
#8
DeepSeek: DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813 · v3.2
87%
R 90%
O 90%
H 75%
A 85%
S 95%
Answer cost $0.1275 Runtime 13m 37s
v3.2 · 28 August 2026
#9
SpaceXAI: Grok 4.6
x-ai/grok-4.6 · v3.2
86%
R 90%
O 90%
H 80%
A 80%
S 90%
Answer cost $0.0459 Runtime 3m 26s
v3.2 · 28 August 2026
#10
Z.ai: GLM 5.3
z-ai/glm-5.3 · v3.2
83%
R 85%
O 80%
H 80%
A 75%
S 95%
Answer cost $0.0306 Runtime 3m 18s
v3.2 · 28 August 2026
#11
Google: Gemini 3.7 Flash
google/gemini-3.7-flash · v3.2
80%
R 95%
O 60%
H 70%
A 80%
S 95%
Answer cost $0.0398 Runtime 2m 42s
v3.2 · 29 August 2026
#12
Perplexity: Sonar
perplexity/sonar · v3.2
76%
R 85%
O 75%
H 65%
A 60%
S 95%
Answer cost $0.0552 Runtime 1m 55s
v3.2 · 29 August 2026
#13
Anthropic: Claude Haiku 4.5
anthropic/claude-haiku-4.5 · v3.1
72%
R 85%
O 70%
H 70%
A 45%
S 90%
Answer cost $0.0309 Runtime 3m 11s
v3.1 · 28 August 2026
#14
Amazon: Nova Premier 1.0
amazon/nova-premier-v1 · v3.2
61%
R 85%
O 80%
H 55%
A 20%
S 65%
Answer cost $0.0420 Runtime 1m 56s
v3.2 · 29 August 2026

Full data

RankModelVersionScoreR / O / H / A / SAnswer costRuntimeRun
#1 Anthropic: Claude Opus 5
anthropic/claude-opus-5
v3.2 98 / 100 (98%) 100% / 100% / 90% / 100% / 100% $0.4864 5m 51s 2026-08-28 10:50:44
#2 OpenAI: GPT-5.6 Sol
openai/gpt-5.6-sol
v3.1 97 / 100 (97%) 95% / 100% / 95% / 100% / 95% $0.1955 6m 4s 2026-08-28 15:35:21
#3 OpenAI: GPT-5.6 Terra
openai/gpt-5.6-terra
v3.1 95 / 100 (95%) 100% / 90% / 95% / 95% / 95% $0.1778 5m 22s 2026-08-28 08:50:44
#4 OpenAI: GPT-5.6 Luna
openai/gpt-5.6-luna
v3.1 94 / 100 (94%) 95% / 95% / 90% / 90% / 100% $0.0189 5m 6s 2026-08-28 17:15:09
#5 Anthropic: Claude Sonnet 5
anthropic/claude-sonnet-5
v3.1 92 / 100 (92%) 95% / 95% / 85% / 90% / 95% $0.1638 5m 6s 2026-08-28 16:44:37
#6 MoonshotAI: Kimi K3
moonshotai/kimi-k3
v3.2 91 / 100 (91%) 95% / 95% / 85% / 85% / 95% $0.1135 2m 39s 2026-08-28 11:27:05
#7 Qwen: Qwen3.8 Max
qwen/qwen3.8-max
v3.2 89 / 100 (89%) 95% / 85% / 85% / 80% / 100% $0.0961 6m 19s 2026-08-28 11:36:23
#8 DeepSeek: DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813
v3.2 87 / 100 (87%) 90% / 90% / 75% / 85% / 95% $0.1275 13m 37s 2026-08-28 11:22:16
#9 SpaceXAI: Grok 4.6
x-ai/grok-4.6
v3.2 86 / 100 (86%) 90% / 90% / 80% / 80% / 90% $0.0459 3m 26s 2026-08-28 11:02:48
#10 Z.ai: GLM 5.3
z-ai/glm-5.3
v3.2 83 / 100 (83%) 85% / 80% / 80% / 75% / 95% $0.0306 3m 18s 2026-08-28 11:41:34
#11 Google: Gemini 3.7 Flash
google/gemini-3.7-flash
v3.2 80 / 100 (80%) 95% / 60% / 70% / 80% / 95% $0.0398 2m 42s 2026-08-29 03:48:55
#12 Perplexity: Sonar
perplexity/sonar
v3.2 76 / 100 (76%) 85% / 75% / 65% / 60% / 95% $0.0552 1m 55s 2026-08-29 03:43:30
#13 Anthropic: Claude Haiku 4.5
anthropic/claude-haiku-4.5
v3.1 72 / 100 (72%) 85% / 70% / 70% / 45% / 90% $0.0309 3m 11s 2026-08-28 09:19:06
#14 Amazon: Nova Premier 1.0
amazon/nova-premier-v1
v3.2 61 / 100 (61%) 85% / 80% / 55% / 20% / 65% $0.0420 1m 56s 2026-08-29 03:58:30