ROHAS Legal AI (India) Benchmark
10 questions test reasoning, authority accuracy, honesty about missing information, applied judgment and instruction fidelity for a total of 100 marks. Two serious failures are tracked separately as disqualifiers.
Run configuration
How to read these results
Performance is measured on a fixed Indian-law benchmark. Results reflect the particular model, provider, configuration, test date and administrator-reviewed grading used in that run. They do not establish a model's accuracy or suitability for every legal task. This is a research benchmark, not legal advice, and does not create a lawyer-client relationship. Law is stated as at the date disclosed in the benchmark. Model names and trademarks belong to their respective owners; ROHAS is not affiliated with or endorsed by any model provider.
ROHAS Legal AI (India) Benchmark Leaderboard
A model that triggers either explicit disqualifier cannot rank #1, however high its raw score.
Full data ▼
| Rank | Model | Version | Score | R / O / H / A / S | Answer cost | Runtime | Run |
|---|---|---|---|---|---|---|---|
| #1 |
Anthropic: Claude Opus 5 anthropic/claude-opus-5 |
v3.2 | 98 / 100 (98%) | 100% / 100% / 90% / 100% / 100% | $0.4864 | 5m 51s | 2026-08-28 10:50:44 |
| #2 |
OpenAI: GPT-5.6 Sol openai/gpt-5.6-sol |
v3.1 | 97 / 100 (97%) | 95% / 100% / 95% / 100% / 95% | $0.1955 | 6m 4s | 2026-08-28 15:35:21 |
| #3 |
OpenAI: GPT-5.6 Terra openai/gpt-5.6-terra |
v3.1 | 95 / 100 (95%) | 100% / 90% / 95% / 95% / 95% | $0.1778 | 5m 22s | 2026-08-28 08:50:44 |
| #4 |
OpenAI: GPT-5.6 Luna openai/gpt-5.6-luna |
v3.1 | 94 / 100 (94%) | 95% / 95% / 90% / 90% / 100% | $0.0189 | 5m 6s | 2026-08-28 17:15:09 |
| #5 |
Anthropic: Claude Sonnet 5 anthropic/claude-sonnet-5 |
v3.1 | 92 / 100 (92%) | 95% / 95% / 85% / 90% / 95% | $0.1638 | 5m 6s | 2026-08-28 16:44:37 |
| #6 |
MoonshotAI: Kimi K3 moonshotai/kimi-k3 |
v3.2 | 91 / 100 (91%) | 95% / 95% / 85% / 85% / 95% | $0.1135 | 2m 39s | 2026-08-28 11:27:05 |
| #7 |
Qwen: Qwen3.8 Max qwen/qwen3.8-max |
v3.2 | 89 / 100 (89%) | 95% / 85% / 85% / 80% / 100% | $0.0961 | 6m 19s | 2026-08-28 11:36:23 |
| #8 |
DeepSeek: DeepSeek V4 Pro 0813 deepseek/deepseek-v4-pro-0813 |
v3.2 | 87 / 100 (87%) | 90% / 90% / 75% / 85% / 95% | $0.1275 | 13m 37s | 2026-08-28 11:22:16 |
| #9 |
SpaceXAI: Grok 4.6 x-ai/grok-4.6 |
v3.2 | 86 / 100 (86%) | 90% / 90% / 80% / 80% / 90% | $0.0459 | 3m 26s | 2026-08-28 11:02:48 |
| #10 |
Z.ai: GLM 5.3 z-ai/glm-5.3 |
v3.2 | 83 / 100 (83%) | 85% / 80% / 80% / 75% / 95% | $0.0306 | 3m 18s | 2026-08-28 11:41:34 |
| #11 |
Google: Gemini 3.7 Flash google/gemini-3.7-flash |
v3.2 | 80 / 100 (80%) | 95% / 60% / 70% / 80% / 95% | $0.0398 | 2m 42s | 2026-08-29 03:48:55 |
| #12 |
Perplexity: Sonar perplexity/sonar |
v3.2 | 76 / 100 (76%) | 85% / 75% / 65% / 60% / 95% | $0.0552 | 1m 55s | 2026-08-29 03:43:30 |
| #13 |
Anthropic: Claude Haiku 4.5 anthropic/claude-haiku-4.5 |
v3.1 | 72 / 100 (72%) | 85% / 70% / 70% / 45% / 90% | $0.0309 | 3m 11s | 2026-08-28 09:19:06 |
| #14 |
Amazon: Nova Premier 1.0 amazon/nova-premier-v1 |
v3.2 | 61 / 100 (61%) | 85% / 80% / 55% / 20% / 65% | $0.0420 | 1m 56s | 2026-08-29 03:58:30 |