5 competencies. 10 questions. 2 automatic disqualifiers. ROHAS is a practical Legal AI screening benchmark that you can run against any AI model in minutes.
Ranked by score on each model's most recent run. Disqualified models failed a hallucination/accuracy trap question and can never rank #1, however high their raw score.
| Rank | Model | Score | Pass/Partial/Fail | Answer cost | Runtime | Latest run |
|---|---|---|---|---|---|---|
| #1 |
xAI: Grok 4.5 x-ai/grok-4.5 |
9.5 / 10 (95%) | 9 / 1 / 0 | $0.1216 | 5m 55s | 2026-07-08 12:18:47 |
| #2 |
Anthropic: Claude Opus 4.8 anthropic/claude-opus-4.8 |
9.5 / 10 (95%) | 9 / 1 / 0 | $0.2557 | 3m 31s | 2026-07-08 11:26:44 |
| #3 |
Anthropic: Claude Fable 5 anthropic/claude-fable-5 |
9.5 / 10 (95%) | 9 / 1 / 0 | $0.6753 | 4m 26s | 2026-07-08 11:38:19 |
| #4 |
OpenAI: GPT-5.6 Sol Pro openai/gpt-5.6-sol-pro |
9.5 / 10 (95%) | 9 / 1 / 0 | $1.2238 | 10m 44s | 2026-07-20 10:11:26 |
| #5 |
OpenAI: GPT-5.6 Luna openai/gpt-5.6-luna |
9 / 10 (90%) | 8 / 2 / 0 | $0.0665 | 2m 9s | 2026-07-13 01:51:10 |
| #6 |
Anthropic: Claude Haiku 4.5 anthropic/claude-haiku-4.5 |
8.5 / 10 (85%) | 7 / 3 / 0 | $0.0288 | 1m 53s | 2026-07-08 11:30:52 |
| #7 |
Z.ai: GLM 5.2 z-ai/glm-5.2 |
8 / 10 (80%) | 6 / 4 / 0 | $0.0210 | 6m 27s | 2026-07-10 09:12:43 |
| #8 |
Qwen: Qwen3.7 Plus qwen/qwen3.7-plus |
8 / 10 (80%) | 7 / 2 / 1 | $0.0389 | 10m 28s | 2026-07-08 12:46:42 |
| #9 |
Google: Gemini 3.5 Flash google/gemini-3.5-flash |
7.5 / 10 (75%) | 6 / 3 / 1 | $0.0842 | 2m 24s | 2026-07-08 12:25:25 |
| #10 |
Anthropic: Claude Sonnet 5 anthropic/claude-sonnet-5 |
7 / 10 (70%) | 4 / 6 / 0 | $0.1320 | 4m 6s | 2026-07-08 11:17:30 |
| #11 |
Meta: Llama 4 Maverick meta-llama/llama-4-maverick |
6.5 / 10 (65%) | 4 / 5 / 1 | $0.0028 | 2m 37s | 2026-07-08 12:29:17 |
| #12 |
DeepSeek: DeepSeek V4 Pro deepseek/deepseek-v4-pro |
6.5 / 10 (65%) | 5 / 3 / 2 | $0.0131 | 7m 56s | 2026-07-08 09:26:09 |
| #13 |
Amazon: Nova 2 Lite amazon/nova-2-lite-v1 |
6.5 / 10 (65%) | 4 / 5 / 1 | $0.0292 | 1m 42s | 2026-07-18 09:45:43 |
| DQ |
OpenAI: GPT-5.5 openai/gpt-5.5 |
7.5 / 10 (75%) | 7 / 1 / 2 | $0.3933 | 9m 55s | 2026-07-07 13:54:29 |
| DQ |
Free Models Router openrouter/free |
5 / 10 (50%) | 3 / 4 / 3 | $0.0000 | 5m 14s | 2026-07-07 09:10:43 |