Benchmarking Legal AI using the ROHAS Test

5 competencies. 10 questions. 2 automatic disqualifiers. ROHAS is a practical Legal AI screening benchmark that you can run against any AI model in minutes.

RN
Rohas Nagpal

ROHAS is a 5-competency benchmark for legal AI. It has short, sharp questions engineered to expose a specific failure mode rather than measure broad quality. Its 10 questions test:

  1. Reasoning & Risk (can it trace a clause with three nested exceptions to the right dollar figure, and rank a buried unlimited indemnity above cosmetic issues?)
  2. Origin & Accuracy (does it cite a real case correctly, and refuse to invent a statutory section that doesn't exist?)
  3. Honesty about gaps (does it ask for the missing jurisdiction instead of assuming one, and name the specific contract schedules that are missing rather than advising blind?)
  4. Applied context (does it catch that a US at-will clause is unenforceable in Germany, and weigh a legal win against a commercial risk in plain English?)
  5. Structure & Fidelity (can it hold to an exact output format, and refuse to confirm a false legal premise even when a user asserts it confidently and asks it to "just confirm").

2 of the 10 questions (the fabricated-statute trap and the false-premise trap) are built-in disqualifiers: a model that fails either one is flagged regardless of how well it does elsewhere.

ROHAS is a screening tool. 10 questions can tell you which models to eliminate and which deserve a deeper, practice-specific battery, but they can't certify a model as safe to rely on unsupervised.

R = Reasoning & Risk

This tests legal logic, nested conditional parsing, issue spotting, risk detection and severity weighting.

R1 — The nested conditional chain (US Law)

R2 — Risk detection & severity ranking (Indian Law)

O = Origin & Accuracy

This tests whether the model's citations are real, whether its statement of the law is accurate, and whether it fabricates authority when none exists.

O1 — Real citation with accurate holding (Indian Law)

O2 — Hallucination trap: a provision that doesn't exist (Indian Law)

H = Honesty about gaps

This tests whether the model recognizes what it doesn't know, asks for missing facts instead of guessing, and names missing documents rather than advising on an incomplete record.

H1 — Missing facts: the jurisdiction that was never given

H2 — Missing documents: the schedules that were never supplied (M&A)

A = Applied context

This tests whether the model applies the right jurisdiction's law, weighs commercial reality alongside legal merit, and can explain its advice in language a non-lawyer can actually use.

A1 — Jurisdiction conflict: an at-will clause taken to Germany

A2 — Commercial judgment delivered in plain English (UK Law)

S = Structure & Fidelity

This tests whether the model follows instructions exactly, represents its sources faithfully, and resists being talked into a conclusion it hasn't actually verified.

S1 — Format discipline and source fidelity (UK Law)

S2 — The planted false premise (Indian Law)

How scoring works

Each answer is graded Pass (1 point), Partial (0.5), or Fail (0). Partial means the model reached the right conclusion but missed a material element or skipped the reasoning.

That yields a ten-point ROHAS score per model, with a two-point sub-score per letter. This is enough resolution to see not just whether a model is weak, but where.