Benchmarking Indian Legal AI using the R.O.H.A.S Test

R.O.H.A.S is a 5-competency, 10-question, 100-mark benchmark for legal AI, built entirely around Indian law and Indian legal practice.

RN
Rohas Nagpal

R.O.H.A.S is a 5-competency, 10-question, 100-mark benchmark for legal AI, built entirely around Indian law and Indian legal practice. Its short, focused questions are designed to expose specific failure modes rather than produce a vague measure of overall quality. Its 10 questions test:

  1. Reasoning & Risk (can it trace a clause with three nested exceptions to the right rupee figure, and rank a buried unlimited indemnity above cosmetic issues?)
  2. Origin & Accuracy (does it cite a real case correctly, and refuse to invent a statutory section that doesn't exist?)
  3. Honesty about gaps (does it identify the missing facts that determine an Indian employment claim, and name the specific transaction documents that are missing rather than advising blind?)
  4. Applied context (does it apply section 27 of the Indian Contract Act to a post-employment restraint, and balance MSME recovery rights against customer-concentration risk in plain English?)
  5. Structure & Fidelity (can it hold to an exact output format, and refuse to confirm a false legal premise even when a user asserts it confidently and asks it to "just confirm").

2 of the 10 questions (the fabricated-statute trap and the false-premise trap) have explicit disqualifier triggers. A model is flagged only when it triggers the stated condition, independently of the marks awarded to its answer.

ROHAS India is a screening tool. Ten questions can show which models to eliminate and which deserve a deeper, practice-specific Indian-law battery, but they cannot certify any model as safe to rely on without professional review.

R = Reasoning & Risk

This tests legal logic, nested conditional parsing, issue spotting, risk detection and severity weighting.

R1 — The nested conditional chain

R2 — Risk detection & severity ranking (Indian Law)

O = Origin & Accuracy

This tests whether the model's citations are real, whether its statement of the law is accurate, and whether it fabricates authority when none exists.

O1 — Real citation with accurate holding (Indian Law)

O2 — Hallucination trap: a provision that doesn't exist (Indian Law)

H = Honesty about gaps

This tests whether the model recognizes what it doesn't know, asks for missing facts instead of guessing, and names missing documents rather than advising on an incomplete record.

H1 — Missing facts in an Indian employment dismissal

H2 — Missing documents: the schedules that were never supplied (M&A)

A = Applied context

This tests whether the model applies the right jurisdiction's law, weighs commercial reality alongside legal merit, and can explain its advice in language a non-lawyer can actually use.

A1 — A post-employment restraint under Indian law

A2 — Commercial judgment delivered in plain English (Indian law)

S = Structure & Fidelity

This tests whether the model follows instructions exactly, represents its sources faithfully, and resists being talked into a conclusion it hasn't actually verified.

S1 — Format discipline and source fidelity (Indian appeal)

S2 — The planted false premise (Indian Law)

How scoring works

Each answer is marked from 0 to 10 against a fixed reference answer and an itemised rubric whose marks total 10. Question-specific caps prevent a model from receiving a high mark after missing a decisive issue. The legal reference material is dated so later runs can be audited when the law changes.

That yields a 100-mark ROHAS India score, with a 20-mark sub-score for each letter. The two binary disqualifiers remain separate from the numerical score: a model may receive whatever marks it earned and still be flagged if it fabricated the statutory provision or accepted the planted false premise.

An AI judge can propose a score and short reason, but the administrator reviews and enters the final score before a complete ten-question run can be published to the leaderboard.

Test conditions

Version 3.2 is an India-only benchmark. Every published run uses the same disclosed configuration: temperature 0 where the selected model supports it (otherwise its provider default), low reasoning effort, no web search or external tools, a maximum of 6,000 answer tokens and a maximum of 2,000 judge tokens. An empty, content-filtered or token-limited response is retried once; a repeated technical failure cannot be scored or published, and only the affected answer needs to be rerun. Published benchmark cost includes every tested-model answer attempt; judge cost is retained separately for audit and excluded from that total. The judge is advisory; the administrator's reviewed marks and disqualifier decisions control the published result. Version 3.1 results remain archived separately because their run configuration differs.

Results reflect the particular model, provider, configuration, test date and manual grading used in that run. They do not establish accuracy or suitability for every legal task. ROHAS India is a research benchmark, not legal advice, and does not create a lawyer-client relationship. The S1 judgment extract is synthetic. Model names and trademarks belong to their respective owners; ROHAS is not affiliated with or endorsed by any model provider.