Skip to main content

Frontier-model intelligence. Normative-database rigour. No one else has both.

Reasoning → Grounding (verifiable sources) → the zone where delegation becomes possible General-purpose models reason, but don't cite Traditional legal-tech sources, but dated intelligence Next-OS reasons and cites
Reasoning × grounding: only Next-OS occupies the quadrant where delegating becomes rational.

107

real-world Italian law tasks

17,000

decisions of the Corte di Cassazione

11

systems put to the test

3

independent dimensions
  • The test
    67 open-ended legal questions across civil, labour and tax law, evaluated criterion by criterion by a panel of three independent judges. A task passes only if every required criterion is met: no partial credit.
  • What it enables
    This is the difference between an AI that writes and an AI you can entrust with analysis: identifying the relevant principles and applying them to the case at hand is the work itself, not its packaging.
  • The result
    Next-OS fully passes 85.1% of tasks, first among the 11 systems. And there is a number that changes the market: the best general-purpose model, Claude Opus 4.8 (79.1%), outperforms every legal-tech platform in the field. Having the right sources is no longer enough if the intelligence using them is a generation behind.
Reasoning: tasks passed in full
All-pass on 67 tasks · higher is better
The line marks where the best general-purpose model overtakes the entire legal-tech field.
  • The test
    Citing a real decision isn’t enough. The citation must be correctly resolved (court, division, number, year) and must genuinely support the issue at stake, in the right orientation: the consolidated one, the one following a reversal, or one of the relevant orientations where the conflict is still open.
  • What it enables
    Professional responsibility cannot be delegated: the signature stays yours. If verifying an answer means redoing the research from scratch, you haven’t saved time: you’ve moved it.
  • The result
    This is where the gap becomes dramatic. Next-OS provides at least one decision that actually covers the issue in 80.6% of cases. The best general-purpose models stop at 3%. GPT-5.5 at zero. With a general-purpose model, source verification stays entirely on your shoulders. With Next-OS, it becomes a check.
Grounding: at least one citation that holds
Coverage on the 67 case-law tasks · higher is better
The best general-purpose models stay at ≤ 3%: they can reason, but they cannot prove.
  • The test
    40 trap questions. Each asks the system to analyse a specific document (a contract, a tax assessment notice, a judgment) that is never attached. Only one behaviour passes the test: stopping, flagging that the document is missing, and asking for it.
  • What it enables
    Real work is not a demo environment: documents go missing, arrive late, arrive incomplete. That is the condition for letting AI work without supervising every step.
  • The result
    Next-OS stops and asks for the documents in 85% of cases. The best general-purpose model, Claude Sonnet 5, does so just over half the time (52.5%). GPT-5.5 and Gemini 3.5 Flash one time in eight. The most common failure isn’t inattention: many systems acknowledge that the document is missing, and then formulate a strategy as if they had read it.
Stopping: it recognises the missing document and asks for it
Global MDD on 40 trap questions · higher is better
Doesn't notice

invents an analysis of the absent document

Answers anyway

notices it, but builds a strategy regardless

Stops and asks

the only behaviour that passes the test

107
real-world tasks
137
evaluation criteria
3
independent judges
κ 0.836
inter-judge agreement