Frontier-model intelligence. Normative-database rigour. No one else has both.
To bring AI into the core of legal work, not just its margins, a professional needs to trust three things: that it reasons correctly, that every answer is grounded in verifiable sources, and that it knows when to stop because the elements are missing. LegalITA, the first three-dimensional benchmark on Italian law, measures all three. Next-OS ranks first on each: and that is what makes the switch to AI-native work possible.
85.1% vs 79.1%
+6.0 points over Claude Opus 4.8
Even when compared to one of the best general-purpose models, Next-OS leads the way in legal reasoning.
80.6% vs 41.8%
+38.8 points over Competitor A
Nearly twice as many answers
with a relevant, verifiable source.
85.0% vs 67.5%
+17.5 points over Competitor C
When an essential document is missing, Next-OS knows when to stop: sooner and better than any other system tested.
THE PROBLEM
AI has entered law firms and legal departments. But it has stayed at the margins of the work.
Most legal and compliance professionals already use AI: to summarise, to prepare a draft, for a first exploration. But the core of the work, case-law research, case analysis, the position to defend, remains manual. Not out of resistance to change: out of sound professional judgement. Until now, you had to choose between intelligence and sources.
Frontier general-purpose models reason extraordinarily well, but when they need to anchor their answers to real case law they invent, or cite authentic decisions that don’t apply: and on a filing, the signature is yours, not the model’s. Traditional legal-tech platforms have the sources, but a previous-generation intelligence, now overtaken on reasoning by the very general-purpose models they were meant to beat.
Delegating the margins of the work to AI requires no trust. Delegating its core does. LegalITA was built to measure exactly the conditions of that trust.
107
17,000
11
3
THE METHOD
The three conditions for delegation.
Condition 01
Does it reason like a professional?
Otherwise you can only delegate the form.
- The test
67 open-ended legal questions across civil, labour and tax law, evaluated criterion by criterion by a panel of three independent judges. A task passes only if every required criterion is met: no partial credit. - What it enables
This is the difference between an AI that writes and an AI you can entrust with analysis: identifying the relevant principles and applying them to the case at hand is the work itself, not its packaging. - The result
Next-OS fully passes 85.1% of tasks, first among the 11 systems. And there is a number that changes the market: the best general-purpose model, Claude Opus 4.8 (79.1%), outperforms every legal-tech platform in the field. Having the right sources is no longer enough if the intelligence using them is a generation behind.
Condition 02
Can it prove what it says?
Otherwise verification costs more than delegation saves.
- The test
Citing a real decision isn’t enough. The citation must be correctly resolved (court, division, number, year) and must genuinely support the issue at stake, in the right orientation: the consolidated one, the one following a reversal, or one of the relevant orientations where the conflict is still open. - What it enables
Professional responsibility cannot be delegated: the signature stays yours. If verifying an answer means redoing the research from scratch, you haven’t saved time: you’ve moved it. - The result
This is where the gap becomes dramatic. Next-OS provides at least one decision that actually covers the issue in 80.6% of cases. The best general-purpose models stop at 3%. GPT-5.5 at zero. With a general-purpose model, source verification stays entirely on your shoulders. With Next-OS, it becomes a check.
Condition 03
Does it know when to stop?
Otherwise you can’t trust it on real cases.
- The test
40 trap questions. Each asks the system to analyse a specific document (a contract, a tax assessment notice, a judgment) that is never attached. Only one behaviour passes the test: stopping, flagging that the document is missing, and asking for it. - What it enables
Real work is not a demo environment: documents go missing, arrive late, arrive incomplete. That is the condition for letting AI work without supervising every step. - The result
Next-OS stops and asks for the documents in 85% of cases. The best general-purpose model, Claude Sonnet 5, does so just over half the time (52.5%). GPT-5.5 and Gemini 3.5 Flash one time in eight. The most common failure isn’t inattention: many systems acknowledge that the document is missing, and then formulate a strategy as if they had read it.
THE TAKEAWAY
It takes two ingredients. No one else has both. We do.
The LegalITA results explain why, until now, the switch wasn’t rational, and the reason is structural, not circumstantial.
Why vertical systems don’t reason at frontier level
Traditional legal-tech platforms are built around a model: they query it, feed it a few retrieved documents, and package the answer. But frontier-model capabilities don’t emerge from calling an API: they emerge from the execution environment the model works in. Wrapping a model caps its capabilities; building the execution environment unlocks them.
Why frontier models can’t cite
Their knowledge of the law is baked into their weights, not anchored to queryable sources. And placing a document archive next to them isn’t enough: for a citation to hold, the underlying database must be machine-readable, with every rule and every decision in a structure the machine can query, verify and trace, with its validity over time and its relationships. It is the infrastructure Aptus.AI has been building for years.
Next-OS is built to sit exactly at that intersection
An execution environment that takes frontier models to their full potential, anchored to Aptus.AI’s machine-readable normative database, where every source is verified, resolved and traceable. It reasons better than the vertical systems, cites better than the general-purpose models, and stops when professional honesty requires it. That combination is what moves AI from the margin to the centre: from drafting to research, from assistant to collaborator.
TRANSPARENCY
A professional decision is made on verifiable evidence.
That is why the full benchmark is documented in a public white paper: method, criteria, evaluation protocol, anonymised legal-tech competitors, and explicitly stated limitations, including sample size and the need for stratified human validation.
Public code repository on GitHub.