Skip to main content

The use of ChatGPT among lawyers is now widespread: document summaries, draft preparation, initial exploration of a legal topic. But when AI needs to sit at the centre of legal work, on case-law research and case-specific analysis, is a generalist model enough? And is a generic legal AI truly sufficient? The LegalITA benchmark by Aptus.AI Research, conducted on 107 real Italian law cases across 11 systems including generalists, legal-tech tools and Aptus.AI, provides the first structured measurement of what changes across these three categories, and on which dimensions the choice carries direct professional implications.

Table of contents

The legal AI market is growing rapidly. ChatGPT and generalist models on one side, vertical legal AI tools on the other, and finally platforms like Aptus.AI that combine frontier intelligence with a proprietary regulatory database. The risk for decision-makers is relying on commercial promises rather than measurable data. Making an informed choice requires a structured comparison, with independent dimensions and explicit criteria.

What a generalist model like ChatGPT can (and cannot) do in legal work

ChatGPT performs well on reformulation, summarisation and initial exploration tasks. The difficulties emerge when the answer needs to rest on a concrete source. A study by Stanford RegLab and Human-Centered AI Institute (Dahl et al., 2024) documented that generalist models hallucinate on verifiable case-law queries at rates equal to or exceeding 58%. Katz et al. (2024) showed that GPT-4 passes the US bar examination: proof that legal fluency exists, but fluency is not the same thing as verifiable grounding in an authentic jurisprudential source.

The three axes of comparison: reasoning, sources, false confidence

An operational comparison between legal AI tools rests on three capabilities that can diverge from one another: legal reasoning (identifying the relevant principles and applying them to the case), the ability to cite pertinent sources that genuinely support the legal issue in the correct orientation, and the ability to stop when an essential element is missing. A system can be strong on one dimension and weak on the others. This applies to generalist models and to generic legal AI tools alike, which is precisely why a single average score would obscure the operational difference. The LegalITA white paper on legal AI measures each of these dimensions separately across 107 real cases.

The LegalITA benchmark measures these three dimensions on 107 real Italian law cases, with 67 open-ended jurisprudential questions and 40 adversarial trap questions designed to test abstention. Source: LegalITA, a proprietary benchmark by Aptus.AI. The comparison involves 11 systems: frontier generalist models, legal-tech competitors specialising in Italian law, and Next-OS by Aptus.AI. Across the three axes, the results reveal three distinct performance tiers.

Legal reasoning on real cases

On legal reasoning, Next-OS by Aptus.AI achieves a full pass on 85.1% of the questions. Frontier generalist models range between 71.6% and 79.1%, with GPT-5.5 (the model powering the latest versions of ChatGPT) at 76.1%. Legal-tech competitors fall in a similar range, between 70.1% and 76.1%. This is the least surprising finding: frontier models now have a high baseline of legal competence, and the gap between the three categories is measured in percentage points, not orders of magnitude. The problem, as the next axis shows, lies elsewhere.

Source existence and pertinence (citation hallucinations)

Here the comparison becomes stark and the gap between the three categories becomes operationally significant. Next-OS provides at least one citation that actually covers the legal issue in 80.6% of cases. Legal-tech competitors range between 17.9% and 41.8%. Generalist models collapse: the best reach 3%, GPT-5.5 produces zero confirmed pertinent citations. A Stanford HAI study (Magesh et al., 2025) documented that even commercial legal research tools with RAG architecture hallucinate between 17% and 33% of citations. A plausible citation is not an existing citation, and an existing citation is not automatically pertinent to the required legal orientation. The distinction has direct implications for the professional liability of the person signing the brief. For a deeper look at how RAG architecture works in legal applications, see our deep dive on generative AI and reliable answer grounding.

Ability to stop when a document is missing (abstention)

The Missing Document Detection test evaluates what happens when the prompt asks the system to analyse a contract or a tax assessment notice that has not been attached. Next-OS stops and requests the document in 85% of cases. Among legal-tech competitors, results vary enormously: from 67.5% for the best to 2.5% for the worst. Generalist models remain low: GPT-5.5 and Gemini 3.5 Flash stop one time in eight. The recurring error is not inattention: the model acknowledges that the document is missing and then proceeds to formulate a strategy as if it had read it. This is the least visible operational risk and the most dangerous one, and labelling yourself as “legal AI” does not make you immune.

Currency, hierarchy and source updates

A legal AI tool can call itself specialised, but what matters is how it manages the hierarchy of norms, currency and revirements (shifts in case-law orientation). A generalist model works on a static knowledge cutoff. A generic legal AI may have updated sources but not necessarily distinguish a consolidated orientation from a superseded one or from an active conflict. Next-OS by Aptus.AI explicitly tracks the status of each jurisprudential orientation and relies on a database covering over 200 regulatory authorities across 7 jurisdictions, updated daily. In professional advisory work, this distinction can change the defensibility of the position taken.

Security, confidentiality and data handling

Security is a practical concern across all three categories. Consumer versions of generalist models rarely offer adequate contractual guarantees. Generic legal AI tools vary: some offer encryption, others do not clearly disclose their data training policies. Aptus.AI enforces a no-training policy on client data, AES-256 and TLS 1.2+ encryption, EU cloud infrastructure with confirmed data residency and ISO/IEC 27001:2022 certification verified by an independent third party. For a professional subject to privilege and GDPR obligations, verifying security policies is not a detail but a contractual prerequisite.

How to compare tools objectively: the scorecard

A scorecard should keep the dimensions separate. Collapsing everything into a single score is the quickest way to build a comparison that looks rigorous but hides the critical errors. The LegalITA benchmark, for instance, does not provide a composite score: it publishes three separate primary metrics, each with its own interpretive threshold.

Minimum metrics and weights

Each criterion needs a test, a measurable metric and an explicit threshold.

CriterionTestMetricMinimum thresholdEvidence
ReasoningOpen-ended questions on real cases% of questions fully passed≥ 80%Independent panel
Pertinent citationVerification of court, division, number, year, orientation% of answers with at least one source that holds≥ 60%Official legal database search
AbstentionPrompts with declared but unattached documents% of answers that stop and request≥ 70%Verifiable trap cases
SecurityContractual policy on training and data residencyCertification + DPA availableISO/IEC 27001:2022Third-party audit
Source coverageRelevant authorities and case lawAuthorities monitored and update frequencyDaily updatePublic documentation

Why a single average hides critical errors

A system can score 90% on reasoning and 5% on the ability to cite pertinent sources. An arithmetic average returns it as a strong system, but for a lawyer who needs to sign a brief, it is a dangerous one. The risk sits in the weakest dimension, not in the average: the comparison should be built on independent dimensions with minimum thresholds on each. This applies to generalist models and to legal AI tools that call themselves specialised alike.

A reliable comparison rests on real cases, not staged demos, and on independent evaluation. The LegalITA benchmark uses 107 real cases, a panel of three automatic judges with an adaptive majority protocol, and an inter-judge agreement measured at κ = 0.836 across 1,477 valid judgments. The limitations of the method, including sample size and the need for stratified human validation, are explicitly disclosed in the white paper. Transparency about limitations is the first signal of a benchmark’s credibility.

LegalITA benchmark results and Next-OS performance

The benchmark’s key figures, useful as a first orientation before an internal pilot:

  • 107 real Italian law cases (67 jurisprudential + 40 adversarial trap questions)
  • 17,000+ Corte di Cassazione decisions in the source corpus
  • 11 systems tested: generalist models, legal-tech competitors and Next-OS
  • 3 independent dimensions: reasoning, grounding, missing document detection
  • κ = 0.836 inter-judge agreement

Next-OS by Aptus.AI is the only system that exceeds the 80% threshold across all three dimensions: 85.1% on reasoning, 80.6% on pertinent citation coverage, 85.0% on the ability to stop when documents are missing. Legal-tech competitors show intermediate performance, with strengths on individual dimensions but none achieving the same cross-dimensional consistency. Generalist models confirm high legal fluency but source grounding near zero.

Source: LegalITA, a proprietary benchmark by Aptus.AI.

Checklist for comparing AI tools before a pilot

Before adopting a tool, the minimum method consists of six verifiable steps.

  • Select 10-15 representative cases from daily practice (a brief, a due diligence review, a case-law search, a legal opinion).
  • Define in advance what constitutes an acceptable answer for each case, criterion by criterion.
  • Run the same set on ChatGPT, on a generic legal AI tool and on Aptus.AI, without adapting prompts to the tool.
  • Have the answers evaluated by two independent professionals, with explicit criteria.
  • Verify security policies, available DPA and data residency for each candidate.
  • Measure total time (generation + source verification), not just the apparent speed of the response.

An internal pilot is irreplaceable: a public benchmark guides the choice but does not replace it.

Conclusion

The choice between ChatGPT, a generic legal AI and Aptus.AI is not an ideological debate but an operational decision. ChatGPT remains a useful tool at the margins of legal work. Generic legal AI tools add a layer of specialisation, but with highly variable results on sources and abstention. Placing AI at the centre of the work, on case-law research and case analysis, requires a complete infrastructure: a verifiable regulatory database, measured abstention capability and contractual security. The LegalITA benchmark shows where these capabilities stand today, who has them and who does not.

Last updated: July 2026. The information in this article does not constitute legal advice.


Frequently asked questions

Can ChatGPT be used for legal work?

Yes, but for limited tasks: text reformulation, summarisation of documents already read, initial exploration of a topic. It is not reliable when the answer needs to rest on a verifiable case-law citation or on a document that has not been uploaded. The risk of source hallucination remains significant.

Why should Aptus.AI be more reliable than ChatGPT or a generic legal AI?

Because it combines a frontier reasoning model with a proprietary structured jurisprudential database, with citation verification and explicit abstention rules. The LegalITA benchmark shows that on the ability to provide at least one pertinent citation the gap between Next-OS and the best generalist models is over 77 percentage points, and the distance from legal-tech competitors reaches up to 39 points.

How do you objectively compare two legal AI tools?

With a scorecard that measures reasoning, sources and abstention separately, on real cases from your own practice, with independent human evaluation and explicit criteria defined before the test. A single average should be avoided: it hides critical errors on the weakest dimension.

Is an existing citation automatically correct?

No. A citation can be authentic but incorrectly resolved (wrong court, division, number or year), or authentic but pertinent to a different legal orientation. In Italian law, a revirement or an active jurisprudential conflict changes the value of the same decision in relation to the legal issue.

Why does it matter that an AI can say “I don’t know”?

Because real legal work is full of cases with incomplete or not yet available documents. A system that formulates a strategy based on documents it has not read creates an unmanageable professional risk. Abstention is an operational capability, not a limitation. The LegalITA benchmark shows that even some specialised legal AI tools fail on this dimension in 97.5% of cases.

Does a benchmark replace an internal pilot?

No. A benchmark guides the choice and highlights the dimensions to evaluate; an internal pilot confirms it on your own daily practice, with the cases you actually encounter in your firm or organisation.