Trust is an interface problem as much as a model problem. Five design decisions that determine whether reviewers accept or quietly ignore AI output.
Sophie DevereuxHead of Legal AI··6 min read
Legal AI deployments fail more often for interface reasons than for model reasons. The findings are good; the workflow makes them expensive to check, so reviewers stop checking, and then stop trusting.
1. Show the source, in place
Every finding should open to its passage in the document, with surrounding context, in one action. If verifying a finding requires searching the PDF, the tool has externalised its cost onto the reviewer — and reviewers respond by accepting findings unread, which is the worst possible outcome.
2. Rank by consequence, not by confidence
A high-confidence finding about a notices clause matters less than a medium-confidence finding about an uncapped indemnity. Ordering by model confidence puts the trivial first. Ordering by potential impact respects the reviewer’s time and matches how lawyers actually triage.
3. Make disagreement cheap and structured
Reviewers should be able to reject a finding in one click and say why in a controlled vocabulary: not present, misclassified, present but immaterial, correct but already known. Free-text rejection notes are never analysed. Structured ones show you exactly where the configuration is wrong.
4. Show what was not found
A list of findings implies the rest of the document is fine. It does not mean that. Display the clause inventory alongside the findings: what was located, what was expected and absent, and what could not be classified. Absence is where the risk hides, and it should be visible without being asked for.
5. Never hide the model version
A report produced in March under one model release and one playbook version is a different artefact from the same report produced in September. Stamp both on the output. When a review is questioned later, this is the first thing anyone will want and the last thing anyone thought to record.
The underlying principle
Design for the reviewer who is sceptical and short of time, because that is every reviewer worth having. A system that makes checking fast earns trust it can then spend. A system that makes checking slow gets trusted anyway, for the wrong reason, and that is when something goes badly wrong.
A necessary note
This article is general information about legal technology and practice, not legal advice, and it does not create a lawyer–client relationship. JuriPro is a technology company, not a law firm. Take advice from a qualified lawyer admitted in the relevant jurisdiction before acting on anything here.
Share
Sophie Devereux
Head of Legal AI, JuriPro
Solicitor of England & Wales turned applied researcher, responsible for how JuriPro models are evaluated against practitioner judgement.
A plain-English account of how a language model reads an agreement, where its judgement is genuinely useful, and the four failure modes every reviewing lawyer should know about.