Field note · 9 November 2026
My Top 10 Questions Before Recommending Any Legal AI Tool
Before I recommend any AI tool to a practice, it has to answer these ten questions honestly. Most tools fail at least two of them.
People ask me to evaluate a lot of AI tools, usually after they have already been pitched one and want a second opinion before committing a practice's data and workflow to it. Over time the questions I actually ask have settled into a fixed list. If a tool cannot answer these ten honestly, I do not recommend it, no matter how good the demo looked.
1. Where does the output actually come from?
Not "what model does it use," but can every claim the tool produces be traced back to a specific source document, or is it generating from general training data dressed up to look case-specific. If the answer is unclear or the vendor gets vague, that is disqualifying on its own. This is the same anchoring standard I hold my own systems to, described fully in how to stop AI hallucination with evidence anchoring.
2. What happens to the data after it leaves the practice?
Where does the input go, is it used to train the vendor's model on other customers' behalf, is it retained after the session ends, and can you actually get a straight answer to this in the vendor's documentation rather than in a sales call where the answer is whatever gets the deal closed. If the data policy requires a call to clarify, that is itself information.
3. Does it know the matter, or only what gets pasted into it?
A general chatbot only knows what a human manually provides each time. A tool worth recommending is wired into the actual matter record, client, deadlines, documents, so it is working from real context rather than from whatever got typed into a box that session. This is the actual difference between a tool and a toy, and it is the dividing line I described in AI agents for lawyers.
4. What does a wrong answer look like, and how would you catch it?
I ask the vendor this directly, and I ask it of myself before recommending anything. If nobody can describe what a plausible-but-wrong output looks like for this specific tool, that is a sign nobody has actually stress-tested it against real failure modes, only against demo cases chosen to succeed.
5. Does it require a review step, or does it invite skipping one?
Some tools are designed with the review step built into the workflow, output is clearly marked as a draft, sources are visible, nothing ships without a click confirming it was checked. Others are designed to feel finished, confident tone, no visible sourcing, which quietly invites the reviewer to skip the step that actually matters. I have watched the second kind cause real problems, described in my top 10 AI mistakes I've watched law firms make.
6. What is the actual failure rate, and who measured it?
A vendor-reported accuracy number, without knowing the test set or the methodology, tells you almost nothing. I ask for specifics, or I run my own small test against a handful of real, messy matters before recommending anything at scale.
7. What happens when it is wrong in a way that matters?
Not the general disclaimer language, the actual operational question: if this tool produces a wrong deadline calculation or a fabricated citation, what is the practice's exposure, and does the tool's design make that failure loud and obvious or quiet and easy to miss.
8. Does it lock you in, or can you leave with your data intact?
Can the practice export its data cleanly if it switches tools later, or is the workflow built in a way that makes leaving expensive enough that the practice stays out of inertia rather than satisfaction. I have opinions about vendor lock-in generally, and this is where they matter most, because a practice's matter data is not something to get stuck inside a tool that stops earning its keep.
9. Is the pricing model aligned with the practice actually using it well, or with usage volume regardless of quality?
Some pricing models reward a vendor for maximum usage regardless of whether the output is good, which creates a quiet incentive misalignment. I prefer models where the vendor's incentive is for the tool to be genuinely trusted and adopted, not merely used as often as possible.
10. Would I actually let this touch a matter I am personally responsible for?
This is the question underneath all the others, and it is the one I actually ask myself before recommending anything. If I would not trust this tool's output on a matter with my own name on it, unreviewed, I do not recommend it to anyone else's practice either, no matter how the pitch went.
Most tools fail at least two of these
In practice, most AI tools I evaluate fail question one or question three outright, no real source anchoring, no real matter context, which means everything past that point is somewhat moot. The tools that pass all ten are rarer than the marketing around this space would suggest, and that scarcity is itself useful information: it tells you the bar most vendors are actually clearing is lower than the confidence of their demos implies. The AI-Native Practice Field Guide goes deeper into how to evaluate and sequence AI adoption across an entire practice, not just a single tool purchase, if this list raised more questions than it answered.
AI tools · vetting · legal technology · due diligence