AI-native practice · 1 October 2026
Choosing Legal AI Tools: What I Wish I'd Known Sooner
The mistakes I made picking tools early on were not about capability. They were about not yet knowing what question I was actually supposed to be asking each vendor.
I have already written up the evaluation process I actually run before recommending a legal AI tool to a client: the worst-document test, the ambiguous-case test, the questions I ask the sales team directly. That process is the correct version of the story. This is the honest one, about what it took to arrive at it, and the specific mistakes along the way that taught me each piece of it the expensive way rather than the easy way.
The first mistake was believing the demo
Early on, I evaluated tools the way I imagine most lawyers still do: I watched the demo, I asked a couple of clarifying questions, and if the answer was competent, I moved forward. What I did not yet understand was that a demo is not evidence about the tool, it is evidence about the specific example the vendor has rehearsed. The first time this actually cost a client something, the tool performed beautifully on the sales call and then degraded badly on a real, slightly messy document three weeks into an actual matter, in a way that was entirely predictable if I had thought to test for it beforehand. I had not thought to, because nobody had told me a demo needed to be treated as marketing rather than data, and it had simply never occurred to me on my own.
The second mistake was optimising for capability instead of failure mode
For a while, my whole evaluation process was really one question dressed up as several: can this tool do the thing. That is the wrong question, or at least a dangerously incomplete one, because almost every serious legal AI tool can do the thing under good conditions. What actually separates a tool worth trusting from one that will eventually embarrass a client is what happens when it cannot do the thing. Does it say so, plainly, or does it produce a fluent, confident answer anyway, dressed in exactly the same tone it uses when it is right. I learned to weight this heavily only after watching a tool confidently misstate something in a way that looked, on the page, indistinguishable from its correct output. Nothing about the interface signalled the difference. That gap between confident and correct is the single most consequential thing I now test for, and I did not know to look for it until I had already been burned by not looking.
The third mistake was letting pricing structure slide past me unexamined
I used to treat pricing as a separate conversation from the actual evaluation, something to negotiate after deciding the tool was good. I no longer separate the two, because the shape of a vendor's pricing tells you something honest about the product that the sales conversation usually does not. A tool priced simply and predictably per unit of work usually reflects a vendor who understands their own cost to serve. A tool with a deceptively low headline price and a long list of usage tiers and add-on modules is very often signalling that the core product is incomplete on its own, and that you will discover the real price once switching costs have already made you captive. I paid that tax more than once before I started reading pricing structure as information rather than as a footnote.
The fourth mistake was not bringing my own data
It took embarrassingly long for me to start insisting on testing tools against a client's actual, messy, real documents rather than whatever clean sample data the vendor supplied. Once I started doing this consistently, the single most useful signal I got was not from any tool's performance, it was from how a vendor reacted to the request. A confident vendor accommodates it without friction. A vendor who resists, or stalls, or insists their sample data is representative enough, is telling you something true about how their tool performs outside its best-case conditions, and I now treat that resistance itself as one of the most reliable data points in the entire evaluation.
What actually changed
None of this came from reading a framework somewhere and applying it. It came from picking wrong a few times, watching exactly how the wrong choice failed, and building a rule specific enough to prevent that particular failure from happening again. That is a slower way to learn than reading a good checklist, and I would genuinely rather a lawyer read the checklist and skip the expensive part, which is exactly why I wrote the more procedural version of this separately. But I think it is worth saying plainly what the checklist does not fully convey: every rule in it exists because something specific and avoidable went wrong first, on a real matter, with a real client's work depending on it, and I would rather you never have that specific moment yourself.
The one thing I would tell myself at the start
Evaluate the tool as though you are going to be the one explaining its wrong answer to a judge or an angry client someday, because eventually you will be, whichever tool you pick. That framing changes what you actually pay attention to during an evaluation. It stops being about which tool is more impressive in a demo, and starts being about which tool will fail in a way you can catch, explain and recover from, because in legal work every tool eventually fails on something. The only real choice you are making is which failure mode you are willing to live with.
legal ai · ai native practice · tool selection · legal tech