← All guides

10 min read · Updated August 2026

How I evaluate a legal AI tool before recommending it to a client

The demo is designed to work. My job is figuring out whether the tool works on the messy, real version of the task, not the clean version they showed you.

Key takeaways

  • Every vendor demo is run on their best-case data. Ask to run it on your worst-case data instead, before you sign anything.
  • The sales pitch answers 'can it do this,' the question that matters is 'what does it do when it can't, and does it tell you.'
  • Pricing structure often reveals more about the real cost of the tool than the feature list does.
  • I weight how a tool fails more heavily than how it succeeds, because failure mode is what determines whether you catch a mistake.
  • Reference calls with existing customers who match your practice size and area beat any spec sheet.

The demo tells you almost nothing

Every vendor I've sat through a demo with has shown me the tool working beautifully, and I don't say that cynically, it's just the nature of a demo. They've tested that specific example, on that specific data, dozens of times before showing it to me. That's not deception, it's just not information about how the tool will behave on your firm's actual, messy, inconsistently formatted documents six months from now.

The single most useful thing I do in any vendor evaluation is ask, before the call even happens, to bring my own sample: a real, anonymized document from the kind of matter the client actually handles, with the kind of formatting quirks and edge cases that real documents have. A surprising number of vendors resist this, and that resistance is itself a data point worth weighing heavily.

What I actually run the tool through

  1. 1. The worst document in the file, not the best one

    I ask the client to pull the messiest real example they have: a scanned document with poor OCR, an inconsistent template, a document missing a section. Vendor tools that only work on clean inputs will tell you a lot by how badly they degrade, and how honestly they signal that degradation, on a rough one.

  2. 2. A deliberately ambiguous case

    I include an edge case where the correct answer genuinely depends on judgment the tool shouldn't have, to see whether it flags uncertainty or confidently produces a wrong answer with the same tone of confidence as a right one. A tool that hedges appropriately is more trustworthy than one that's right more often but never admits doubt.

  3. 3. The same input twice

    I run the identical input through the tool on two separate days and compare outputs. Meaningful drift on an identical input, with no clear reason, tells you something about how much you can rely on the tool for anything that needs to be reproducible, which in legal work is most things.

  4. 4. What happens at the edges of the stated scope

    I deliberately give the tool something slightly outside what it claims to do, to see whether it declines gracefully or produces a confident, wrong answer anyway. This matters more than almost anything else, because the tools that fail badly are the ones that don't know they're failing.

The questions I actually ask the sales team

  • What happens to our data, specifically: is it used for training, how long is it retained, and can you show me the contract language, not just the marketing page, that says so.
  • Who are three customers of a similar size and practice area I can call directly, not a case study, an actual phone number.
  • What's the failure rate on your own internal testing, and how do you define a failure. If they don't have an answer ready, that's an answer.
  • What happens when the underlying model you're built on changes. Some tools quietly degrade or shift behavior when the vendor upgrades an underlying model, and I want to know if there's a change log and a way to catch that.
  • How does pricing scale as our usage grows, specifically at two and five times current volume, because the pitch price is rarely the number that matters a year in.

Reading the pricing structure for information

Pricing tells you almost as much as the product demo does. A tool priced per document processed, cleanly and predictably, usually reflects a vendor confident in a consistent cost to serve. A tool priced with a low headline number and a lot of usage tiers, overage charges, and add-on modules is often signaling that the core product isn't complete on its own, and you'll discover the real price once you're already dependent on it.

I also pay close attention to how a vendor prices seats versus usage. Per-seat pricing for a tool a small practice would use occasionally is often a bad fit, because you end up paying for access nobody uses most weeks. I'd rather see usage-based pricing that scales with actual value delivered, even if the per-unit number looks higher on paper.

A useful vendor red flag

If a sales rep can't clearly explain, in plain language, what their tool does when it's uncertain, walk away from that specific evaluation regardless of how good the demo looked. A tool that can't describe its own failure mode hasn't been tested rigorously enough by the people who built it, and you'll be the one discovering that mode for them.

How this fits into building something durable

When I'm helping a client assemble what I call a practice pack, a set of tools and workflows built specifically for how their practice actually runs, tool selection is the step people want to rush, because it feels like shopping rather than building. I push back on that instinct every time. The tool you choose here is the foundation everything else gets built on top of, and swapping it out after six months of workflows are built around it is expensive in a way that's easy to underestimate up front.

My rule of thumb: spend as much calendar time evaluating the tool as you expect to spend using it in the first month. If that sounds like too much time, it usually means the evaluation was too shallow, not that the rule is wrong.

Questions

How many vendors should I actually evaluate before deciding?
Three is usually enough to see real variation in approach and pricing, without turning the process into a research project on its own. More than five and you're usually procrastinating on the decision rather than gathering meaningfully new information.
Is it reasonable to ask a vendor for a free trial with real firm data?
Yes, and I'd treat pushback on this as a real signal. A confident vendor will let you pilot with representative, appropriately anonymized data before you commit. If they insist you can only test with their sample data, ask yourself why.
What should I do if two tools tie on every test I run?
Break the tie on support and roadmap transparency rather than features. Ask each vendor directly what they're building next and how often they ship, then call a reference customer and ask how support actually responded the last time something broke.
How much should I trust a vendor's own benchmark numbers?
Treat them as a floor, not a ceiling, and always as measured on their chosen data, not yours. I use them to decide which tools are worth the time to test directly, never as the basis for a final decision on their own.
Should I involve the whole team in the evaluation, or just decide myself?
Involve whoever will use the tool daily in the actual hands-on testing, even if the final decision sits with one person. The people doing the work will surface friction a spec sheet never will, and their buy-in matters for adoption after the fact.

Want this built for your practice, not just read about it?

Book an intro call