← All writing

Field note · 16 October 2026

AI Agents for Lawyers: What They Can and Can't Do Yet

An honest capability assessment, not a demo reel. What agentic AI actually handles reliably in legal work today, and what still needs a lawyer's judgment squarely in the loop.

The word "agent" has started doing the same work "AI" did two years ago, meaning almost anything a vendor wants it to mean. I build agentic systems for legal work, MatterOS runs agents on intake, drafting and analysis, so I have a direct stake in the category, which is exactly why I want to be specific rather than promotional about what an agent actually does reliably today and what it does not.

What an agent actually is, for this purpose

An agent, in the sense that matters for legal work, is AI that takes a multi-step task, executes a sequence of actions toward it, checks its own intermediate results against something, and produces an output for a human to review, rather than answering a single question and stopping. That is a meaningfully different thing from a chatbot, and the gap between the two is where most of the hype and most of the real capability both live.

What agents can do reliably right now

Structured intake. An agent can run an intake conversation, ask adaptive follow-up questions based on prior answers, and produce a first-pass matter summary and conflicts check before a lawyer spends a minute on it. This is the single most mature use of agentic AI in legal work I have seen, because intake is bounded, repeatable, and the cost of an imperfect first pass is low, a lawyer reviews it before anything happens.

First-draft assembly from structured facts. Given a matter's structured record and a template, an agent can produce a genuinely useful first draft, not a finished one, pulling facts from the matter rather than requiring them typed in by hand. This works well specifically because the facts already exist as structure. It works much worse when the agent is asked to also extract those facts reliably from messy source material in the same pass, which is why anchoring the extraction step separately matters.

Document extraction against a bounded pool. Pulling specific factual claims out of a defined, closed set of documents, each tied to its source, is something agents do well today, provided the pool is genuinely bounded and the claims are the kind of discrete fact that has one correct answer. This is the extraction step underneath evidence anchoring, and it is mature enough that I trust it as a first pass on every build.

Deadline calculation from triggering facts. Once a triggering fact exists in structured form, an incident date, a notice received, calculating the resulting deadline is close to a solved problem. This is quietly one of the highest-value things agentic AI does in legal work, because missed deadlines from uncalculated dates are one of the most common and most expensive failures I have seen across 200 systems.

What agents cannot do reliably yet

Judgment calls with no clear right answer. Anything requiring weighing genuinely competing considerations, how aggressively to argue a contested point, whether a settlement offer is actually good given a client's full situation, is not something I trust an agent to resolve unsupervised, and I do not expect that to change soon. This is not a current limitation waiting on a better model. It is a category difference: judgment work is precisely the work that resists being reduced to a repeatable procedure, which is the MATTER Method's own definition of the line between assembly work and judgment work.

Working reliably outside a bounded context. Ask an agent to extract facts from a genuinely open-ended source, "review everything relevant in the client's files," with no defined pool, and reliability drops meaningfully. Agents are strong within a boundary and noticeably weaker at deciding where the boundary should be. That decision still belongs to a person.

Knowing what it does not know, without being told to check. Left unprompted, an agent will produce a fluent, confident answer to a question it cannot actually answer reliably, rather than surfacing its own uncertainty. This is exactly why the no-source-found check in evidence anchoring is a deliberate, separate step rather than something you can trust an agent to volunteer on its own.

Anything where the cost of an undetected error is high and the review step gets skipped. This is less a limitation of the technology and more a limitation of how firms deploy it. Agentic AI that runs first, with a lawyer supervising every step, is a meaningfully different risk profile from agentic AI that runs first and gets skimmed rather than reviewed because the review step is where the actual time savings were supposed to come from. Skipping supervision to capture speed is the most common way I see firms turn a genuine capability into a genuine liability.

The honest summary

Agents are reliably strong at bounded, repeatable, well-defined tasks, exactly the category the MATTER Method calls assembly work, and they remain unreliable at judgment, at deciding their own boundaries, and at flagging their own uncertainty without being told to. That is not a disappointing state of the technology. It is close to exactly where you would expect it to be if you took the assembly-versus-judgment distinction seriously from the start, which is the whole argument for naming that distinction before deploying anything.

The full ninety-day sequence for moving a practice onto this model, what to systematise first, which work AI should carry and with what guardrails, is in the AI-Native Practice Field Guide. It is not a hype document. It is the version of this piece with the implementation detail included.

ai agents · legal ai · ai in law · agentic ai

Want this working inside your practice?

Book a call