← All guides

9 min read · Updated August 2026

Turning research you've already done into something that compounds

Most practices redo the same research every eighteen months because nobody can find the last time they did it. Here's how I fix that without a big software project.

Key takeaways

  • The research you've already paid for is usually your practice's most underused asset, sitting in memos nobody can locate.
  • A knowledge base only works if retrieval is easier than redoing the research from scratch, that's the entire bar it has to clear.
  • Structure beats volume: a hundred well-tagged memos beat a thousand loose ones in a shared drive.
  • Build the retrieval layer before you worry about anything fancier, because that's where almost all the value sits.
  • Someone has to own keeping it current, or it becomes an expensive historical archive within a year.

The research you're already sitting on

I did an exercise recently with a small regulatory practice where we pulled every substantive research memo from the past three years into one place, just to see what was there. It came to just over two hundred documents. When we actually looked through them, at least thirty were near-duplicates of each other, the same or nearly the same legal question researched from scratch multiple times because nobody remembered, or could find, that it had already been answered.

That's not a knowledge management failure in some abstract sense, it's billable hours spent redoing work that already existed, invisible because it happened in different matters, months apart, handled by different people who had no way to know the other research existed. This is the single most common inefficiency I find when I go looking, more common than bad intake processes or slow document review.

What I've watched fail first

  • A shared drive folder named 'research,' sorted by date, with no tagging. This is where research goes to be technically preserved and practically unfindable.
  • A big-bang project to migrate everything into a fancy new system before anyone has proven the basic retrieval concept works on a small slice.
  • Requiring detailed metadata entry for every document at the point of filing, which nobody keeps up with under deadline pressure, so the system degrades within a month.
  • Building the system around whoever's personal filing habits are best, rather than around how the next person searching for something will actually think to look for it.

How I actually build this with a client

  1. 1. Start with the questions people actually ask, not the documents you have

    Before touching the archive, I ask the team to list the ten questions they've each had to research more than once. That list becomes the organizing structure. It's a completely different exercise than starting with 'here are our documents, how do we file them,' and it produces a system organized around retrieval instead of storage.

  2. 2. Tag by the question, not just the practice area

    'Employment law' as a tag is nearly useless when you're trying to find one specific memo. 'Whether a remote employee in state X triggers registration requirements' is what someone actually searches for. I push clients toward specific, question-shaped tags even though it takes longer up front.

  3. 3. Build a thin retrieval layer before anything else

    This is often just a well-structured, searchable index with good tags and a one-line summary of the conclusion for each memo, sometimes backed by a simple AI-assisted search over the collection. I resist building anything more elaborate until this basic layer is proven to actually get used.

  4. 4. Pressure-test retrieval with real questions

    Once it exists, I hand the system to someone who wasn't involved in building it and give them three real questions the practice has faced before, timing how long it takes them to find the relevant prior research. If it's not meaningfully faster than starting from scratch, the structure needs work before it's worth rolling out.

  5. 5. Assign an owner for upkeep, explicitly

    Someone has to be responsible for tagging new research as it's produced, and for periodically checking that older conclusions still hold given changes in law. Without a named owner, this decays within a year, and everyone quietly reverts to redoing research from scratch.

Where AI genuinely helps here, and where it doesn't

AI is useful in this specific project for two things: helping surface and tag the backlog of old, unstructured memos so a human doesn't have to read all two hundred of them by hand, and powering a natural-language search layer over the collection so someone can ask a question in plain terms rather than guessing the exact tag. Both of those are strong, practical uses.

Where I'd stop is letting AI summarize legal conclusions without a lawyer reviewing the summary against the source memo before it goes into the knowledge base. The whole point of this system is that people trust the conclusion enough to rely on it without redoing the research, and a subtly wrong AI-generated summary sitting in the archive is worse than no archive at all, because it's wrong with the appearance of authority.

The real success metric

I don't measure this project by how many documents got tagged. I measure it by whether, three months later, someone can point to a specific instance where they found an existing answer instead of billing time to redo the research. One good example of that is worth more than a fully populated but unused system.

Keeping it from going stale

The knowledge base's biggest long-term risk isn't that it goes unused, it's that it gets used with outdated conclusions nobody flagged as superseded. I build in a simple rule: any memo older than eighteen months in a fast-moving area gets a visible flag prompting a quick recheck before anyone relies on it, rather than trusting people to remember the law might have changed underneath a conclusion they're pulling off the shelf.

This is a small addition, but it's the difference between a knowledge base that compounds in value over years and one that quietly becomes a liability the first time someone relies on a conclusion that's no longer correct.

Questions

How long does it take to build a basic version of this for an existing archive?
For a practice with a few hundred documents, I've done a working first version in two to three weeks of part-time effort, mostly spent on tagging the backlog. The retrieval layer itself is usually the fast part once the tagging structure is settled.
Do I need a dedicated software platform for this, or can it start simpler?
Start simpler than you think. A well-tagged, well-indexed set of documents with decent search, even built on tools the practice already has, beats a partially adopted, expensive platform. Prove the retrieval concept before spending on infrastructure.
What's the biggest mistake practices make when starting this project?
Trying to organize by document type or practice area instead of by the actual questions people ask. It feels intuitive but it recreates the same 'technically filed, practically unfindable' problem the archive already has.
Should paralegals and associates be able to add to the knowledge base directly?
Yes, encourage it, but route additions through a light review step so tagging stays consistent and conclusions get a second look before they're trusted by the whole team.
How do I know if this is even worth the investment for a small practice?
Do the quick version of the audit I described: pull research from the last year or two and count near-duplicate questions. If you find more than a handful, the redone work alone usually justifies the project.

Want this built for your practice, not just read about it?

Book an intro call