I spent more of my fourteen years running a firm arguing about document filing than I ever did about software, and at the time it felt like housekeeping. It is not housekeeping now. How your documents are named, foldered and dated has become the thing that decides what an AI tool can answer for you, and what each answer costs you to get.

A benchmark published on 18 August 2026 puts numbers against that. NetDocuments ran the same AI agent over the same questions and the same documents twice, once with structured information about the matter attached and once without it. Cost for each correct answer fell by 48 per cent and token use for each answer by 52 per cent. Where the saving went back into quality rather than into the bill, answer quality rose by 7 per cent while overall cost still fell by 18 per cent. The test put 300 questions across ten real matters holding 874 documents and around 60 million characters, taken from public regulatory filings and court dockets.

Read the study for what it is

The supplier ran the benchmark on its own product and nobody independent has repeated it, so treat the headline as an advertisement until somebody does. The money figure deserves less than that. It models a firm of two thousand people asking four million questions a year and arrives at savings near a million dollars, which tells a practice of eight fee earners nothing at all. What survives the discount is the comparison underneath it. The model stayed the same, the questions stayed the same and the documents stayed the same. The only thing that moved was the information wrapped around them, and it moved the result on both quality and price.

What it means for a firm your size

You are not going to buy a context graph. You already hold the raw material one is built from, being matter numbers, client names, document types, dates, who signed what and which version went out of the door. The question is whether any of it is readable by a machine, or whether it sits in the heads of two long-serving staff and in folders called new folder two and final v3.

Point an AI tool at a shared drive of scans named by the scanner and it has to read everything and guess at what each thing is. Point it at the same matter with names that say what a document is and when it was signed, and it reads less, guesses less and gives you an answer you can check against the file. The benchmark measures that gap in tokens and dollars. Your fee earners already measure it in the twenty minutes they spend hunting for the executed version.

Where to start

Take one practice area and open the last twenty closed matters. Ask whether you can tell, from the file names and the folder structure alone, which document is the executed agreement, which is the client instruction and which is counsel advice. If you cannot, that is your AI project, and it sits ahead of any tool you were thinking of buying.

Then do three things. Agree a naming convention and apply it to new matters rather than trying to repair the archive. Run character recognition over your scans, because a picture of a page holds no text for a tool to read. Ask your document system supplier, in writing, what matter information it passes to any AI you connect and what it keeps back.

NetDocuments set out the benchmark, the method and every figure in its own open announcement of the Legal Context Engineering Benchmark, which asks nothing of you to read.

If you want your filing and your matter data put straight before you buy an AI tool rather than afterwards, that is work we do with firms: talk it through with us.