Back to Blog
Published:
Last Updated:
Fresh Content
Legal Contract ReviewChapter 10

What AI contract review will not decide for you

7 min read
1,499 words
high priority
Ali Ahmed

Ali Ahmed

AI Solutions Engineer, Cognilium AI

TL;DR

Five decisions an AI contract reviewer does not make, each bounded to what we actually built — including the one that matters most, which is that nothing in the system ever questions the playbook it was given.

What AI contract review will not decide for you

Five decisions, and a system like the one we built makes none of them.

That is not a complaint about the technology. A contract reviewer that quietly substituted its own judgement for a legal team's would be a much worse product — and much harder to defend when somebody asks why a clause was flagged. But knowing exactly where assessment stops and judgement begins is the difference between a tool a legal team trusts correctly and one they either over-rely on or abandon.

For the in-house counsel deciding whether to buy this, and the engineer deciding whether to build it. 8 minute read.

How to read this, and what it does do well

Every limit below is bounded to a system we built and can describe. Not to the category, not to every product on the market, and not to what a model could theoretically do. Where something was not researched, §6 says so rather than implying it does not exist.

Now the credit, because the rest means nothing without it.

It applies your rules rather than general legal knowledge. The playbook a legal team already maintains becomes the standard, so findings are arguable against a document the team owns — not against a model's impression of what contracts usually say.

It is consistent in a way people are not. The fortieth contract on a Friday is assessed by the same criteria as the first on a Monday. That is the single clearest advantage over human review at volume, and it is not a small one.

It reaches every clause. A reviewer under time pressure skims; the pipeline scores every chunk across every category regardless of where it sits in the document.

And it shows its work. Each finding carries a risk level, a justification and a suggested revision, attached to the specific text that produced it — so disagreeing with it is possible, which is the property that makes it usable at all.

None of that is trivial and all of it is hard to build. The limits below are the boundary of the job, not failures at it.

It will not tell you the playbook is wrong

This is the one that matters most, and it is the direct consequence of the design's greatest strength.

The system's authority comes from applying the customer's own rules. Which means it applies them faithfully — including the ones that are wrong.

A playbook rule that was correct three years ago and has since been overtaken by legislation will be applied to every contract, confidently, with a justification and a suggested revision, until somebody notices.

A rule written for a supplier relationship and applied to a customer one produces findings that are internally consistent and commercially backwards. Nothing in the pipeline evaluates whether a rule is a good rule.

Every stage assumes the playbook is correct. Extraction pulls out the terms. Scoring measures relevance to those terms. Analysts assess clauses against those rules. At no point does anything ask whether the standard being applied still makes sense — and the more capable the system gets at applying rules, the more efficiently it propagates a bad one.

That is not a defect to be fixed with a better model. It is structural. The system is a very fast, very consistent applier of a standard it did not write and cannot evaluate, and the standard is the part that decides whether the output is any good.

Which relocates the real work. Buying this does not remove the need for someone to own the playbook — it makes that ownership considerably more consequential, because a rule that used to be applied inconsistently by tired humans is now applied identically to everything.

It will not decide whether to sign

The system finds clauses, assesses them against rules and proposes revisions. Ranking those findings against a commercial relationship is a different activity entirely.

A high-risk indemnity in a contract with your largest customer is not the same decision as the identical clause with a new supplier. The risk is the same; the answer is not. What you concede depends on what you want, what you are worth to them, what they are worth to you, and what you conceded last time.

None of that is in the document, so none of it is available to a system that reads documents. The findings are an input to the decision. They are not the decision, and a product that presented them as one would be exceeding its evidence.

It will not remember the counterparty

Bounded to what we built: the pipeline processes a contract against a playbook. It does not carry history across reviews.

So it does not know that this counterparty always redlines clause eleven, that their standard paper has softened since last year, or that the last three negotiations with them settled in the same place. A reviewer with five years of experience knows all of that and it materially changes what they focus on.

Note this is a scope statement, not a technical impossibility. Nothing prevents a system from maintaining counterparty history — it is a different product with different data retention questions attached, and it is not what this one does. Anyone comparing tools should ask about it explicitly rather than assuming either answer.

It will not survive a bad taxonomy

The routing that makes the system affordable depends on categories being distinct.

Our published piece on this pipeline's routing documents the failure directly: when categories overlap, chunks score highly on several at once, the router sends work to all of them, and the saving degrades back toward analysing everything with everything.

The system reports no error while this happens. The output is still correct — arguably more complete. It just costs what it would have cost with no routing at all.

And the fix is not in the software. It is in how somebody defined the categories in a playbook. That makes it a configuration problem presenting as a cost problem, which is the hardest kind to attribute.

Two things we have not researched, stated as gaps

Comparison against other contract-review products. We have opened no competitor's primary documentation, run no benchmark and hold no evaluation data. We are not saying this system compares favourably or unfavourably — we are saying we have not done that work, and anything we wrote about it would be marketing rather than evidence.

Model selection for legal-domain reasoning. Which model performs best on clause-level legal assessment is an empirical question we have not tested at the level a published claim would require.

A declared research gap is not a finding, and presenting one as an absence is how technical writing starts going past its sources.

Why this is not a criticism, and the check for Monday

The pattern across all five limits is the same, and it is a design property rather than a defect.

The system holds assessments. It does not hold judgements. It will tell you a liability cap is unusually low against your own standard. It will not tell you whether to accept it, whether this counterparty is worth it, or whether your standard was sensible in the first place.

That is the correct division. A reviewer that inferred commercial strategy from clause text would be unauditable, and the first time it was wrong nobody could explain why.

How we would build around it. Make the playbook a first-class, versioned, reviewable artefact rather than an uploaded file — so a change to the standard is as visible as a change to the code applying it. The system will never question the rules; the least it can do is make them easy to question.

Status, stated plainly. This is a system we built. It runs, it is demonstrable on a call, and it has zero delivered engagements — no legal team is running it today, and nothing in this cluster describes anyone's business but our own. We build these on request, against your documents and your playbook.

The check worth running, and it needs nothing from us. Take one rule from your playbook and answer three questions: when was it written, who owns it now, and what would tell you it had gone stale? If the answer to the third is "somebody would notice", that is the assumption automation removes — because consistency means a wrong rule is now applied perfectly, every time, to everything.

About Cognilium Cognilium is an AI engineering company — agent systems, retrieval, knowledge graphs, voice AI and the production plumbing that makes them survive contact with real workloads. https://cognilium.ai · https://www.linkedin.com/company/37180269/

Want to know where the boundary sits for your documents? Book a 15-minute call — we will walk what the system decides and what it hands back, on your playbook if you bring one. No deck.

Sources

This article draws on our own build and on the published sibling covering this pipeline's routing, cited in §5.

Share this article

The work behind this series

The review pipeline these articles describe — extraction, category routing, parallel scoring and human escalation — as an engagement.

Ali Ahmed

Ali Ahmed

AI Solutions Engineer, Cognilium AI

Ali Ahmed is an AI Solutions Engineer at Cognilium AI.

Applied AI AgentsAgentic SystemsRetrieval-Augmented Generation (RAG)LLM Product Engineering
In short

Key takeaways

  • A playbook-driven reviewer applies the customer's own rules, which is what makes its findings arguable — and it applies the wrong ones just as faithfully.
  • Nothing in the pipeline evaluates whether a rule is a good rule. Every stage assumes the playbook is correct.
  • The findings are an input to a commercial decision, not the decision — the same clause warrants different answers with different counterparties.
  • As built, the system does not carry history between reviews, so it has no memory of how a counterparty behaves.
  • Automation makes playbook ownership more consequential, not less, because a bad rule is now applied identically to everything.
What goes wrong

Common mistakes to avoid

  • Treating consistency as correctness. Applying the same standard every time is only valuable if the standard is right.
  • Reading findings as a recommendation to sign or refuse. They are assessments against rules, not commercial judgement.
  • Assuming counterparty history is included. Ask explicitly; it is a scope question, not a given.
  • Blaming the model when routing costs rise. Overlapping categories are a taxonomy problem that surfaces as a bill.

Still have a question this did not answer?

The person who wrote this article answers these. Describe your setup and what you are stuck on — you will get a straight answer, including where we think the approach is wrong.