Topical authority hub

Engineering Legal Contract Review

What happens between a contract arriving and a lawyer seeing it: extraction, category routing, parallel scoring, and the escalation rules that decide which clauses a person must read.

AI contract review
Articles
12
Total read
79m
Pillar
Set
Start here — foundational guide
PillarFoundational guide

How do you review a contract against your own legal playbook with AI?

A contract review system that reads your own playbook is four services and three queues. This is the whole pipeline — what each stage costs, where it degrades, and the one design decision the architecture is built around.

Ali Ahmed8 minAug 24, 2026
Read the guide

What a contract review system is actually made of

Engineering legal contract review means building the path a document takes from an inbox to a reviewer's screen, and being explicit about every point on it where a judgement gets made. The interesting work is not the model call. It is deciding what happens to a 60-page master services agreement in the ninety seconds before anyone looks at it.

A working pipeline has four stages, and they are worth separating because they fail in different ways. Text comes out of the document. Clauses are sorted into categories. Each category is scored against a position the business has already taken. Then a rule — not a model — decides which findings a human must see. Collapse any two of those stages into one prompt and you lose the ability to say why a clause was flagged.

Everything below is one of those stages. Each links to the writeup that covers it in full.

Stage one: getting reliable text out of the document

Contracts arrive as Word files with tracked changes, as scans of signed originals, and as PDFs exported from something that lost the heading structure on the way out. Before anything can be reasoned about, the clause boundaries have to be recovered — and clause boundaries are structural, not semantic.

This is the stage where reaching for a language model first is the expensive mistake. Numbered headings, defined-term tables and cross-references are deterministic: a parser finds them the same way every time, for no per-page cost, and can be tested against a fixture. A model asked to do the same job will be right most of the time and silently wrong on the document that matters, because a missed clause boundary does not produce an error — it produces a review that never mentions the clause.

The rule that holds up in production: parse what has structure, and spend model calls only on the text where meaning is genuinely ambiguous.

Stage two: routing clauses to the right category

A limitation-of-liability clause and a data-processing clause are not reviewed by the same standard, and they are not reviewed by the same person. Routing is what makes the rest of the pipeline affordable: instead of asking one general question of the whole agreement, the system asks a specific question of each part.

Categories come from the playbook, which means they come from the business rather than from a taxonomy someone downloaded. Indemnity, liability cap, termination, assignment, governing law, data protection, payment terms, IP ownership — the list is short, stable, and the same across most agreements a company signs. What differs is the position the company holds on each one.

Routing also produces the audit trail. When a reviewer asks why a clause was not flagged, the answer is either "it was routed to payment terms and scored inside policy" or "it was never routed", and those are very different bugs.

Stage three: scoring each category in parallel

Once clauses are routed, each category is scored independently against the playbook position for that category. Independently matters for two reasons: the scorers run at the same time, so a long agreement does not take proportionally longer, and a scorer that is wrong about indemnity cannot contaminate the reading of governing law.

The awkward part of parallel scoring is duplication. Contracts repeat themselves — a liability concept appears in the main body, again in an annex, and again in a definitions table — and independent scorers will each report it. A reviewer handed the same finding eleven times stops reading the list, which defeats the point of the system. Deduplication is therefore part of the scoring stage, not a cosmetic pass afterwards.

Cost lands here too. Scoring is where the token spend is, and it is the stage where the difference between a system that is expensive and one that is not comes down to how much text each scorer is given rather than which model is chosen.

Stage four: deciding what a human has to read

Escalation is the stage that determines whether the system is trusted, and it should be the least clever part of the build. A finding reaches a person because it crossed a threshold the business wrote down — a liability cap above an agreed figure, an assignment clause without consent, any clause in a category marked always-review. Those are rules. They are readable, they are arguable in a meeting, and they do not change when a model is swapped.

The failure mode is a system that escalates everything, which is a search interface wearing a review system's clothes, or one that escalates nothing, which is an unreviewed contract with a green tick on it. Both come from letting a model decide significance instead of letting it decide content.

Getting this boundary right is also what makes the output defensible. The system reports what a clause says and how it compares to the stated position. Whether that difference is acceptable is a commercial decision, and it belongs to a person.

Where the reviewer actually works

Lawyers redline in Word. A review system that produces findings anywhere else has added a second place to look, and the second place loses. The practical consequence is that output has to land as tracked changes and comments in the document itself, with the reasoning attached to the clause it concerns rather than collected in a report nobody opens.

That constraint shapes the build more than it sounds. It decides the timing — pre-signature review has to be fast enough to happen during drafting — and it decides what a finding is allowed to be, because a comment anchored to a clause cannot be three paragraphs of hedging.

Contract review that has to agree with the ERP

A reviewed contract is only useful if the terms it settled are the terms the business then operates on. The renewal date, the agreed price, the payment terms and the liability position all exist a second time inside Dynamics 365 — on the vendor record, the purchase agreement, the sales order — and the two copies drift. A contract says one price and a purchase order carries another; a renewal passes because the date lived in a document and the reminder lived nowhere.

This is the join that makes contract review an ERP problem rather than a document problem, and it is the reason the pipeline above has to write back as well as read. It is also where the decision between dedicated contract-lifecycle software and the ERP you already run gets made on something other than a feature list.

Build it, or buy it

The honest version of this question is not build-versus-buy but which parts are worth building. Extraction and storage are commodities and should be bought. The playbook — the set of positions a company holds and the thresholds that escalate — cannot be bought, because it is the thing that makes the review yours rather than generic. Most of the value sits in that layer, and most of the effort in a failed project went into the layers below it.

The pipeline described on this page is what we build as an engagement: extraction, category routing, parallel scoring and the escalation rules, wired into the systems that already hold the commercial record.

Continue the path

Ordered by chapter. Each post stands alone but builds on the one before it.

Chapter 1

Twelve scorers, eleven analysts — what happens to the category with no specialist?

A contract-review pipeline scores every chunk across twelve legal categories and routes it to eleven specialists. The missing twelfth is not an oversight — it is the control signal the whole cost model depends on.

8 minRead
Chapter 2

How do you stop eleven agents reporting the same clause eleven times?

Parallel specialist agents all read the same text, so they all report it. Deduplicating across agents is not tidying — it is what decides whether the output is usable, and it is a different problem from deduplicating messages.

6 minRead
Smart Category Routing for Contract Review — Cognilium AI
Chapter 2

Smart Category Routing for Contract Review

A focused application of the LLMOps routing pattern to legal contract analysis — the analyst-selection logic that ships fewer clauses to fewer agents and finishes a 3,300-call review in 154 seconds.

6 minRead
Chapter 3

Where should retry live when your worker reads from a queue?

A queue worker can retry in-process or let the queue redeliver. Doing both is correct — and the two tiers multiply rather than add, which is the cost nobody budgets for.

6 minRead
Chapter 4

What stops a degraded model from taking the whole queue with it?

Retry is how a degraded model provider becomes an outage. A circuit breaker fails fast instead — and in a container fleet it protects far less than you think, because every replica has its own.

6 minRead
Chapter 5

Is partial failure an outage, or a setting?

When three of twelve agents fail, does the job fail? Making that a configured threshold rather than an exception is the right call at scale — and it produces results that carry no record of how complete they are.

6 minRead
Chapter 6

Deterministic parsing or a model — when is the cheapest call the one you never make?

The same pipeline ingests a spreadsheet with no model calls and a PDF with dozens. The format your customer sends decides your unit economics — which makes the upload form a cost lever most teams never touch.

7 minRead
Chapter 7

Your container is mid-message and autoscaling just killed it

Amazon ECS documents task scale-in protection and names queue-processing workloads in its first example. It works — and it will block your next deployment in ways AWS spells out and most teams discover during a release.

7 minRead
Chapter 8

Can you change an agent's prompt without a redeploy?

Hot-reloading agent prompts and thresholds from a parameter store is operationally excellent and a governance problem — a behaviour change with no diff, no review and no rollback story.

6 minRead
Chapter 9

Why does the contract reviewer live inside Word?

Putting AI review inside Word rather than in a portal is a distribution decision, and it cascades — into how you authenticate, what your output can be, and who owns the surface you ship on.

6 minRead
Chapter 10

What AI contract review will not decide for you

Five decisions an AI contract reviewer does not make, each bounded to what we actually built — including the one that matters most, which is that nothing in the system ever questions the playbook it was given.

7 minRead
Build it for real

Read the writeup. Now ship the system.

Cognilium engineers ship the architectures behind these articles for enterprise teams. If you're mid-build on legal contract review, talk to us.

AI for your ERP

Dynamics 365 & SAP

Schedule a call