What a contract review system is actually made of
Engineering legal contract review means building the path a document takes from an inbox to a reviewer's screen, and being explicit about every point on it where a judgement gets made. The interesting work is not the model call. It is deciding what happens to a 60-page master services agreement in the ninety seconds before anyone looks at it.
A working pipeline has four stages, and they are worth separating because they fail in different ways. Text comes out of the document. Clauses are sorted into categories. Each category is scored against a position the business has already taken. Then a rule — not a model — decides which findings a human must see. Collapse any two of those stages into one prompt and you lose the ability to say why a clause was flagged.
Everything below is one of those stages. Each links to the writeup that covers it in full.
Stage one: getting reliable text out of the document
Contracts arrive as Word files with tracked changes, as scans of signed originals, and as PDFs exported from something that lost the heading structure on the way out. Before anything can be reasoned about, the clause boundaries have to be recovered — and clause boundaries are structural, not semantic.
This is the stage where reaching for a language model first is the expensive mistake. Numbered headings, defined-term tables and cross-references are deterministic: a parser finds them the same way every time, for no per-page cost, and can be tested against a fixture. A model asked to do the same job will be right most of the time and silently wrong on the document that matters, because a missed clause boundary does not produce an error — it produces a review that never mentions the clause.
The rule that holds up in production: parse what has structure, and spend model calls only on the text where meaning is genuinely ambiguous.
Stage two: routing clauses to the right category
A limitation-of-liability clause and a data-processing clause are not reviewed by the same standard, and they are not reviewed by the same person. Routing is what makes the rest of the pipeline affordable: instead of asking one general question of the whole agreement, the system asks a specific question of each part.
Categories come from the playbook, which means they come from the business rather than from a taxonomy someone downloaded. Indemnity, liability cap, termination, assignment, governing law, data protection, payment terms, IP ownership — the list is short, stable, and the same across most agreements a company signs. What differs is the position the company holds on each one.
Routing also produces the audit trail. When a reviewer asks why a clause was not flagged, the answer is either "it was routed to payment terms and scored inside policy" or "it was never routed", and those are very different bugs.
Stage three: scoring each category in parallel
Once clauses are routed, each category is scored independently against the playbook position for that category. Independently matters for two reasons: the scorers run at the same time, so a long agreement does not take proportionally longer, and a scorer that is wrong about indemnity cannot contaminate the reading of governing law.
The awkward part of parallel scoring is duplication. Contracts repeat themselves — a liability concept appears in the main body, again in an annex, and again in a definitions table — and independent scorers will each report it. A reviewer handed the same finding eleven times stops reading the list, which defeats the point of the system. Deduplication is therefore part of the scoring stage, not a cosmetic pass afterwards.
Cost lands here too. Scoring is where the token spend is, and it is the stage where the difference between a system that is expensive and one that is not comes down to how much text each scorer is given rather than which model is chosen.
Stage four: deciding what a human has to read
Escalation is the stage that determines whether the system is trusted, and it should be the least clever part of the build. A finding reaches a person because it crossed a threshold the business wrote down — a liability cap above an agreed figure, an assignment clause without consent, any clause in a category marked always-review. Those are rules. They are readable, they are arguable in a meeting, and they do not change when a model is swapped.
The failure mode is a system that escalates everything, which is a search interface wearing a review system's clothes, or one that escalates nothing, which is an unreviewed contract with a green tick on it. Both come from letting a model decide significance instead of letting it decide content.
Getting this boundary right is also what makes the output defensible. The system reports what a clause says and how it compares to the stated position. Whether that difference is acceptable is a commercial decision, and it belongs to a person.
Where the reviewer actually works
Lawyers redline in Word. A review system that produces findings anywhere else has added a second place to look, and the second place loses. The practical consequence is that output has to land as tracked changes and comments in the document itself, with the reasoning attached to the clause it concerns rather than collected in a report nobody opens.
That constraint shapes the build more than it sounds. It decides the timing — pre-signature review has to be fast enough to happen during drafting — and it decides what a finding is allowed to be, because a comment anchored to a clause cannot be three paragraphs of hedging.
Contract review that has to agree with the ERP
A reviewed contract is only useful if the terms it settled are the terms the business then operates on. The renewal date, the agreed price, the payment terms and the liability position all exist a second time inside Dynamics 365 — on the vendor record, the purchase agreement, the sales order — and the two copies drift. A contract says one price and a purchase order carries another; a renewal passes because the date lived in a document and the reminder lived nowhere.
This is the join that makes contract review an ERP problem rather than a document problem, and it is the reason the pipeline above has to write back as well as read. It is also where the decision between dedicated contract-lifecycle software and the ERP you already run gets made on something other than a feature list.
Build it, or buy it
The honest version of this question is not build-versus-buy but which parts are worth building. Extraction and storage are commodities and should be bought. The playbook — the set of positions a company holds and the thresholds that escalate — cannot be bought, because it is the thing that makes the review yours rather than generic. Most of the value sits in that layer, and most of the effort in a failed project went into the layers below it.
The pipeline described on this page is what we build as an engagement: extraction, category routing, parallel scoring and the escalation rules, wired into the systems that already hold the commercial record.