TL;DR
Redlining is proposing the edit, not flagging the risk. That distinction decides whether a system saves a reviewer time or simply moves the work elsewhere.
Redlining is proposing the edit — not flagging the risk. It is the tracked change in the margin that says this clause should say thirty days, not sixty, and here is the wording.
Yes, AI can do it. Most contract AI does not — it flags and stops, which is a different and much easier job.
What does redlining actually mean?
The name comes from the pen. A reviewer marks the document in red: strike this, insert that, and a comment explaining why.
Three things happen in a real redline, and a system has to do all three or it has not redlined anything:
| Part | What it is |
|---|---|
| The deletion | The words that must go, marked exactly |
| The insertion | The words that replace them — a concrete alternative, not a description of one |
| The reason | Why, in terms the counterparty can negotiate with |
The insertion is the part that separates redlining from review. "This liability cap is below our standard" is a finding. "Replace 'not exceed the fees paid' with 'not exceed 150% of fees paid in the preceding twelve months'" is a redline.
One of those saves a lawyer time. The other adds a step.
Why is flagging so much easier than redlining?
Because flagging only needs to recognise a problem. Redlining has to commit to an answer.
A system that flags can be usefully vague. A system that redlines cannot — it has to produce specific words that are legally coherent, consistent with your other positions, and defensible if the counterparty pushes back.
That is why most products stop at flagging and call it review. The demo looks similar. The work left on the reviewer's desk is completely different:
- A flag hands back a question. The reviewer still has to decide the position, draft the wording, and check it against the rest of the agreement.
- A redline hands back a decision. The reviewer accepts, edits or rejects — which is a minute, not fifteen.
Ask any vendor for the insertion text, not the finding. It is the fastest way to tell the two apart, and it is a question that cannot be answered with a slide.
What does an AI system need in order to redline?
Four inputs. Missing any one turns the redline back into a flag:
- Your standard position, as words. Not a principle — the actual clause you want. A rule that says "limit liability where commercially reasonable" cannot generate an insertion, because there is nothing to insert.
- Your fallback position. Redlining is negotiation, and a system that only knows your ideal ask will propose it every time, including where you would have settled.
- Your red line. So the system knows when to stop proposing and escalate instead.
- The rest of the contract. An insertion that contradicts another clause is worse than no insertion.
This is why the playbook decides whether redlining is possible at all — and why playbook-driven contract review is the prerequisite rather than a feature. A playbook of principles produces flags. A playbook of positions produces redlines.
Where does AI redlining go wrong?
Four failure modes, and none of them is bad language:
- Proposing the ideal every time. Technically correct, commercially exhausting. The counterparty learns that your first position is theatre.
- Internally inconsistent insertions. A liability change that contradicts the indemnity clause two pages later. The model saw the clause; it did not see the contract.
- Redlining the unredlinable. Some terms are not negotiable positions — governing law on a counterparty's standard paper, for instance. Proposing an edit there wastes a negotiation round.
- Silent confidence. A proposed insertion that looks identical whether the system was certain or guessing. A redline the reviewer cannot calibrate is a redline they must fully re-check, which removes the whole benefit.
The fix for the last one is not a score. It is saying which trigger fired — no rule covered this, two rules disagreed, this crosses a red line — because those tell a reviewer what kind of check they owe.
Where should redlining actually happen?
In the document, as tracked changes. Not in a portal, not in a report, not as a list of recommendations to apply by hand.
The reason is mechanical rather than aesthetic. A redline that is not in the document has to be transcribed into it, and transcription is where positions get softened, insertions get paraphrased and one of the four proposed edits gets forgotten.
A redline delivered as tracked changes is already the deliverable. The reviewer works in Word, where the contract already is, and the accept/reject decision is the same gesture they already use.
This is how we build it. Paralegent AI proposes against the clause inside Word — one of Cognilium's AI optimization apps for Microsoft Dynamics 365, and we build these on request against a customer's own playbook, contracts and Dynamics environment.
About Cognilium Cognilium builds AI optimization apps for Microsoft Dynamics 365 — companion apps that optimize the pricing, inventory, warehouse and planning decisions your ERP manages but can't optimize. Dynamics is your system of record. Cognilium is your system of intelligence. https://cognilium.ai · https://www.linkedin.com/company/37180269/
Legal AI Ops. We transform legal workflows with agentic AI, copilots, agentic workflows and decision intelligence — built into core workflows rather than beside them, to raise productivity and cut operational overhead. Contract Review Copilot is the contract-review app in that family. It ships as Paralegent AI, in production today. How we build Legal AI Ops — custom AI capabilities on top of legal work, against your playbook and your Dynamics 365.
Bring a contract you have already redlined by hand to a 15-minute call, and we will show you which of your edits a playbook could have proposed and which ones needed you.
Sources
No external source is cited. Every claim is either our own design position, labelled as such, or a definition of standard legal-industry practice.
Sources and fact-check
| # | § | Claim | Tier | Source | Verdict |
|---|---|---|---|---|---|
| 1 | 1 | Redlining comprises deletion, insertion and reason | T2 — ours, a definition of common practice, stated as ours | Internal definition | PASS |
| 2 | 2 | Flagging names a risk; redlining commits to replacement wording | T2 — ours, the article's central argument | Internal definition | PASS |
| 3 | 2 | Most contract AI stops at flagging | T2 — ours, and bounded. Stated as our reading of the market, not as a claim about any named product. No vendor is named | _research/2026-09-09/01-RESEARCH-A… §2 SERP reading | PASS |
| 4 | 2 | "a minute, not fifteen" | T2 — ours, illustrative and clearly rhetorical. Not a measured figure, not attached to any customer | — | PASS |
| 5 | 3 | Four inputs are required to redline | T2 — ours | Same profile §1, §11 | PASS |
| 6 | 3 | A playbook of principles yields flags; positions yield redlines | T2 — ours | Internal definition | PASS |
| 7 | 4 | The four failure modes | T2 — ours | Internal definition | PASS |
| 8 | 5 | Tracked changes in the document; transcription loses precision | T2 — ours | Same profile §6 Word add-in | PASS |
| 9 | 5 | Paralegent AI proposes against the clause inside Word | T2 — capability. No outcome, no customer, no measure | Same profile §1 Delivery Method, §6 | PASS |
Tier summary: 0 × T1, 9 × T2 — 0 × T4.
Entirely T2, and that is correct here. This article defines industry practice and states design positions. It makes no vendor claim and names no competitor — row 3 is bounded to our own reading and deliberately names nobody.
One rhetorical figure, flagged. "a minute, not fifteen" in §2 is illustrative contrast, not a measurement. It is attached to no customer and no benchmark. If `/fact-check` wants it cut, it can go without damaging the argument.
No figures from the technical profile. No call counts, no durations, no accuracy.
