Back to Blog
Published:
Last Updated:
Fresh Content
Legal AI in the ERPChapter 7

When should contract AI escalate to a human reviewer?

7 min read
1,510 words
high priority
Ali Ahmed

Ali Ahmed

AI Solutions Engineer, Cognilium AI

TL;DR

On four triggers, not on a confidence score. A score in the middle is the least informative number in the system, and it fills queues with nothing useful.

On four triggers, and a confidence score is not one of them. A score in the middle is the least informative number the system produces, and building the escalation rule on it is why review queues fill with items nobody needed to see.

The useful question is not how sure is it. It is what kind of uncertainty is this.

Why is a confidence score a poor trigger on its own?

Because it collapses several different situations into one number.

A clause can score in the middle because it is genuinely borderline, because the playbook rule covering it is vague, because the clause is unusual, or because it touches two rules that disagree. Those need four different responses. A single threshold gives them the same one.

And the practical failure is predictable. Set the threshold high and the queue is unmanageable, so people stop working it. Set it low and the interesting cases are buried among routine ones, which produces the same result more slowly.

The deeper problem is that a middling score reads as mild concern, and mild concern gets accepted on a busy afternoon. A system should not produce a number whose most likely effect is to be waved through.

What should actually trigger escalation?

Four conditions. They are ours, and each maps to a different kind of uncertainty:

TriggerWhy it is different from a low score
No rule covers this clauseNot uncertainty — absence. The playbook has nothing to say, and only a person can decide whether that matters
Two rules disagreeThe contradiction is in your policy, not in the contract. Resolving it improves the playbook, not just this review
The clause crosses a red lineNot a judgement call at all. It should stop regardless of how confidently it was identified
The term reaches the ERPPayment terms, delivery terms, prices and discounts become operative in a purchase agreement. A wrong call here prices real orders

The fourth is the one most systems miss, because it is not a property of the clause — it is a property of where the clause ends up. A liability cap and a payment term may score identically; only one of them changes what a purchase order charges.

Escalate on consequence, not only on confidence.

What about the clause nobody wrote a rule for?

This is the trigger worth dwelling on, because it is the one people configure away first.

A contract arrives with a clause your playbook has never contemplated — a novel data-processing term, an unusual audit right, a jurisdiction nobody has dealt with. The system has no rule to measure it against. There are two honest options and one dishonest one.

The dishonest option is to score it against the nearest rule. It produces a number, the number looks like the other numbers, and nothing on the screen says the measurement was approximate. That is how a review becomes confidently wrong.

The two honest options are to escalate it as uncovered, or to state that it was not assessed. Either is defensible; both are visible. What matters is that "we had no rule for this" never renders the same way as "this was checked and it is fine."

And there is a compounding benefit. Every uncovered clause is a gap in the playbook that the business has now discovered for free. Escalations of this kind should be feeding rule changes, not just decisions.

How Paralegent AI handles it. Our contract-review app escalates on the trigger rather than on a threshold, and carries the clause and the rule with it. It is one of Cognilium's AI optimization apps for Microsoft Dynamics 365, and we build these against a customer's own playbook.

What should an escalation actually carry?

Enough for the reviewer to decide without re-reading the contract. A flag alone moves work rather than reducing it.

Four things, every time:

  • The clause, quoted. Not a summary of it.
  • The rule it was measured against, named. If the system cannot cite the rule, it should not raise the flag — that is the difference between a finding and an opinion.
  • Which trigger fired, so the reviewer knows what kind of question they are answering.
  • The suggested position, so the common case is accept rather than think from scratch.

A good escalation takes under a minute to resolve. If it takes ten, the system has handed over its problem rather than solved it.

What happens to everything the system does not escalate?

It is still visible, and that is not a detail.

Anything not escalated has been silently accepted, and silence is a decision the organisation is making at volume. Three rules make that defensible:

  • Nothing disappears. Every clause the system considered stays inspectable, with its score and its rule. A reviewer who wants to check the quiet ones can.
  • Nothing scores as "fine" by default. A clause no rule covers is unresolved, not acceptable. That distinction survives right into the audit eighteen months later.
  • The unresolved ones are counted. A review that resolved most of a contract and quietly skipped the rest should say so on its face.

Who reviews the reviewer?

Somebody, on a sample, forever. This is the part that gets dropped after the first month and it is the part that keeps the system honest.

Take a small random sample of non-escalated clauses each week and have a person read them cold. You are not looking for the system to be right — you are looking for the case it did not think was interesting. Escalations are self-checking, because a human already looked. The silent acceptances are not.

And when the sample finds something, fix the rule, not the review. A single corrected review helps one contract. A corrected playbook rule helps every contract after it — which is why the playbook is the real asset.

On Business Central, the escalation rule lives in the agent's own instructions — Microsoft's example tells an agent to "mark the task for manual review" when a field is missing or unclear: what you can and cannot build there.

About Cognilium Cognilium builds AI optimization apps for Microsoft Dynamics 365 — companion apps that optimize the pricing, inventory, warehouse and planning decisions your ERP manages but can't optimize. Dynamics is your system of record. Cognilium is your system of intelligence. https://cognilium.ai · https://www.linkedin.com/company/37180269/

Legal AI Ops. We transform legal workflows with agentic AI, copilots, agentic workflows and decision intelligence — built into core workflows rather than beside them, to raise productivity and cut operational overhead. Contract Review Copilot is the contract-review app in that family. It ships as Paralegent AI, in production today. How we build Legal AI Ops — custom AI capabilities on top of legal work, against your playbook and your Dynamics 365.

Bring a contract your team argued about to a 15-minute call and we will show you which of the four triggers it would have fired.

Sources

Sources and fact-check
#§ClaimTierSourceVerdict
11A single score collapses distinct kinds of uncertaintyT2 — ours, an argumentInternal definitionPASS
22The four escalation triggersT2 — ours, and §2 states they are oursInternal definition; Paralegant_TECHNICAL_PROFILE.md §4.3PASS
32Payment terms, delivery terms, prices and discounts become operative in a purchase agreementT1purchase-agreements, fetched 2026-09-10 — those terms are copied to the PO header and agreement prices are applied to covered linesPASS — load-bearing
43What an escalation must carryT2 — oursSame profile, §11 output schemaPASS
4b2aScoring an uncovered clause against the nearest rule produces a confidently wrong reviewT2 — ours, a design argument stated as oursInternal definitionPASS
4c2aUncovered clauses are playbook gaps and should feed rule changesT2 — oursInternal definitionPASS
54Unescalated clauses stay visible; "no rule" is unresolved not acceptableT2 — ours, a design positionInternal definitionPASS
65Sample non-escalated clauses; fix the rule not the reviewT2 — oursInternal definitionPASS

Tier summary: 1 × T1, 7 × T2 — 0 × T4.

No figures. No threshold value, no score, no accuracy claim, no volume. The article is about which conditions trigger escalation, never about how often they fire — which would be a performance claim about our own system.

Disclosure: no client, no measured outcome, no deployment claim.

Share this article

Ali Ahmed

Ali Ahmed

AI Solutions Engineer, Cognilium AI

Ali Ahmed is an AI Solutions Engineer at Cognilium AI.

Applied AI AgentsAgentic SystemsRetrieval-Augmented Generation (RAG)LLM Product Engineering
Next in this series
How do negotiated terms reach a purchase agreement in Dynamics 365?
Chapter 8 · 7 min
In short

Key takeaways

  • A confidence score is a poor escalation trigger because it collapses several different kinds of uncertainty into one number, and the middle of the range is the least informative part of it.
  • Escalate on four conditions: no rule covers the clause, two rules disagree, a red line is crossed, or the term becomes operative in the ERP.
  • Escalate on consequence, not only on confidence. A payment term and a liability cap can score the same; only one prices real purchase orders.
  • An escalation must carry the quoted clause, the named rule, the trigger and a suggested position — or it moves work instead of reducing it.
  • Silence is a decision. Unescalated clauses must stay visible, and "no rule covered this" must never render as "acceptable".
  • Sample the non-escalated clauses weekly. Escalations are self-checking; silent acceptances are not.
What goes wrong

Common mistakes to avoid

  • Tuning one confidence threshold and calling it the escalation policy. It produces either an unworkable queue or a buried one.
  • Raising a flag the system cannot justify. If it cannot name the rule, it has an opinion, not a finding.
  • Letting "no rule applies" render as approval. That is the single most expensive default in the system.
  • Dropping the weekly sample after the first month. The silent acceptances are the only part nobody else is checking.

Frequently Asked Questions

Find answers to common questions about the topics covered in this article.

Still have questions?

Get in touch with our team for personalized assistance.

Contact Us

Still have a question this did not answer?

The person who wrote this article answers these. Describe your setup and what you are stuck on — you will get a straight answer, including where we think the approach is wrong.