Every finding traces to a highlighted span on a source page, judged against a versioned industry baseline pinned before the work started. Nothing is asserted that cannot be shown.
A team reads thousands of documents, works through a checklist inherited from a partner's spreadsheet, and writes a memo under deadline. Four structural failures follow — and the real cost is not the fee. It is the missed finding discovered after close.
Typical transaction diligence. Vendor assessments queue for months behind it.
Per transaction in fees. Enterprises run thousands of vendor reviews a year at $2k–15k each.
The same data room, materially different findings. No reproducibility, no measurable quality bar.
“What good looks like here” walks out at retirement, and does not scale to a new sector.
Anything a parser, a lookup, a calculation or a rule can do reliably is not given to a language model. The model's job is judgement, synthesis and narrative — the part that genuinely requires reasoning. Facts come from components that can be tested to 100%; only judgement is probabilistic.
Step 05 is the only step a model touches — and it reads structured facts, never raw documents.
What is true about this target, and what should be true in this industry, are different kinds of statement. They are stored apart, retrieved apart, and presented to the model in separately-labelled blocks.
The target's own documents and the facts drawn from them, scoped to one engagement. It never defines a standard.
Practices, obligations, benchmarks and known failure patterns — versioned and cited. It never asserts anything about this target.
The system encodes the weight of a source rather than leaving it to the reader. A finding resting only on tier 4–5 material cannot be rated material without an explicit reviewer override — and the report says so.
| Tier | Source | Weight |
|---|---|---|
| 1 | Audited financials, regulatory filings, court and registry records | Highest |
| 2 | Executed contracts, certificates, licences, insurance policies | High |
| 3 | Management accounts, board minutes, internal policies | Medium |
| 4 | Management presentations, forecasts, self-assessment questionnaires | Low — corroboration required |
| 5 | Verbal statements, undated drafts, unsigned documents | Lowest — flagged unverified |
Many, shallow, recurring. Thousands a year — automation ratio matters more than depth.
Few, deep, deadline-driven. 500–5,000 documents against the clock.
Find the issues before the buyer does.
A screening tier, then a full tier once it clears.
Gap assessment against a named standard.
Recurring self-review against industry practice.
Re-run on a schedule, or on a trigger event.
A diligence tool that claims everything is a tool you cannot rely on for anything. These exclusions are design decisions, not gaps in the roadmap.
Create it, drop in the data room, and watch the gap list build itself before a single model is called.