Strong: extraction, normalisation, contradiction-finding
Pulling structured claims out of a data room, normalising three years of inconsistently formatted financials, and flagging where the deck disagrees with the contracts are tasks with objective answers and immediate verification. This is where the hours go and where models return them.
- Claim extraction with a page-level citation for each item
- Cross-document contradiction detection
- Competitor and pricing sweeps from public sources
- First-draft diligence questions derived from gaps in the material
Weak: anything requiring an unavailable fact
Asked for a market size that no public source establishes, a model will produce a plausible number rather than refuse. The failure is not that it guessed; it is that the guess arrives with the same tone as a verified figure. Every downstream reader then inherits false precision.
The control is structural: require a source URL or document reference for any figure labelled verified, and render everything else as an estimate with its basis stated. A system that cannot say "unknown" is not usable for diligence.
Controls that make model output auditable
| Risk | Control | Observable signal |
|---|---|---|
| Fabricated figures | Source required for verified status | Share of claims with live citations |
| Silent drift | Re-run scoring on a fixed test deal | Score variance between runs |
| Over-confidence | Confidence band tied to evidence coverage | Confidence falls as coverage falls |
| Unexplainable scores | Stored data → evidence → reasoning chain | Every score opens to its inputs |
The division of labour that works
Let the model compress the material and surface what disagrees. Let the analyst decide what it means. The moment a system starts issuing verdicts whose inputs cannot be inspected, it has stopped being research infrastructure and started being an unaccountable opinion.