Evidence, confidence, and verification
Read high-stakes answers correctly and understand exactly what “verified” means on the platform.
Confidence follows evidence
Missing temperature, boundary, reference state, time basis, or cost scope must lower confidence or produce a next-measurement recommendation. The agent should refuse requests to hide these limitations or to convert a screening result into a bankable conclusion.
A good answer states both sides of the evidence boundary: what the current data can support and what it cannot prove.
What verified means
- Verified: a deterministic calculation, parser, source check, or other independent procedure re-established the claim.
- Unverified: the claim may be plausible, sourced, or model-generated, but no eligible independent check completed.
- Refuted or conflicting: the checker found evidence inconsistent with the claim and the answer should not present it as established.
- Unavailable: the verifier could not be reached or lacked the required input; the UI should say so rather than showing an empty verified section.
Model agreement is not verification
A second language model can challenge or refute a claim, but agreement between models cannot promote it to verified.
High-stakes review checklist
- 1
Trace the input
Confirm the number, unit, boundary, timestamp, and source file or external source.
- 2
Inspect the method
Distinguish a database lookup, deterministic equation, empirical model, and language-model synthesis.
- 3
Read the limitations
Check accuracy class, excluded uses, missing measurements, and uncertainty basis.
- 4
Try to break it
Change the dominant assumption, introduce a conflicting source, or ask what observation would reverse the recommendation.
