Skip to the document
Madhuopen lab

Reading documents you can trust: the numbered findings · part 3 of 4

adopted6 August 2026cost: part of the $3.50 total

What works: the trust machinery

Second readings, deterministic rules, zoomed re-asks, honest abstention

The question

What did 9 measured findings establish about what works: the trust machinery?

adopted

Diversity beats repetition..

Findings9
NumberedL13–L21

13. Diversity beats repetition.

Cross-channel agreement (image reader vs OCR-text reader — different eyes, different mistakes): 93.2% exact when green vs 91.4% for same-model unanimity. This is the key external signal.

14. Fused signals make the best router.

Channel-agree AND vote-unanimous: 84% of fields auto-fillable @ 94.1% exact / 0.9856 ANLS; the red 16% is only ~69–72% right → human review. The trust sorter is the product.

15. Agreement-path arithmetic:

agreement (86% of fields) = 93% exact → auto-fill; disagreement (14%) = 72% → review. Answer accuracy comes from the reader; trust comes from the machinery.

16. Table-structure injection fixes the wrong-row class — when routed.

Phase-1 run (2026-08-05): DocLayout-YOLO + SLANet-plus → r=/c=-indexed HTML skeletons injected for the 107/300 table-page questions. Overall 0.9525 → 0.9536; the baseline-imperfect slice jumped 0.665 → 0.822 (textbook wrong-cell fix: '50.0'→'30/60'); but the baseline-perfect slice slipped 1.000 → 0.992 (formatting/distraction tax). Structure injection is a routed action for disputed/table-relevant fields, not a default. Published backing: +7.1pp table QA (2505.17625), DocTags OOD F1 0.37→0.92 (2605.19866).

17. Crop-and-re-ask resolves consensus errors that sampling can't

— and when even the crop is ambiguous, cross-copy/external checks decide, and if those fail too the honest output is a REVIEW FLAG, not a guess. (W-2 case: box 17 read "501.13" twice confidently; truth was 101.13 — visible only in the higher-resolution original. Flagging beat trusting the re-ask.)

18. Deterministic validators are free evidence nobody else ships.

W-2 case: SS tax = exactly 6.2% × wages and Medicare = 1.45% × wages confirmed the right cells were read; Indiana state rate ≈ 3% flagged the implausible 501.13. Arithmetic, format (SSN/EIN/date/money), checksums, cross-field rules — wire them into routing (see trust_layer_survey.md: no competitor does this).

19. A missing citation is a signal, not a failure.

The aligner refuses to draw a box when the VLM's reading and OCR can't be reconciled (fuzzy score below threshold). Wrong box would be worse than no box. Known gaps: short/ duplicate values snap to twin cells (P2 in queue); squished OCR text ("2203RDAVENE") blocks multi-line matches — both deterministic aligner fixes, not model problems.

20. Null has three meanings and we must split them

(logged 08-04): "page says no" vs "page doesn't say" vs "we failed to extract". Most cross-prompt disagreement in a clinical-extraction study sat on exactly this distinction. Detectors can't fire on an ambiguous null.

21. Ground-truth-free stability measurement exists

(2606.05970): κ agreement across prompt paraphrases, stratified — runs on unlabeled documents, bounds stability (not accuracy). Our path to calibrating on thousands of unlabeled docs.

extractionfindingsC