Skip to the document
Madhuopen lab

Reading documents you can trust: the numbered findings · part 1 of 4

insight6 August 2026cost: part of the $3.50 total

What sets answer quality

The reader decides; machinery around it cannot add accuracy it does not have

The question

What did 4 measured findings establish about what sets answer quality?

insight

Reader quality is set by the reader — machinery can't add to it..

Findings4
NumberedL1–L4

1. Reader quality is set by the reader — machinery can't add to it.

Qwen3-VL-235B @ 2400px = 0.9525 ANLS. Every inference-time scheme we tried (judges, voting, crops, three readers) landed 0.94–0.95. The DocVQA inference-time phase is CLOSED at 0.9525; the leaderboard 0.9725 gap is protocol/ambiguity, not fixable cheaply.

2. Resolution before intelligence.

1540px → 2400px inputs: +0.7 ANLS, the single largest cheap win. Confirmed live on the W-2 case (below): a 620px page produced 0 readable fields and two wrong-cell errors after upscale; the 1536px original of the same form was simply correct, first pass.

3. Don't swap the reader for a frontier model.

Qwen3-VL-235B is #1 on DocVQA-test (0.971) above GPT-5.x/Gemini-3/Opus-class at 4–24× less cost, and wrong-row errors persist across ALL model families (arXiv 2606.32029) — paying 24× buys the same class of mistake.

4. Model size matters at the extraction task only for open-vocabulary fields.

(arXiv 2606.05970, logged 08-04): for closed-value fields (yes/no, dates, amounts) prompt/schema choices dominate and model size is near-neutral; for many-class open fields, model choice changes the answer on ~half the documents. The field's value space is known before any call — so it's a routing feature.

extractionfindingsA