AI Insights: Latency & Cost Optimization · part 8 of 17
insight27 July 2026cost: ~$5 (full A100 deploy→delete + judge API)
An independent judge: does the quality story survive?
Re-scoring everything with Claude Opus 5 + Sonnet 5 instead of GPT-5.1 grading itself
The question
Every score so far came from GPT-5.1 grading answers — including its own. Do the conclusions hold under a different vendor's judge?
insight
The ranking survives, absolute scores drop 1.3–1.5 points (the old scores were inflated), and the independent judges caught a fabrication the self-judge missed.
In one breathThe key methodology check in the project: every score before this point was self-graded. Re-judging with a different vendor kept the ranking but lowered the absolute numbers — treat any single-judge score as relative, not absolute.
GPT-5.1 single8.7 → 7.35old self-judge → Opus 5
Qwen single8.0 → 6.45gap widens 0.7 → ~1.0
Judge agreement±0.5Opus ≈ Sonnet everywhere
Rankingunchangedno conclusion flips
Results (median of 10 runs, Opus 5 / Sonnet 5)
| Model & mode | Old score (GPT-5.1 judging) | Claude Opus 5 | Claude Sonnet 5 |
|---|---|---|---|
| GPT-5.1 single | 8.7 | 7.35 | 7.0 |
| GPT-5.1 parallel | 8.0 | 6.0 | 6.0 |
| Qwen single | 8.0 | 6.45 | 6.0 |
| Qwen parallel | 5.5 | 4.9 | 5.0 |
Three findings
- The ranking survives — GPT-5.1 ahead of Qwen, single ahead of parallel, exactly as before. No conclusion flips.
- Independent judges are 1.3–1.5 points stricter across the board — the old 8.x scores were inflated for both models roughly equally; self-judge bias itself was small (~0.2).
- The Claude judges caught a real defect the self-judge missed: Qwen occasionally invents an itemized-deduction insight with a made-up dollar figure ($12,300 "excess") for a client who takes the standard deduction. The GPT-5.1 judge had scored those answers 8/10; the independent judges flagged the fabrication specifically.