Skip to the document
Madhuopen lab

AI Insights: Latency & Cost Optimization · part 8 of 17

insight27 July 2026cost: ~$5 (full A100 deploy→delete + judge API)

An independent judge: does the quality story survive?

Re-scoring everything with Claude Opus 5 + Sonnet 5 instead of GPT-5.1 grading itself

The question

Every score so far came from GPT-5.1 grading answers — including its own. Do the conclusions hold under a different vendor's judge?

insight

The ranking survives, absolute scores drop 1.3–1.5 points (the old scores were inflated), and the independent judges caught a fabrication the self-judge missed.

In one breathThe key methodology check in the project: every score before this point was self-graded. Re-judging with a different vendor kept the ranking but lowered the absolute numbers — treat any single-judge score as relative, not absolute.
GPT-5.1 single8.7 → 7.35old self-judge → Opus 5
Qwen single8.0 → 6.45gap widens 0.7 → ~1.0
Judge agreement±0.5Opus ≈ Sonnet everywhere
Rankingunchangedno conclusion flips

Results (median of 10 runs, Opus 5 / Sonnet 5)

Results (median of 10 runs, Opus 5 / Sonnet 5)
Model & modeOld score (GPT-5.1 judging)Claude Opus 5Claude Sonnet 5
GPT-5.1 single8.77.357.0
GPT-5.1 parallel8.06.06.0
Qwen single8.06.456.0
Qwen parallel5.54.95.0

Three findings

  • The ranking survives — GPT-5.1 ahead of Qwen, single ahead of parallel, exactly as before. No conclusion flips.
  • Independent judges are 1.3–1.5 points stricter across the board — the old 8.x scores were inflated for both models roughly equally; self-judge bias itself was small (~0.2).
  • The Claude judges caught a real defect the self-judge missed: Qwen occasionally invents an itemized-deduction insight with a made-up dollar figure ($12,300 "excess") for a client who takes the standard deduction. The GPT-5.1 judge had scored those answers 8/10; the independent judges flagged the fabrication specifically.
LLM-as-judgeClaude Opus 5Claude Sonnet 5evaluation bias