§3 · Experiments · the lab notebook
A claim, its evidence, a verdict.
Every experiment is written into the same notebook: the question it asked, what it cost, what was measured, and whether the idea was adopted, rejected, or needs more data. Rejected ideas stay — they are the most expensive thing I own. 31 parts so far: 15 adopted, 4 rejected, 11 insights, 1 waiting on more data.
AI Insights: Latency & Cost Optimization
Cutting the insights panel from a 24 s wait to under a second, at about a tenth of the cost
HanuDB, tuned while I slept
A document database improved by an autonomous agent: one change at a time, two benchmark runs, hard gates, keep or discard.
Reading documents you can trust: the numbered findings
39 findings from eight days of measured runs on scanned documents — what sets accuracy, what is a dead end, and what makes an answer trustworthy.
Public data, kept honest
An autonomous loop that makes the Department of Labor's H-1B disclosure files more accurate one fix at a time — and has to prove each fix did not move the numbers for the wrong reason.
Next in the notebook
Retrieval, measured
The RAG micro-experiments behind the measurement series: chunking, hybrid, reranking, the semantic cache.