Skip to the document
Madhuopen lab

§3 Experiments · product

paused12 September 202613 September 2026 · 4 parts

Public data, kept honest

An autonomous loop that makes the Department of Labor's H-1B disclosure files more accurate one fix at a time — and has to prove each fix did not move the numbers for the wrong reason.

public datadata qualityH-1Bautonomous loop
In one breath22 runs, 21 kept. Every LCA filing since fiscal 2010, resolved to a metro through ZIP, county and CBSA; employers clustered by tax ID. The dataset behind bigImmigrationHub.

Summary

Disclosure files are messy in specific, repeatable ways: the same city spelled three ways, one employer under several tax identifiers, a legacy schema that reads dates differently every few years. The loop attacks one at a time. Its guardrail is the headline itself: if Kansas City's filing count moves by more than three per cent, the run must explain exactly why or it is discarded.

Rows resolved to a metro area — higher is better% of rows
00000experiment 00: 99.802%experiment 01: 99.981%experiment 02: 99.98%experiment 03: 99.98%experiment 04: 99.98%experiment 05: 99.98%experiment 06: 99.98%experiment 07: 99.98%experiment 08: 99.98%experiment 09: 99.98%experiment 10: 99.98%experiment 11: 99.98%experiment 12: 99.98%experiment 13: 99.98%experiment 14: 99.98%experiment 15: 99.98%experiment 16: 99.628%experiment 17: 99.537%experiment 18: 99.537%experiment 19: 99.537%experiment 20: 99.537%experiment 21: 99.537%best 99.981% · 010021

The protocol, in order