Public data, kept honest · part 3 of 4
adopted13 September 2026cost: local CPU; runtime ≈ 2.5 minutes a run
History and coverage: every file since 2010
5 runs
The question
Did the change make the dataset more accurate without moving the headline numbers for a reason nobody can explain?
adopted
History: DOL LCA files FY2010-FY2022 (4 legacy schemas mapped in scripts/convert_lca_legacy.py) + USCIS hub FY2010-FY2021 via Chrome agent; fiscal_years now 2010-2026; ZIP zero-padding — 9.64M rows -> 7.17M cases; KC last-3 headline…
Runs5
Kept5
Every run
| # | what changed | verdict | KC filings (3 y) | geo resolved % | notes |
|---|---|---|---|---|---|
| 6 | Spec (user): 'last 3 years' = 36-month window counted monthly; per-year buckets + KC by_month series | kept | 5977 | 99.98 | window first Sep2023-Aug2026 then user aligned to Jul2023-Jun2026 |
| 8 | Spec (user): data from start of 2023 -> calendar-year views (CY2022 partial..CY2026 YTD) + full-scope monthly series; USCIS Employer Data Hub FY2023-FY2026 integrated… | kept | 5977 | 99.98 | USCIS KC: 3959 approvals FY2023-26YTD, 551 petitioners; FY2023 archived CSV looks under-complete |
| 11 | USCIS hub: FY2023 (and FY2022) re-exported from Tableau via Chrome agent; FY2023 now 57,415 rows vs 33,332 in archived CSV; uscis_hub.py uses Tableau files, min FY 2023 | kept | 5977 | 99.98 | README caveat about FY2023 under-coverage removed |
| 15 | Year attribution: case counted in FY/month of its FIRST decision (status = latest). Matches h1bdatahub per-year counts exactly (Children's Mercy 28/44/31 FY2023-25 nationally) | kept | 5949 | 99.98 | KC 3-yr 5977->5949 (cases first decided before Jul 2023 leave the window); totals unchanged |
| 16 | History: DOL LCA files FY2010-FY2022 (4 legacy schemas mapped in scripts/convert_lca_legacy.py) + USCIS hub FY2010-FY2021 via Chrome agent; fiscal_years now 2010-2026; ZIP… | kept | 5905 | 99.628 | 9.64M rows -> 7.17M cases; KC last-3 headline unchanged at 5949; runtime 854s; FY2010-14 have no worksite ZIP (city-based geo) |