for judges
The 30‑second path
PlumeRank ranks which wildfire-camera tiles a dispatcher should open next, calibrated to a fixed budget of 2 alerts per camera‑night (one 81-minute camera window, the length of a FIgLib sequence), and measures itself on median minutes after ignition until the right camera reaches the top of that ranking.
1. Click these, in order
- The Wall — the live ranking, no login, replaying development fires and fire-free confusion-set windows. Its threshold is labelled on screen.
- /build-info — real
cv2.getBuildInformation()output, the instance id, the AMI id and the git commit this box is running, read live off the box. - /results — the build-info block in-page and the held-out headline: systems on one clock, operating curve, sensitivity strip and the ablation grid, from the committed run.
2. The receipt
- 734 passed, 1 skipped locally (CI, without the corpus, runs 722 and skips 13), 97% coverage on
plumerank/ - Property-based tests (
tests/test_invariants.py, hypothesis, 1,500 generated cases) checking that the fused hazard stays in[0, 1]and the alert budget is never overspent ruff check .andmypyboth clean;scripts/check_claims.py(insidemake test) mechanically checks every route, dependency, Makefile target and OpenCV symbol the README and the other judge-facing documents name against the real source- Development set (in-sample, never the claim): 17/29 detected, p50 32.0 min (results/arm_dev_2026-10-10.json; rates and baselines on /results)
- Held-out set — the headline: 27/54 detected, median 38.0 min at 2.0 alerts per camera-night; like-for-like baselines 18/54 and 20/54, neither reaching a median (results/bench_2026-10-10.json) — see §4
- One AWS
t4g.small(Graviton) and one S3 bucket. No Lambda, no managed vision API.
3. Reproduce it yourself
python3 -m venv .venv && .venv/bin/pip install -r requirements-dev.lock
make test # ~2 min: 735-test offline suite + coverage + the claim-consistency check
make verify # ≤ 3 min: rebuilds the venv, re-runs the suite, recomputes the
# micro-corpus metrics against the committed results. Does NOT
# reproduce the published headline, and says so on screen.The pair that reproduces the held-out headline itself: make fetch then make bench — about 35 minutes of fetch, then the evaluation (12m29s on the machine that produced the published file). Nobody needs to run that to evaluate this submission.
4. How the headline was produced — and how far to trust it
The held-out run happened once, on 2026-10-10: the constants were frozen and committed first (results/constants.sha256), and the single read is logged with its commit (results/hold_access.log). 54 sequences were measured — 60 in the manifest, 6 excluded by the published rule.
27 of 54 is exactly half, so the median is on a knife edge: against an ignition clock shifted by −5 to 0 minutes it reads 32–38 minutes, and from +1 to +5 minutes it does not exist. The detection counts — 27 against 18 and 20 — are the steadier comparison.
5. Honest limitations
- Daylight only — no IR hardware, no night detection.
- Cloud, not fog, produced the most false alarms on the confusion set (measured before the per-camera baseline was adopted). A blind census found the “fog” class picks a marine layer no more often than chance (29% vs 29%), so it is mostly overcast; the threshold uses every confusion window, so only per-class claims change.
- The soft-boundary statistic the design started from (S3) is not predictive on its own — AUC 0.44 on the development frames where it is defined — yet it carries the largest fusion weight, and no S3-off ablation has run. The measured result comes from the per-camera baseline and the budget-calibrated ranking, not from that one signal.
- Stabilisation registers each frame to its raw predecessor, and under moving cloud its whole-frame estimate follows the sky: checked against the ground, 48 of DEV’s 55 warps of 1 px or more are one genuinely shaking camera, the rest follow cloud (results/registration_2026-10-10.json). Replaying the fire-free set without those warps moves its false alarms from 64 to 63.
- The evidence crop is the largest moving component, not a fire localisation: at each DEV fire’s first rank-1 or alert tick it contains the plume on 9 of 21, cloud on 7 (results/evidence_first_hits_2026-10-10.json). A cloud in the crop means “look at the tile”, not “all clear”.
- One instance, no HA — every published number is committed in the repository.
- No dispatch, no suppression, no upload path — architecturally absent.
6. Links
Wall · /results · /build-info · /health · source on GitHub · pitch deck