Jun 2026
A closed-loop architecture for building domain models on the high-fidelity data your operation already discards. Capture and verification precede training; training is the last, smallest move. On-policy distillation, the synthetic-data question, reverse-KL, early-trajectory selection, and the as-built closeout as both reward and eval — a field note and whitepaper read against a real Division 9 takeoff.