d858226294
CI / Format (ruff format) (push) Successful in 28s
CI / Lint (ruff check) (push) Successful in 35s
CI / Sync project version with tag (push) Has been skipped
CI / Type check (ty) (push) Successful in 28s
CI / Lint (ruff check) (pull_request) Successful in 30s
CI / Sync project version with tag (pull_request) Has been skipped
CI / Type check (ty) (pull_request) Successful in 52s
CI / Tests (push) Successful in 4m24s
CI / Tests (pull_request) Successful in 3m35s
CI / Format (ruff format) (pull_request) Successful in 28s
A fixed comparison point for future architecture variants, so each experimental axis (routed trunk, WGAN generators, attention history, shared conditioning) is a single edit away from one known config. flow/flow autoregressive, hidden_dim 512 / 6 blocks per stage, physical conditioning, no router, 7.70M params. Chosen by ranking the five runs in analysis_runs/ by mean Jensen-Shannon divergence against the Geant4 reference: unrouted flow wins (0.172) over routed flow (0.197/0.200) and both WGAN runs (0.218/0.234), with the lead concentrated in per-event total deposited energy and the per-PDG marginals. batch_size 36864 is sized for one L40S on deepthought2 from a measured linear fit of this config's training step (reserved MiB = 0.9736 * bs + 115), giving ~36 GiB, 78% of the card. The comments record two measured facts that are easy to get wrong: WGAN is slower to *train* than flow (n_critic plus the gradient-penalty double-backward), its advantage being inference-only; and sample_secondaries_ar loops over all k_max slots unconditionally rather than short-circuiting on n_sec, which is what makes the autoregressive decoder the dominant cost on both axes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>