The whole point of conditioning="physical" is generalizing to a
species/material outside the training menu, but two independent code
paths still hard-required training-vocab membership:
- giant/data/transforms.py: build_cond_features unconditionally raised
KeyError on an out-of-vocab pdg/material. _vectorized_map_lookup
gains a strict=False mode (dummy index instead of raising), used only
under conditioning="physical" where ConditionEncoder never reads
cond_cat anyway; "embedding" mode is untouched and still raises,
since cond_cat IS the conditioning signal there.
- giant/rollout.py: the known_pdg termination gate still killed a track
on step 1 for any pdg outside pdg_map, regardless of conditioning
mode. Now skipped entirely under conditioning="physical".
- giant/model/network.py: PdgRouter/ProcessRouter always build their
own training-vocab nn.Embedding independent of conditioning, silently
reintroducing the same limitation at the routing layer. build_models
now raises loudly if conditioning="physical" is paired with either
router type, rather than silently building a model that can't
generalize the way it claims to.
This unblocks the held-out-species/material generalization experiment
against the multi-material dataset (see CLAUDE.md roadmap). Each fix
has a regression test, including an end-to-end rollout test seeded
with a resolvable-but-out-of-vocab PDG code.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- make_seed_frontier only resolves particle mass/charge in "physical"
mode, so "embedding"-mode rollouts no longer crash on a seed PDG code
giant.particles can't resolve (the TERM_UNKNOWN_PDG gate now handles it).
- nearest_known_pdg skips unresolvable candidate PDG codes instead of
raising and killing the whole rollout/predict run.
- predict/rollout fail with a clear message when a checkpoint predates
the sec_phys normalizer, instead of a bare KeyError.
- validate_marginals' phys_kl degrades to NaN (matching the
energy_fraction_kl pattern) instead of crashing when a validated batch
has zero secondaries on either side.
- Correct CLAUDE.md's stale claim that the materials table is unfilled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds model.conditioning = "physical" | "embedding": physical mode routes
particle mass/charge and material Z_eff/A_eff/density/X0/lambda_int through
small MLPs to replace the learned PDG/material embedding tables, so the
surrogate generalizes to PDG codes/materials outside the training vocab
instead of memorizing it. "embedding" stays available as the comparison
baseline (old checkpoints without the key default to it).
Stage 2 now regresses a secondary's mass/charge directly against a fixed
physics-derived target instead of a learned/snapped embedding, and uses no
snapping at inference — the model's raw predicted (mass, charge) is the
secondary's physical identity, including for its own further rollout steps.
A separate reporting-only nearest-known-PDG lookup (never fed back into the
model) populates output pdg columns / the embedding-mode rollout fallback.
giant/materials.py's table is populated with Geant4's own built-in NIST
constants (Z_eff, A_eff, density, X0, lambda_int), extracted directly from
the Geant4 11.4.1 build vendored in minicalosim via G4NistManager rather
than hand-typed literature values. G4_LYSO is left unfilled: confirmed (both
by runtime lookup and by searching minicalosim's history) that it's never
actually a constructed Geant4 material there, only documentation/UI color-map
text.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
_Recorder previously accumulated every generated step across all events/
tracks/steps in Python lists, materialised once at the end and written
via a single pq.write_table — memory scaled with n_events * max_steps *
avg_tracks_per_event. rollout() now takes an optional on_chunk callback
that streams each non-empty batch immediately (fixed per-key dtypes via
_RECORD_DTYPES keep every chunk's table schema identical, which
pq.ParquetWriter requires across writes); giant rollout wires this to an
incrementally-written ParquetWriter, mirroring the row-group streaming
giant predict already does on its input side. Without on_chunk, rollout()
keeps its old buffered return for existing callers/tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
miniCaloSim's detector is a stack of planar layer slabs along one axis, so
material/layer_id are a pure function of depth. The new "slab" method
exploits this with an exact O(log #segments) binary search over
depth-axis segment boundaries, instead of a nearest-neighbour search over
hundreds of thousands of reference points — much cheaper per call, which
matters since the oracle is queried on every autoregressive rollout step.
"knn"/"svm" remain as fallbacks for non-slab geometries.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Closes the loop from single-step prediction into full showers:
- giant/geometry.py + `dwarf build-geometry-oracle`: learn position ->
(material, layer_id) from data (KNN/SVM) to supply the conditioning the
surrogate does not predict; flag detector escape by NN distance.
- giant/rollout.py: breadth-first batched frontier that steps all active
tracks, spawns secondaries as new tracks, and terminates on energy cutoff,
per-track max steps, escape, or natural end. Energy is deposited locally on
every stop except escape (leakage), so showers conserve energy exactly.
- `giant rollout` CLI: seed from real events (argmax pre_E), load checkpoint,
write a world-frame steps parquet + YAML sidecar.
- giant/analysis.py: compute_rollout_observables + plot_rollout_* for
single-sided longitudinal/transverse/total-energy shower profiles;
analysis/export_rollout_observables.py driver.
- scikit-learn added as an optional `geometry` extra (lazy-imported).
- Tests: tests/test_geometry.py, tests/test_rollout.py.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>