Files
giant/CLAUDE.md
T
lars 8ff70e3c87
CI / Lint (ruff check) (push) Failing after 4s
CI / Format (ruff format) (push) Failing after 4s
CI / Type check (ty) (push) Failing after 3s
CI / Tests (push) Failing after 3s
CI / Bump version, build & publish wheel (push) Has been skipped
analyze: drive prep/submit from the rollout YAML sidecar
`giant analyze prep` / `submit` now take the `giant rollout` YAML sidecar as
their only positional input instead of explicit --rollout/--reference/--out-dir.
The YAML's `output`/`dataset` keys name the rollout parquet and its seed file
(the reference truth), and the rest of the sidecar (checkpoint, geometry oracle,
cutoffs) flows into every plot's gallery metadata.

prep derives its own run directory next to the rollout parquet
(<...>/analysis_<id>/) holding shared.json, run_meta.json, reduced/, plots/.
compute-one and render now take just --run-dir / a run-dir argument and read the
resolved paths + metadata from run_meta.json, so the condor wrapper no longer
threads file paths. open_side scans a directory of reference shards via glob.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 17:56:36 +02:00

12 KiB
Raw Blame History

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Commands

uv sync --extra cpu                # install dependencies with CPU-only torch (standard/default)
uv sync --extra cuda               # install dependencies with CUDA 11.8 torch
uv sync --extra cpu --extra dev    # add dev extras (pytest, etc.)
uv sync --extra cpu --extra geometry  # add scikit-learn for the geometry oracle (giant rollout)
pytest               # run tests
giant train path/to/steps.parquet --mode flow   # train (flow matching)
giant train path/to/steps.parquet --mode ddpm   # train (DDPM baseline)
giant predict path/to/steps.parquet --checkpoint ckpt/best.pt   # per-step predictions
giant rollout path/to/steps.parquet --checkpoint ckpt/best.pt --geometry oracle.pkl  # full showers
giant analyze submit rollout.yaml --accounting-group cms   # parallel rollout-vs-reference analysis on HTCondor
giant analyze render <run_dir> --gallery                   # render PDFs + HTML gallery (run_dir from prep/submit)
dwarf --help                                     # dataset/tooling CLI: convert, migrate, bump-gen,
                                                  # bump-schema, status, update-manifest, create-manifest,
                                                  # make-root, build-geometry-oracle, hparam-scan
                                                  # (see scripts/dwarf.py)

cpu and cuda are mutually exclusive — pick one to select the torch build (pinned to 2.3.x; newer torch requires newer NVIDIA drivers). Plain uv sync with no extra will not install torch at all; uv has no concept of a "default extra", so --extra cpu should always be included unless you need GPU support.

Lint and type checking

uv run ruff check .     # lint
uv run ruff format .    # format
uv run ty check .       # type check

Part of the dev extra. Run these periodically (not just at commit time) to catch drift early.

Architecture

GIANT is a conditional generative surrogate for the Geant4 step function. It replaces the stochastic physics engine: given a pre-step particle state (conditioning), it samples a post-step outcome — now including the variable-length list of secondary particles the step produces (Phase 2, see Roadmap).

Data pipeline (giant/data/): parquet files from miniCaloSim are loaded into numpy arrays (loader.py), then log-transformed and rotated into a local coordinate frame where pre_dir = ẑ (transforms.py), before being wrapped in a PyTorch Dataset (dataset.py). Train/val split is by event_id to avoid leaking correlated steps from the same shower.

Stage-1 output space (9D, giant/constants.py:LOCAL_TARGET_NAMES): log_step_length, two additive-log-ratio (ALR) coordinates edep_logit/sec_logit of a deposit / secondary / post-energy simplex, post_dir (post-scattering momentum direction, unit vector in the local frame), and travel_dir (direction of post_pos - pre_pos, unit vector in the local frame). The energy simplex decodes via softmax over [edep_logit, sec_logit, 0] × pre_E so edep + e_sec + post_E == pre_E holds by construction — energy conservation is architectural, not learned (see energy_simplex_decode). post_pos is not a raw target — it's reconstructed at inference as pre_pos + step_length * world_frame(travel_dir), since step_length already encodes that displacement's magnitude and duplicating it would let the two become inconsistent.

Conditioning vector (15D continuous, COND_DIM): pre-step position, log(pre-energy), pre-step direction, layer ID (COND_DIM_BASE=8) — plus, since particle/material physical-property conditioning (model.conditioning, see below), 7 more columns: particle log(mass)/charge (PARTICLE_PHYS_DIM=2, giant/particles.py) and material Z_eff/A_eff/log(density)/log(X0)/log(λ_int) (MATERIAL_PHYS_DIM=5, giant/materials.py). n_sec and e_sec are not conditioning inputs (that was Phase 1 / the energy-conservation PoC); the model predicts them.

ConditionEncoder/SecondaryConditionEncoder (giant/model/network.py) support two mutually exclusive conditioning modes, selected per-checkpoint (model_config["conditioning"], defaulting to "embedding" for old checkpoints without the key, "physical" for new giant train runs — see --conditioning):

  • "embedding" (original Phase 2 design): a learned nn.Embedding per PDG code / material name, indexed by a dataset-scoped dense vocab (pdg_map/mat_map). Memorizes the training menu.
  • "physical" (default): the 7 physical-property columns above are each routed through a small MLP (particle_mlp/material_mlp) to the same emb_dim width the embedding tables would have produced — a drop-in replacement computable for any PDG code / material name, not just ones seen in training, which is what lets the surrogate generalize to a held-out material or species. giant/particles.py decodes nuclear/ion PDG codes (the 10LZZZAAAI scheme) via the scikit-HEP particle package with a Z/A-digit-decode fallback for isomer codes the package's ground-state-only table misses. giant/materials.py ships real Geant4-11.4.1-derived z_eff/a_eff/density/x0/lambda_int values for every material the detector geometry actually produces; the sole exception is G4_LYSO (not a stock Geant4 NIST material, never actually constructed by the geometry — see the module docstring), which stays MaterialProperties(None, ...) and raises loudly (MaterialPropertiesNotFilledError) rather than silently defaulting if it's ever requested.

Model (giant/model/network.py): a two-stage model, both checkpointed together.

  • Stage 1 — DenoisingMLP: ResBlock stack with a SinusoidalEmbedding for the flow/diffusion time variable and a ConditionEncoder fusing the conditioning. Predicts the 9D primary vector field, plus an n_sec_head classifier over {0..K_MAX} (K_MAX=15) that runs on the condition encoding alone (no diffusion noise), callable via predict_n_sec.
  • Stage 2 — SecondaryDecoder: a second flow-matching net (SecondaryConditionEncoder fuses the pre-step conditioning with the Stage-1 outcome) that generates all K_MAX secondary slots at once. Each slot is (stick-breaking energy logit, local-frame direction 3D, log-mass, charge) = SEC_SLOT_DIM=6, ordered by descending energy; slots beyond the predicted n_sec are masked. Secondary energies are a stick-breaking partition of the e_sec budget from Stage 1 (they sum to it), so the whole chain conserves energy. A secondary's mass/charge are regressed directly against a fixed physics-derived target (its ground-truth PDG code's giant.particles.particle_mass_charge) — not a learned/moving embedding target, so nothing needs detaching. No snapping at inference: the predicted (mass, charge) are used as-is as the secondary's physical identity, including for its own future conditioning if it goes on to take further steps in a rollout. A separate, reporting-only nearest-known-PDG lookup (giant.particles.nearest_known_pdg) is used purely to populate a nominal pdg label for output rows / "embedding"-mode fallback conditioning — it never feeds back into the model.

schedule.py provides both a CosineSchedule for DDPM and the flow matching loss utilities (Lipman et al. 2022 conditional flow matching).

Samplers (giant/sample.py): DDPM, DDIM, and flow matching (ODE integration, ~10 steps). Flow matching is the primary mode.

Validation (giant/validate.py): step-level marginal comparisons.

Analysis (giant/analysis/, giant analyze CLI): a lean, streaming rollout-vs-reference plotting pipeline that compares one autoregressive giant rollout (for a given checkpoint) against a held-out miniCaloSim reference steps file, and produces publication-styled PDFs assembled into an HTML gallery. It exploits the fact that rollout output and a raw reference file share a world-frame physical column subset under identical names (pre_*/post_*/edep/step_length/pdg/material/event_id), so no ALR/local-frame decode is needed — everything is world-frame mm/MeV. Structure: sources.py (canonical LazyFrames + synthetic-termination-row filtering + the secondary view, which is generation>0 & step_no==0 rollout tracks vs exploded sec_*_list reference columns), reduce.py (the streaming primitives — a single hist1d group_by([group,bin]).len() pass, per-event scalars, edep-weighted depth/transverse profiles, species share, leakage), grouping.py/context.py (fixed bin edges + energy-quantile/pdg/material group sets resolved once by prep into shared.json, so every compute job is one pass with no range scan), catalog.py (the declarative PlotSpec registry — marginals × {overall,energy,pdg,material}, per-event totals, shower profiles, species/leakage, secondaries), and render.py (the only module importing ETPlot's plotstyle/LaTeX; dispatches on Reduced.kind, writes PDFs + metadata.yaml). Input is a giant rollout YAML sidecar (condor.py:load_rollout_yaml): its output/dataset keys name the rollout parquet and the seed file (= the reference truth), and the rest of the YAML (checkpoint, geometry oracle, cutoffs) flows into each plot's gallery metadata. prep derives its own run directory next to the rollout parquet (<...>/analysis_<id>/) holding shared.json, run_meta.json, reduced/, plots/. Compute/render split: giant analyze submit rollout.yaml runs prep then submits one HTCondor job per plot (compute-one --run-dir, polars/numpy only — no LaTeX on workers), each writing a small reduced/<id>.json; the local giant analyze render <run_dir> turns those into the styled PDF/gallery tree. See giant/analysis/__init__.py.

Shower rollout (giant/rollout.py, giant rollout CLI): autoregressively steps the two-stage model into a full shower — each primary post-step becomes the next pre-step, secondaries are pushed as new tracks, and per-step material/layer_id come from a GeometryOracle (giant/geometry.py, built via dwarf build-geometry-oracle) that learns position → (material, layer_id) from data and flags detector escape by nearest-neighbour distance. Tracks terminate on energy cutoff, per-track max steps, escape, or natural end; energy is deposited locally on every stop except escape (leakage), so showers conserve energy by construction.

Roadmap

Phase 1 (done): number of secondaries and their total energy were conditioning inputs; the model predicted only the 9D primary post-step (energy-conservation PoC).

Phase 2 (implemented — baseline): the two-stage model above jointly predicts n_sec, the energy simplex (e_sec falls out of it), and each secondary's energy/direction/species, so a rollout is self-contained (no ground-truth secondary counts injected). This is the "get a baseline out" track agreed with Jan & Tobias (2026-07-07).

Physical-property conditioning (implemented): model.conditioning = "physical" | "embedding" (see above) replaces the learned PDG/material embeddings with a small MLP over particle mass/charge and material Z_eff/A_eff/density/X0/λ_int, and Stage 2 predicts a secondary's mass/charge directly instead of a snapped species embedding. "embedding" stays available as the generalization-comparison baseline. giant/materials.py's table is already filled with real values for every material the geometry produces. Not yet done: the actual held-out-material/species generalization comparison against the "embedding" baseline is unrun — the 34GB multi-material dataset at the repo root (6 materials, 237 PDG codes including nuclear/ion codes) is the natural dataset for that experiment.

Next directions (parallel, not yet built): faster-eval architectures measured against a ~10× native-Geant4 budget — a Wasserstein-GAN throwaway (single-pass eval) and a mixture-of-experts / routing tree of small nets selected per call (pdg / energy / process), with soft/differentiable gating on continuous routing axes; a sampling-calorimeter (multi-material) dataset. See the knowledge base (/home/lars/knowledge-base/meta/roadmap.md).