lars 1115451c8e
CI / Format (ruff format) (push) Successful in 27s
CI / Lint (ruff check) (push) Successful in 28s
CI / Sync project version with tag (push) Has been skipped
CI / Type check (ty) (push) Successful in 23s
CI / Tests (push) Successful in 53s
CI / Format (ruff format) (pull_request) Successful in 28s
CI / Lint (ruff check) (pull_request) Successful in 28s
CI / Sync project version with tag (pull_request) Has been skipped
CI / Type check (ty) (pull_request) Successful in 22s
CI / Tests (pull_request) Successful in 52s
Make default checkpoint out_dir name reflect only non-default hyperparams
Previously the same fixed 7 fields (mode/hidden_dim/n_blocks/emb_dim/
conditioning/lr/batch_size) were always baked into the name, even for a
vanilla run, and router config wasn't represented at all. Now
default_out_dir_name only includes fields that differ from
DEFAULT_CONFIG, adds router/seed/epochs as candidates, and caps at 6
shown fields with a hashed overflow suffix for heavily-swept configs.
2026-07-29 11:21:02 +02:00
2026-07-27 11:19:06 +02:00

giant

Geant4 Inference via Autoregressive Neural sTep surrogate — a play on Geant4 and the step function being the computationally heaviest part of the simulation.

Proof-of-concept surrogate model for the Geant4 step function. Given a pre-step particle state, the model samples a physically plausible post-step outcome — the primary's continuation plus the secondary particles it produces — replacing the stochastic Geant4 physics engine with a trained conditional generative model.

Training is driven entirely from parquet files of the miniCaloSim steps tree. No Geant4 runtime dependency.

Architecture

A two-stage conditional flow matching model (Lipman et al. 2022): a small MLP learns a vector field mapping noise → step outcomes in ~10 ODE steps per sample. Falls back to DDPM for comparison.

Stage 1 — primary (9D, diffused):

Index Variable Encoding
0 step_length [mm] log
12 edep_logit, sec_logit ALR coords of the deposit/secondary/post-energy simplex
35 post_dir in local frame unit vector
68 travel_dir (post_pos pre_pos) in local frame unit vector

The two energy logits decode via softmax over [edep_logit, sec_logit, 0] × pre_E, so edep + e_sec + post_E == pre_E exactly — energy conservation is built into the parametrization, not left to the loss. Stage 1 also has a classifier head predicting the number of secondaries n_sec ∈ {0..15} from the conditioning alone.

Both post_dir and travel_dir are expressed in the coordinate frame where pre_dir = ẑ, making the scattering distribution nearly azimuthally symmetric. post_pos itself is not a raw target — it's reconstructed at inference as pre_pos + step_length * world_frame(travel_dir), so the two stay consistent by construction instead of being learned (and potentially diverging) independently.

Stage 2 — secondaries (SecondaryDecoder): conditioned on the pre-step state and the Stage-1 outcome, a second flow net generates all K_MAX = 15 secondary slots at once. Each slot carries a stick-breaking energy fraction, a local-frame direction, and a continuous particle-type embedding (snapped to the nearest PDG at inference), ordered by descending energy; slots beyond the predicted n_sec are masked. The secondary energies are a stick-breaking partition of the e_sec budget from Stage 1, so the full chain conserves energy. Each secondary's momentum is reconstructed afterward from (energy, direction, species) rather than predicted.

Conditioning: PDG code (embedding), pre-step position, log(pre-energy), pre-step direction, material (embedding), layer ID. (n_sec / e_sec are outputs now, not inputs.)

Roadmap

Phase 1 (done): n_sec and total secondary energy e_sec were conditioning inputs; the model predicted only the 9D primary post-step (energy-conservation PoC).

Phase 2 (implemented — baseline): the two-stage model above predicts n_sec and each secondary's energy, direction, and species jointly with the primary, so a shower rollout is fully self-contained.

Next directions: faster-eval architectures against a ~10× native-Geant4 budget (Wasserstein-GAN, mixture-of-experts routing tree), a multi-material sampling-calorimeter dataset, and physical-property conditioning over learned embeddings.

Data

Input: parquet files produced by miniCaloSim, or converted from a ROOT file via uv run dwarf convert. Each row is one Geant4 step. Train/val split is by event_id (not row shuffle) to avoid leaking correlated steps from the same shower.

Project structure

giant/
├── giant/
│   ├── data/
│   │   ├── loader.py       # parquet → numpy arrays (incl. streaming/chunked reads)
│   │   ├── transforms.py   # log transforms, local-frame rotation, energy simplex, secondary encode/decode
│   │   └── dataset.py      # StepsDataset / StreamingStepsDataset (PyTorch)
│   ├── model/
│   │   ├── network.py      # SinusoidalEmbedding, ConditionEncoder, DenoisingMLP, SecondaryDecoder
│   │   └── schedule.py     # CosineSchedule (DDPM) and flow matching utilities
│   ├── constants.py        # output/conditioning dims, K_MAX, secondary slot layout, schema keys
│   ├── config.py           # default hyperparameters, TOML config merging, device autodetect
│   ├── pipeline.py         # builds datasets/normalizers and kicks off a training run
│   ├── train.py            # two-stage training loop, checkpointing, graceful shutdown
│   ├── sample.py           # DDPM / DDIM / flow matching samplers + secondary sampling
│   ├── geometry.py         # GeometryOracle: position → (material, layer_id, escaped) for rollout
│   ├── rollout.py          # autoregressive shower rollout driver
│   ├── validate.py         # step-level marginal + KL-divergence validation
│   ├── analysis.py         # step- and shower-level diagnostics: marginals, correlations, rollout observables
│   └── cli.py              # `giant train` / `predict` / `rollout` Typer app
├── scripts/                            # dataset/tooling logic, unified under the `dwarf` CLI (`uv run dwarf --help`)
│   ├── dwarf.py                        # Typer app: convert, migrate, bump-gen, bump-schema, status,
│   │                                    #   update-manifest, create-manifest, make-root, hparam-scan
│   ├── steps_to_parquet.py             # ROOT → parquet conversion (uproot/awkward/polars) — `dwarf convert`
│   ├── steps_to_parquet_parallel.py    # fan out conversion over several ROOT files — `dwarf convert --jobs N`
│   ├── migrate_geant_steps.py          # one-time move into the raw/processed/pools/derived layout — `dwarf migrate`
│   ├── bump_dataset_version.py         # cut a new raw gen or parquet schema, with a logged reason —
│   │                                    #   `dwarf bump-gen` / `bump-schema` / `status` / `update-manifest` / `create-manifest`
│   ├── create_root_files.py            # generate new ROOT shards via a minicalosim executable — `dwarf make-root`
│   ├── geometry_oracle.py              # fit a position → (material, layer_id) oracle — `dwarf build-geometry-oracle`
│   └── hparam_scan.py                  # hyperparameter grid scan over `giant train` runs — `dwarf hparam-scan`
└── tests/

Setup

uv sync --extra cpu               # CPU-only torch (use --extra cuda for CUDA 11.8 instead)
uv sync --extra cpu --extra dev   # add dev tools (pytest, ruff, ty)

cpu and cuda are mutually exclusive extras selecting the torch build; plain uv sync installs no torch at all. See CLAUDE.md for details.

Training, prediction, and rollout

giant train path/to/steps.parquet --mode flow
giant predict path/to/steps.parquet --checkpoint checkpoints/.../best.pt

# Full-shower rollout needs a geometry oracle (position → material/layer_id):
uv sync --extra cpu --extra geometry
dwarf build-geometry-oracle path/to/steps.parquet --out oracle.pkl
giant rollout path/to/steps.parquet --checkpoint checkpoints/.../best.pt --geometry oracle.pkl

train/predict accept a TOML config file (--config) and CLI overrides for hyperparameters; see --help on any command for the full option list. giant train --wandb logs per-epoch metrics (the same ones written to metrics.csv) to Weights & Biases; requires uv sync --extra wandb. giant rollout seeds showers from the highest-energy entry step per event, then autoregressively steps the two-stage model to completion — pushing secondaries as new tracks and looking up material/layer_id from the oracle at each step. Tracks terminate on energy cutoff, per-track max steps, detector escape, or natural end; energy is deposited locally on every stop except escape (leakage), so showers conserve energy by construction.

Validation

giant.validate.validate_marginals runs step-level marginal and KL-divergence checks during training (--validate-every). For deeper diagnostics on a trained checkpoint — stratified marginals, correlation structure, physical-constraint violations, and shower-level rollout observables (longitudinal/transverse profiles, PDG energy shares) — see giant.analysis.

Development

uv run pytest            # run tests
uv run ruff check .      # lint
uv run ruff format .     # format
uv run ty check .        # type check
S
Description
GIANT — conditional generative surrogate for the Geant4 step function. Two-stage flow-matching model that replaces Geant4's stochastic shower physics, predicting post-step outcomes and secondary particles while conserving energy by construction.
Readme 14 MiB
Languages
Python 100%