lars 8ff70e3c87
CI / Lint (ruff check) (push) Failing after 4s
CI / Format (ruff format) (push) Failing after 4s
CI / Type check (ty) (push) Failing after 3s
CI / Tests (push) Failing after 3s
CI / Bump version, build & publish wheel (push) Has been skipped
analyze: drive prep/submit from the rollout YAML sidecar
`giant analyze prep` / `submit` now take the `giant rollout` YAML sidecar as
their only positional input instead of explicit --rollout/--reference/--out-dir.
The YAML's `output`/`dataset` keys name the rollout parquet and its seed file
(the reference truth), and the rest of the sidecar (checkpoint, geometry oracle,
cutoffs) flows into every plot's gallery metadata.

prep derives its own run directory next to the rollout parquet
(<...>/analysis_<id>/) holding shared.json, run_meta.json, reduced/, plots/.
compute-one and render now take just --run-dir / a run-dir argument and read the
resolved paths + metadata from run_meta.json, so the condor wrapper no longer
threads file paths. open_side scans a directory of reference shards via glob.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 17:56:36 +02:00

giant

Geant4 Inference via Autoregressive Neural sTep surrogate — a play on Geant4 and the step function being the computationally heaviest part of the simulation.

Proof-of-concept surrogate model for the Geant4 step function. Given a pre-step particle state, the model samples a physically plausible post-step outcome — the primary's continuation plus the secondary particles it produces — replacing the stochastic Geant4 physics engine with a trained conditional generative model.

Training is driven entirely from parquet files of the miniCaloSim steps tree. No Geant4 runtime dependency.

Architecture

A two-stage conditional flow matching model (Lipman et al. 2022): a small MLP learns a vector field mapping noise → step outcomes in ~10 ODE steps per sample. Falls back to DDPM for comparison.

Stage 1 — primary (9D, diffused):

Index Variable Encoding
0 step_length [mm] log
12 edep_logit, sec_logit ALR coords of the deposit/secondary/post-energy simplex
35 post_dir in local frame unit vector
68 travel_dir (post_pos pre_pos) in local frame unit vector

The two energy logits decode via softmax over [edep_logit, sec_logit, 0] × pre_E, so edep + e_sec + post_E == pre_E exactly — energy conservation is built into the parametrization, not left to the loss. Stage 1 also has a classifier head predicting the number of secondaries n_sec ∈ {0..15} from the conditioning alone.

Both post_dir and travel_dir are expressed in the coordinate frame where pre_dir = ẑ, making the scattering distribution nearly azimuthally symmetric. post_pos itself is not a raw target — it's reconstructed at inference as pre_pos + step_length * world_frame(travel_dir), so the two stay consistent by construction instead of being learned (and potentially diverging) independently.

Stage 2 — secondaries (SecondaryDecoder): conditioned on the pre-step state and the Stage-1 outcome, a second flow net generates all K_MAX = 15 secondary slots at once. Each slot carries a stick-breaking energy fraction, a local-frame direction, and a continuous particle-type embedding (snapped to the nearest PDG at inference), ordered by descending energy; slots beyond the predicted n_sec are masked. The secondary energies are a stick-breaking partition of the e_sec budget from Stage 1, so the full chain conserves energy. Each secondary's momentum is reconstructed afterward from (energy, direction, species) rather than predicted.

Conditioning: PDG code (embedding), pre-step position, log(pre-energy), pre-step direction, material (embedding), layer ID. (n_sec / e_sec are outputs now, not inputs.)

Roadmap

Phase 1 (done): n_sec and total secondary energy e_sec were conditioning inputs; the model predicted only the 9D primary post-step (energy-conservation PoC).

Phase 2 (implemented — baseline): the two-stage model above predicts n_sec and each secondary's energy, direction, and species jointly with the primary, so a shower rollout is fully self-contained.

Next directions: faster-eval architectures against a ~10× native-Geant4 budget (Wasserstein-GAN, mixture-of-experts routing tree), a multi-material sampling-calorimeter dataset, and physical-property conditioning over learned embeddings.

Data

Input: parquet files produced by miniCaloSim, or converted from a ROOT file via uv run dwarf convert. Each row is one Geant4 step. Train/val split is by event_id (not row shuffle) to avoid leaking correlated steps from the same shower.

Project structure

giant/
├── giant/
│   ├── data/
│   │   ├── loader.py       # parquet → numpy arrays (incl. streaming/chunked reads)
│   │   ├── transforms.py   # log transforms, local-frame rotation, energy simplex, secondary encode/decode
│   │   └── dataset.py      # StepsDataset / StreamingStepsDataset (PyTorch)
│   ├── model/
│   │   ├── network.py      # SinusoidalEmbedding, ConditionEncoder, DenoisingMLP, SecondaryDecoder
│   │   └── schedule.py     # CosineSchedule (DDPM) and flow matching utilities
│   ├── constants.py        # output/conditioning dims, K_MAX, secondary slot layout, schema keys
│   ├── config.py           # default hyperparameters, TOML config merging, device autodetect
│   ├── pipeline.py         # builds datasets/normalizers and kicks off a training run
│   ├── train.py            # two-stage training loop, checkpointing, graceful shutdown
│   ├── sample.py           # DDPM / DDIM / flow matching samplers + secondary sampling
│   ├── geometry.py         # GeometryOracle: position → (material, layer_id, escaped) for rollout
│   ├── rollout.py          # autoregressive shower rollout driver
│   ├── validate.py         # step-level marginal + KL-divergence validation
│   ├── analysis.py         # step- and shower-level diagnostics: marginals, correlations, rollout observables
│   └── cli.py              # `giant train` / `predict` / `rollout` Typer app
├── scripts/                            # dataset/tooling logic, unified under the `dwarf` CLI (`uv run dwarf --help`)
│   ├── dwarf.py                        # Typer app: convert, migrate, bump-gen, bump-schema, status,
│   │                                    #   update-manifest, create-manifest, make-root, hparam-scan
│   ├── steps_to_parquet.py             # ROOT → parquet conversion (uproot/awkward/polars) — `dwarf convert`
│   ├── steps_to_parquet_parallel.py    # fan out conversion over several ROOT files — `dwarf convert --jobs N`
│   ├── migrate_geant_steps.py          # one-time move into the raw/processed/pools/derived layout — `dwarf migrate`
│   ├── bump_dataset_version.py         # cut a new raw gen or parquet schema, with a logged reason —
│   │                                    #   `dwarf bump-gen` / `bump-schema` / `status` / `update-manifest` / `create-manifest`
│   ├── create_root_files.py            # generate new ROOT shards via a minicalosim executable — `dwarf make-root`
│   ├── geometry_oracle.py              # fit a position → (material, layer_id) oracle — `dwarf build-geometry-oracle`
│   └── hparam_scan.py                  # hyperparameter grid scan over `giant train` runs — `dwarf hparam-scan`
└── tests/

Setup

uv sync --extra cpu               # CPU-only torch (use --extra cuda for CUDA 11.8 instead)
uv sync --extra cpu --extra dev   # add dev tools (pytest, ruff, ty)

cpu and cuda are mutually exclusive extras selecting the torch build; plain uv sync installs no torch at all. See CLAUDE.md for details.

Training, prediction, and rollout

giant train path/to/steps.parquet --mode flow
giant predict path/to/steps.parquet --checkpoint checkpoints/.../best.pt

# Full-shower rollout needs a geometry oracle (position → material/layer_id):
uv sync --extra cpu --extra geometry
dwarf build-geometry-oracle path/to/steps.parquet --out oracle.pkl
giant rollout path/to/steps.parquet --checkpoint checkpoints/.../best.pt --geometry oracle.pkl

train/predict accept a TOML config file (--config) and CLI overrides for hyperparameters; see --help on any command for the full option list. giant rollout seeds showers from the highest-energy entry step per event, then autoregressively steps the two-stage model to completion — pushing secondaries as new tracks and looking up material/layer_id from the oracle at each step. Tracks terminate on energy cutoff, per-track max steps, detector escape, or natural end; energy is deposited locally on every stop except escape (leakage), so showers conserve energy by construction.

Validation

giant.validate.validate_marginals runs step-level marginal and KL-divergence checks during training (--validate-every). For deeper diagnostics on a trained checkpoint — stratified marginals, correlation structure, physical-constraint violations, and shower-level rollout observables (longitudinal/transverse profiles, PDG energy shares) — see giant.analysis.

Development

uv run pytest            # run tests
uv run ruff check .      # lint
uv run ruff format .     # format
uv run ty check .        # type check
S
Description
GIANT — conditional generative surrogate for the Geant4 step function. Two-stage flow-matching model that replaces Geant4's stochastic shower physics, predicting post-step outcomes and secondary particles while conserving energy by construction.
Readme 14 MiB
Languages
Python 100%