lars 4ba419ebe4 Merge energy-conservation-poc into phase2-secondary-prediction
Brings the energy-conservation PoC work (dwarf CLI unification, dwarf
status improvements, predict --comment, ODE-step comparison scripts,
predict-parquet-only analysis refactor) onto the Phase 2 branch.

Conflict resolution:
- giant/analysis.py: took the energy-conservation-poc version wholesale.
  That branch deliberately removed the live checkpoint+sampler diagnostics
  path (ModelBundle/load_model_bundle/make_val_loader/collect_samples) in
  favor of reading `giant predict --coord local` parquet output. Phase 2's
  only edits to this file adapted the removed path to the new dataset API,
  so nothing Phase-2-specific is lost; no external code called those funcs.

Fixes for pre-existing breakage surfaced by the merge (both predate it):
- giant/cli.py: predict's `_process` unpacked build_features into 5 values,
  but Phase 2 made it return 8 (added n_sec/sec_cont/sec_pdg_idx). Expanded
  the unpack; `giant predict --coord local` would have crashed otherwise.
- tests/test_steps_to_parquet.py: Phase 2 renamed _add_secondary_energy ->
  _add_secondary_attributes without updating this test. Renamed the calls
  and extended the fixture with the pdg/pre_d{x,y,z} columns the expanded
  function reads; e_sec assertions unchanged.
- analysis/compare_ode_steps_energy_conservation.py: E731 lambda assignment
  (added in the un-linted final PoC commit) rewritten as a def.

ruff, ty, and pytest (179 passed) all green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 12:18:30 +02:00

giant

Geant4 Inference via Autoregressive Neural sTep surrogate — a play on Geant4 and the step function being the computationally heaviest part of the simulation.

Proof-of-concept surrogate model for the Geant4 step function. Given a pre-step particle state, the model samples a physically plausible post-step outcome — replacing the stochastic Geant4 physics engine with a trained conditional generative model.

Training is driven entirely from parquet files of the miniCaloSim steps tree. No Geant4 runtime dependency.

Architecture

Conditional flow matching model (Lipman et al. 2022): a small MLP learns a vector field mapping noise → step outcomes in ~10 ODE steps per sample. Falls back to DDPM for comparison.

Output space (9D, diffused):

Index Variable Transform
0 step_length [mm] log
1 ΔE = pre_E post_E [MeV] log
2 edep [MeV] log
35 post_dir in local frame unit vector
68 travel_dir (post_pos pre_pos) in local frame unit vector

Both post_dir and travel_dir are expressed in the coordinate frame where pre_dir = ẑ, making the scattering distribution nearly azimuthally symmetric. post_pos itself is not a raw target — it's reconstructed at inference as pre_pos + step_length * world_frame(travel_dir), so the two stay consistent by construction instead of being learned (and potentially diverging) independently.

Conditioning: PDG code (embedding), pre-step position, log(pre-energy), pre-step direction, material (embedding), layer ID, number of secondaries (Phase 1 only — see Roadmap below).

Roadmap

The model is developed in two phases:

Phase 1 (current): The number of secondaries produced in each step is passed as a conditioning input. This makes training easier because the model has direct access to multiplicity information and can focus on learning the continuous post-step kinematics.

Phase 2 (target): The number of secondaries is not given — the model must predict it jointly with all secondary properties (energy, direction, species) for each step. This requires extending the output space and likely an autoregressive or set-based generative approach for the variable-length secondary list.

Data

Input: parquet files produced by miniCaloSim, or converted from a ROOT file via uv run dwarf convert. Each row is one Geant4 step. Train/val split is by event_id (not row shuffle) to avoid leaking correlated steps from the same shower.

Project structure

giant/
├── giant/
│   ├── data/
│   │   ├── loader.py       # parquet → numpy arrays (incl. streaming/chunked reads)
│   │   ├── transforms.py   # log transforms, local-frame rotation, normaliser
│   │   └── dataset.py      # StepsDataset / StreamingStepsDataset (PyTorch)
│   ├── model/
│   │   ├── network.py      # SinusoidalEmbedding, ConditionEncoder, DenoisingMLP
│   │   └── schedule.py     # CosineSchedule (DDPM) and flow matching utilities
│   ├── config.py           # default hyperparameters, TOML config merging, device autodetect
│   ├── pipeline.py         # builds datasets/normalizers and kicks off a training run
│   ├── train.py            # training loop, checkpointing, graceful shutdown
│   ├── sample.py           # DDPM / DDIM / flow matching samplers
│   ├── validate.py         # step-level marginal + KL-divergence validation
│   ├── analysis.py         # notebook diagnostics: marginals, correlations, constraint checks
│   └── cli.py              # `giant train` / `giant predict` Typer app
├── scripts/                            # dataset/tooling logic, unified under the `dwarf` CLI (`uv run dwarf --help`)
│   ├── dwarf.py                        # Typer app: convert, migrate, bump-gen, bump-schema, status,
│   │                                    #   update-manifest, create-manifest, make-root, hparam-scan
│   ├── steps_to_parquet.py             # ROOT → parquet conversion (uproot/awkward/polars) — `dwarf convert`
│   ├── steps_to_parquet_parallel.py    # fan out conversion over several ROOT files — `dwarf convert --jobs N`
│   ├── migrate_geant_steps.py          # one-time move into the raw/processed/pools/derived layout — `dwarf migrate`
│   ├── bump_dataset_version.py         # cut a new raw gen or parquet schema, with a logged reason —
│   │                                    #   `dwarf bump-gen` / `bump-schema` / `status` / `update-manifest` / `create-manifest`
│   ├── create_root_files.py            # generate new ROOT shards via a minicalosim executable — `dwarf make-root`
│   └── hparam_scan.py                  # hyperparameter grid scan over `giant train` runs — `dwarf hparam-scan`
└── tests/

Setup

uv sync --extra cpu               # CPU-only torch (use --extra cuda for CUDA 11.8 instead)
uv sync --extra cpu --extra dev   # add dev tools (pytest, ruff, ty)

cpu and cuda are mutually exclusive extras selecting the torch build; plain uv sync installs no torch at all. See CLAUDE.md for details.

Training

giant train path/to/steps.parquet --mode flow
giant predict path/to/steps.parquet --checkpoint checkpoints/.../best.pt

Both accept a TOML config file (--config) and CLI overrides for hyperparameters; see --help on either for the full option list.

Validation

giant.validate.validate_marginals runs step-level marginal and KL-divergence checks during training (--validate-every). For deeper, notebook-driven diagnostics on a trained checkpoint — stratified marginals, correlation structure, physical constraint violations — see giant.analysis.

Development

uv run pytest            # run tests
uv run ruff check .      # lint
uv run ruff format .     # format
uv run ty check .        # type check
S
Description
GIANT — conditional generative surrogate for the Geant4 step function. Two-stage flow-matching model that replaces Geant4's stochastic shower physics, predicting post-step outcomes and secondary particles while conserving energy by construction.
Readme 14 MiB
Languages
Python 100%