1 Commits

Author SHA1 Message Date
lars bf3271f09e docs: rewrite README, keep version/test badges live via Gitea Actions
CI / Sync project version with tag (push) Has been skipped
CI / Lint (ruff check) (push) Successful in 1m6s
CI / Format (ruff format) (push) Successful in 1m6s
CI / Type check (ty) (push) Successful in 1m12s
CI / Tests (push) Successful in 3m15s
CI / Publish package to Gitea package registry (push) Has been skipped
CI / Bump version, tag, and update changelog on merge to master (push) Successful in 28s
CI / Update README badges (version, test count) (push) Successful in 1m21s
Rewrite README.md from scratch as a scannable landing page (hero, one
mermaid pipeline diagram, quick start, deep detail folded into
collapsible sections) instead of the old flat prose dump duplicating
CLAUDE.md.

Swap the static "CI" badge for a live Gitea Actions status badge, and
add an update-badges job to ci.yml that recomputes the version and
test-count badges on every push to master and pushes an update only
when they actually changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JL1hFbhLv5uwjTXqkWTLnH
2026-09-02 15:58:02 +02:00
2 changed files with 299 additions and 161 deletions
+46
View File
@@ -173,6 +173,52 @@ jobs:
git push origin "refs/tags/$TAG" git push origin "refs/tags/$TAG"
fi fi
update-badges:
name: Update README badges (version, test count)
needs: [ruff-check, ruff-format, type-check, test, bump-version]
if: github.ref == 'refs/heads/master' && github.event_name == 'push'
runs-on: ubuntu-latest
container:
image: docker.gitea.com/runner-images:ubuntu-latest
volumes:
- /srv/act-runner-cache/uv:/uv-cache
steps:
# ref: master (not the triggering SHA) so this picks up whatever
# bump-version just pushed, rather than badging the pre-bump commit.
- uses: actions/checkout@v4
with:
token: ${{ secrets.CI_TOKEN }}
ref: master
- uses: astral-sh/setup-uv@v5
with:
enable-cache: false
- run: |
echo "UV_CACHE_DIR=/uv-cache" >> "$GITHUB_ENV"
echo "UV_LINK_MODE=copy" >> "$GITHUB_ENV"
- run: uv sync --extra cpu --extra dev
- name: Compute version and test count
id: stats
run: |
VERSION=$(uv version --short)
TEST_COUNT=$(uv run pytest --collect-only -q 2>/dev/null | grep -oE '^[0-9]+ tests? collected' | grep -oE '^[0-9]+')
echo "version=$VERSION" >> "$GITHUB_OUTPUT"
echo "test_count=$TEST_COUNT" >> "$GITHUB_OUTPUT"
- name: Rewrite badge lines in README.md
run: |
sed -i -E "s|badge/version-[^-]+-informational|badge/version-${{ steps.stats.outputs.version }}-informational|" README.md
sed -i -E "s|badge/tests-[0-9]+%20passing-brightgreen|badge/tests-${{ steps.stats.outputs.test_count }}%20passing-brightgreen|" README.md
- name: Commit and push if the badges actually changed
run: |
git config user.name "gitea-actions"
git config user.email "actions@git.larsbogner.de"
git add README.md
if ! git diff --cached --quiet -- README.md; then
git commit -m "chore: update README badges (version ${{ steps.stats.outputs.version }}, ${{ steps.stats.outputs.test_count }} tests)"
git push origin HEAD:master
else
echo "Badges already up to date"
fi
sync-version-on-tag: sync-version-on-tag:
name: Sync project version with tag name: Sync project version with tag
if: startsWith(github.ref, 'refs/tags/') if: startsWith(github.ref, 'refs/tags/')
+253 -161
View File
@@ -1,203 +1,295 @@
# giant <div align="center">
**G**eant4 **I**nference via **A**utoregressive **N**eural s**T**ep surrogate. # GIANT
A conditional generative model that replaces the Geant4 step function: given a pre-step particle state it samples a post-step outcome — the primary's continuation plus its secondary particles — and autoregressively rolls that out into full showers. Trained entirely from parquet dumps of the miniCaloSim steps tree; no Geant4 runtime dependency. ### **G**eant4 **I**nference via **A**utoregressive **N**eural s**T**ep surrogate
A conditional generative model that replaces the Geant4 step function — sample a
post-step outcome instead of simulating one, then roll that out into full
calorimeter showers.
[![python](https://img.shields.io/badge/python-3.12%2B-3776AB?logo=python&logoColor=white)](pyproject.toml)
[![torch](https://img.shields.io/badge/torch-2.3.x-EE4C2C?logo=pytorch&logoColor=white)](pyproject.toml)
[![version](https://img.shields.io/badge/version-0.3.17-informational)](CHANGELOG.md)
[![tests](https://img.shields.io/badge/tests-1131%20passing-brightgreen)](tests/)
[![CI](https://git.larsbogner.de/lars/giant/actions/workflows/ci.yml/badge.svg?branch=master)](https://git.larsbogner.de/lars/giant/actions)
[![license](https://img.shields.io/badge/license-unlicensed-lightgrey)](#license)
</div>
---
## The idea
Geant4's step function is the innermost loop of detector simulation — for every
particle, at every step, it stochastically samples where the particle goes next,
how much energy it deposits, and what secondaries it spawns. GIANT learns that
function instead of running it: given a pre-step particle state (position,
energy, direction, particle species, material), a two-stage model samples a
post-step outcome — including the variable-length list of secondaries — and
autoregressively rolls that out into whole showers. It trains entirely from
parquet dumps of a Geant4 steps tree; nothing downstream needs a Geant4 runtime.
Two guarantees are architectural, not learned:
- **Energy is conserved by construction.** Stage 1 decodes deposit / secondary /
post-step energy through a softmax simplex that sums to the pre-step energy
exactly; Stage 2's secondaries stick-break that same energy budget.
- **No shower leaks across the train/val split.** Steps are split by `event_id`,
never by row, so correlated steps from the same shower can't appear on both
sides.
## How a step becomes a shower
```mermaid
flowchart LR
A["pre-step state\nposition · energy · direction\nspecies · material"] --> B["ConditionEncoder\nphysical / embedding / onehot"]
B --> C["Stage 1\n9D post-step outcome"]
C --> D["Stage 2 (autoregressive)\nsecondaries, descending energy"]
D --> E["rollout step"]
E -->|"primary continues"| F["GeometryOracle\nposition to material, layer"]
E -->|"secondaries pushed"| G["track queue"]
F --> A
G --> A
E -->|"terminated"| H["deposited shower"]
```
## Quick start ## Quick start
```bash ```bash
uv sync --extra cpu # install deps (CPU torch; use --extra cuda for GPU) uv sync --extra cpu # install deps (CPU torch; --extra cuda for GPU)
giant new-run --hidden-dim 512 --lr 3e-4 # scaffold config.toml + run dir
giant model summary --config config.toml # parameter counts + which config keys actually bite
giant train path/to/steps.parquet # train (flow + wgan by default)
giant predict path/to/steps.parquet --checkpoint checkpoints/.../best.pt
giant new-run --hidden-dim 512 --lr 3e-4 # scaffold config.toml + run dir
giant train path/to/steps.parquet # train (flow stage 1 + wgan stage 2, by default)
dwarf build-geometry-oracle path/to/steps.parquet --out oracle.pkl # needed for rollout dwarf build-geometry-oracle path/to/steps.parquet --out oracle.pkl # needed for rollout
giant rollout path/to/steps.parquet --checkpoint checkpoints/.../best.pt --geometry oracle.pkl giant rollout path/to/steps.parquet --checkpoint checkpoints/.../best.pt --geometry oracle.pkl
giant analyze prep rollout.yaml && giant analyze render <run_dir> --gallery # rollout-vs-Geant4 diagnostics
``` ```
Every command takes `--help` for the full flag list, and `--config config.toml` for anything not exposed as a flag. Every command takes `--help` for its full flag list, and `--config config.toml`
for anything not exposed as a flag.
## Architecture ---
A **two-stage model**, checkpointed together. Either stage's outcome can be produced by one of three interchangeable generative objectives (`--stage1-generator`/`--stage2-generator`, or `--mode` to set both at once): `flow` (conditional flow matching, ODE-sampled in ~10 steps), `ddpm` (denoising diffusion), or `wgan` (single-pass WGAN-GP generator/critic). <details>
<summary><h2 style="display:inline">Architecture</h2></summary>
**Stage 1 — primary step.** Predicts the 9D post-step outcome (`giant/constants.py:LOCAL_TARGET_NAMES`) from the pre-step conditioning: **Stage 1 — primary step.** Predicts the 9D post-step outcome
(`giant/constants.py:LOCAL_TARGET_NAMES`) from the pre-step conditioning:
| Index | Variable | Encoding | | Index | Variable | Encoding |
|-------|----------|----------| |-------|----------|----------|
| 0 | `step_length` [mm] | log | | 0 | `step_length` [mm] | log |
| 12 | `edep_logit`, `sec_logit` | ALR coords of the deposit/secondary/post-energy simplex | | 12 | `edep_logit`, `sec_logit` | ALR coordinates of the deposit / secondary / post-energy simplex |
| 35 | `post_dir` in local frame | unit vector | | 35 | `post_dir` | unit vector, local frame (`pre_dir = ẑ`) |
| 68 | `travel_dir` (`post_pos pre_pos`) in local frame | unit vector | | 68 | `travel_dir` (`post_pos pre_pos`) | unit vector, local frame |
- Energy logits decode via softmax over `[edep_logit, sec_logit, 0]` × `pre_E`, so `edep + e_sec + post_E == pre_E` exactly — conservation is architectural, not learned. Energy logits decode via `softmax([edep_logit, sec_logit, 0]) × pre_E`, so
- `post_dir`/`travel_dir` live in the frame where `pre_dir = ẑ`. `post_pos` isn't a target — it's reconstructed as `pre_pos + step_length * world_frame(travel_dir)`. `edep + e_sec + post_E == pre_E` holds exactly. `post_pos` is not itself a
target — it's reconstructed as `pre_pos + step_length · world_frame(travel_dir)`,
since duplicating that magnitude in a second target would let the two drift out
of sync.
**Stage 2 — secondaries.** Conditioned on the pre-step state and Stage 1's outcome, it generates the variable-length list of secondary particles. Two decoding strategies (`--stage2-decoder`): **Stage 2 — secondaries.** Conditioned on the pre-step state and Stage 1's
outcome, it generates the variable-length secondary list, one token at a time in
descending-energy order (`autoregressive`, default) or all `K_MAX` slots in one
masked pass (`one_shot`). Autoregressive tokens condition on a running history —
`markov` (previous token only) or `attention` (causal self-attention, KV-cached
at inference). Either way, secondary energies stick-break the `e_sec` budget
handed down from Stage 1, so the whole chain conserves energy. A secondary's
species is represented `onehot` (categorical, top-N PDG codes + "other"),
`physical` (continuous log-mass/charge), or `embedding` (nearest-neighbour
lookup).
- `autoregressive` — emits secondaries one at a time in descending-energy order, each token conditioned on a running history of prior tokens (`markov`: previous token only, or `attention`: causal self-attention, KV-cached at inference) **Conditioning (15D).** Pre-step position / energy / direction / layer, plus
- `one_shot` — all `K_MAX` slots generated in a single forward pass, masked past the predicted `n_sec` particle mass/charge and material Z_eff/A_eff/density/X0/λ_int — encoded the
same three ways as secondary species above, configured *independently* per axis
(`conditioning.particle.type` / `conditioning.material.type`). `physical`
computes rather than looks up, so it generalizes to species and materials
outside the training menu; that's the default. `n_sec`/`e_sec` are always model
outputs, never conditioning inputs.
Either way, secondary energies stick-break the `e_sec` budget handed down from Stage 1, so the full chain conserves energy. A secondary's particle identity is represented as `onehot` (categorical, top-N PDG codes + "other"), `physical` (continuous log-mass/charge), or `embedding` (nearest-neighbour lookup). **Composable by design** every stage assembles from small registries, so
swapping one axis doesn't touch the others:
**Conditioning.** Pre-step position/energy/direction/layer, plus particle mass/charge and material Z_eff/A_eff/density/X0/λ_int, encoded the same three ways as particle identity above. The particle and material axes are configured independently (`conditioning.particle.type` / `conditioning.material.type`; `--conditioning` sets both at once) and may mix — the `physical` representation generalizes to species/materials outside the training menu since it's computed rather than looked up. `n_sec`/`e_sec` are always model outputs, never conditioning inputs. | Registry | Choices |
|---|---|
| Objective | `flow` (matching, ~10-step ODE sample) · `ddpm` (denoising diffusion) · `wgan` (single-pass GAN) |
| Trunk | `resmlp` · `none`, optionally MoE-routed (`RoutedTrunk`) |
| Router | `energy` · `pdg` · `process` · `composed` · `none` — soft-mixed at train time, **top-1 dispatched at eval time**, which is the actual inference-speed win |
| History (stage 2 AR) | `markov` · `attention` · `none` |
**MoE routing** (`--router`): a pluggable `Router` (`energy`/`pdg`/`process`/`composed` axes) top-1-dispatches each row to one of several small expert trunks at eval time, instead of running one monolithic trunk. The CLI flags configure Stage 1's router; Stage 2 has its own `stage2_model.router` block, config-file only. </details>
## Data <details>
<summary><h2 style="display:inline">Configuration</h2></summary>
- Input: parquet files produced by [miniCaloSim](https://gitlab.etp.kit.edu/lbogner/minicalosim), or converted from ROOT via `dwarf convert`. One row = one Geant4 step. Every default lives in one place: frozen dataclasses in `giant/config.py`,
- **Conditioning (pre-step) columns:** `event_id`, `pdg`, `pre_x`/`pre_y`/`pre_z`, `pre_E`, `pre_dx`/`pre_dy`/`pre_dz` (direction), `material`, `layer_id`. composed into `GiantConfig` (`conditioning` / `stage1_model` / `stage2_model` /
- **Primary outcome (post-step) columns:** `post_x`/`post_y`/`post_z`, `post_E`, `post_dx`/`post_dy`/`post_dz`, `step_length`, `edep` (energy deposited in this step), `e_sec` (total energy carried off by secondaries), `child_track_ids` (its length gives `n_sec`). `train`). `DEFAULT_CONFIG` is *generated* from `GiantConfig().to_dict()` rather
- **Secondary columns**, one variable-length list per step: `sec_pdg_list`, `sec_E_list`, `sec_dx_list`/`sec_dy_list`/`sec_dz_list` — padded/truncated to `K_MAX` (15) slots on load, ordered by descending energy. than hand-maintained, so the dataclasses can't drift from what actually gets
- **Optional:** `process` — the physics process that produced the step (e.g. `compt`, `phot`, `eBrem`); a post-step label used only as classifier supervision (`ProcessRouter`), never as conditioning. merged. TOML config keys are validated against that shape — an unknown key is
- Train/val split is by `event_id` (`--seed`-controlled), not row shuffle, so correlated steps from the same shower never leak across the split. rejected with a did-you-mean suggestion. Precedence: CLI flag > `--config` file
- Loading a directory or `.manifest` of multiple parquet files (each one Geant4 job, `event_id` restarting from 0) offsets each file's `event_id`s by a fixed per-file stride so ids stay globally unique across files. > default.
## Project structure ```toml
# config.toml — resolved shape of the four blocks
[conditioning]
particle.type = "physical"
material.type = "physical"
``` [stage1_model]
giant/ generator = "flow"
├── giant/
│ ├── data/ [stage2_model]
│ │ ├── loader.py # parquet → numpy arrays (incl. streaming/chunked reads) generator = "wgan"
│ │ ├── transforms.py # log transforms, local-frame rotation, energy simplex, secondary encode/decode decoder = "autoregressive"
│ │ ├── dataset.py # StepsDataset / StreamingStepsDataset (PyTorch)
│ │ └── setup_cache.py # sidecar cache for the pre-epoch setup scan (vocab/split/normalizers) [train]
│ ├── model/ epochs = 100
│ │ ├── models.py # Stage1Model, Stage2OneShot, Stage2Autoregressive, CriticModel batch_size = 4096
│ │ ├── builders.py # build_models / build_critics — config dict → assembled stage models lr = 3e-4
│ │ ├── encoders.py # ConditionEncoder (physical / embedding / onehot, per axis)
│ │ ├── layers.py # ResBlock/AdaLNResBlock registry, SinusoidalEmbedding, MLP heads
│ │ ├── trunks.py # trunk registry (resmlp, none) + RoutedTrunk (MoE expert bodies)
│ │ ├── routers.py # Router registry: energy / pdg / process / composed / none
│ │ ├── history.py # stage-2 AR history encoders: markov / attention (KV-cached) / none
│ │ ├── objectives.py # flow / ddpm / wgan objective registry
│ │ ├── schedule.py # CosineSchedule (DDPM) and flow matching utilities
│ │ ├── wgan.py # WGAN-GP gradient penalty / critic / generator losses
│ │ ├── summary.py # build-only introspection behind `giant model summary`
│ │ ├── _legacy.py # v0.2 checkpoint model_config/state-dict migration
│ │ └── network.py # re-export shim over all of the above
│ ├── constants.py # output/conditioning dims, K_MAX, secondary slot layout, schema keys
│ ├── cond_layout.py # single source of truth for the cond_cont/cond_cat column layout
│ ├── particles.py # PDG → (mass, charge) decode, incl. nuclear/ion codes; onehot/embedding secondary-identity decode
│ ├── materials.py # material name → (Z_eff, A_eff, density, X0, λ_int)
│ ├── config.py # default hyperparameters, TOML config merging, device autodetect
│ ├── pipeline.py # builds datasets/normalizers and kicks off a training run (with setup-stage caching)
│ ├── training/ # two-stage training: loop, per-stage trainers, metrics, checkpointing
│ │ ├── loop.py # epoch loop, graceful shutdown, best-checkpoint selection
│ │ ├── trainers.py # StageSpec + flow/ddpm and WGAN-GP per-stage trainers
│ │ ├── stage2_inputs.py# ground-truth stage-2 targets + autoregressive/teacher-forcing inputs
│ │ ├── metrics.py # MetricsCollector: metrics.csv columns, W&B logging, progress/summary
│ │ ├── amp.py # bf16 autocast (`train.precision`)
│ │ ├── plots.py # training-progress plots (`giant analyze metrics`)
│ │ └── checkpoint.py # checkpoint assembly/restore (format unchanged since v0.2)
│ ├── sample.py # DDPM / DDIM / flow matching / WGAN samplers + secondary sampling
│ ├── checkpoint_io.py # checkpoint → ready-to-run models/normalizers (predict + rollout)
│ ├── geometry.py # GeometryOracle: position → (material, layer_id, escaped) for rollout
│ ├── rollout.py # autoregressive shower rollout driver
│ ├── validate.py # step-level marginal + KL-divergence validation
│ ├── _migration.py # shared v0.2 → v0.3 facts used by both migration surfaces
│ ├── analysis/ # rollout-vs-reference analysis pipeline (see `giant analyze` below)
│ │ ├── sources.py # canonical LazyFrames + secondary view
│ │ ├── variables.py # per-step value expressions shared by range sizing and the catalog
│ │ ├── reduce.py # streaming reduction primitives (hist1d, per-event scalars, profiles, ...)
│ │ ├── grouping.py # fixed bin edges + energy/pdg/material group sets
│ │ ├── context.py # resolves grouping into `shared.json` once per run
│ │ ├── reduced.py # Partial/Reduced — the compact JSON a compute job emits
│ │ ├── catalog.py # declarative PlotSpec registry (`giant analyze list`)
│ │ ├── router_gating.py / type_embedding_distance.py # checkpoint-bound diagnostics
│ │ ├── runtime_estimate.py # per-(plot, chunk) walltime estimates for submit
│ │ ├── condor.py # prep / compute-one / merge / submit-description plumbing
│ │ └── render.py # PDFs + HTML gallery (only module importing plotstyle/LaTeX)
│ └── cli.py # `giant train` / `new-run` / `model summary` / `predict` / `rollout` / `analyze`
├── giant/tools/ # dataset/tooling logic, unified under the `dwarf` CLI (`dwarf --help`)
│ ├── dwarf.py # Typer app: convert, migrate, bump-gen, bump-schema, status,
│ │ # update-manifest, create-manifest, make-root,
│ │ # build-geometry-oracle, warm-cache, hparam-scan
│ ├── steps_to_parquet.py # ROOT → parquet conversion (uproot/awkward/polars) — `dwarf convert`
│ ├── steps_to_parquet_parallel.py # fan out conversion over several ROOT files — `dwarf convert --jobs N`
│ ├── migrate_geant_steps.py # one-time move into the raw/processed/pools/derived layout — `dwarf migrate`
│ ├── bump_dataset_version.py # cut a new raw gen or parquet schema, with a logged reason —
│ │ # `dwarf bump-gen` / `bump-schema` / `status` / `update-manifest` / `create-manifest`
│ ├── create_root_files.py # generate new ROOT shards via a minicalosim executable — `dwarf make-root`
│ ├── geometry_oracle.py # fit a position → (material, layer_id) oracle — `dwarf build-geometry-oracle`
│ ├── warm_setup_cache.py # precompute `giant train`'s setup-stage sidecar — `dwarf warm-cache`
│ ├── hparam_scan.py # hyperparameter grid scan over `giant train` runs — `dwarf hparam-scan`
│ └── profile_analysis_costs.py # profiling helper for the `giant analyze` reduction pipeline
└── tests/
``` ```
## Setup Some knobs only exist in the config file, with no CLI flag:
`stage2_model.autoregressive.teacher_forcing`/`.history`,
`stage2_model.particle_type.target`/`.class_weighting`,
`stage2_model.n_sec.mode`/`.owner`, `conditioning.share_stages`,
`stage2_model.router.*`, and the finer `router` knobs (`lambda_balance`,
`gumbel`, `learn_width`, …).
`configs/` holds kept reference configs — `baseline.toml` is the fixed
comparison point every experimental variant (routed trunk, WGAN, attention
history, embedding conditioning) is a single edit away from. v0.2 flat-schema
configs and checkpoints load and auto-migrate.
</details>
<details>
<summary><h2 style="display:inline">Data</h2></summary>
Input is parquet — one row per Geant4 step — from
[miniCaloSim](https://gitlab.etp.kit.edu/lbogner/minicalosim), or converted from
ROOT via `dwarf convert`.
| Group | Columns |
|---|---|
| **Conditioning (pre-step)** | `event_id`, `pdg`, `pre_x`/`pre_y`/`pre_z`, `pre_E`, `pre_dx`/`pre_dy`/`pre_dz`, `material`, `layer_id` |
| **Primary outcome (post-step)** | `post_x`/`post_y`/`post_z`, `post_E`, `post_dx`/`post_dy`/`post_dz`, `step_length`, `edep`, `e_sec`, `child_track_ids` (length → `n_sec`) |
| **Secondaries** (variable-length lists) | `sec_pdg_list`, `sec_E_list`, `sec_dx_list`/`sec_dy_list`/`sec_dz_list` — padded/truncated to `K_MAX = 15` slots, descending energy |
| **Optional** | `process` — physics-process label, classifier supervision only (`ProcessRouter`), never conditioning |
Train/val split is by `event_id` (`--seed`-controlled), not row shuffle, so a
shower's correlated steps never straddle the split. Loading a directory or
`.manifest` of several parquet files offsets each file's `event_id`s by a
per-file stride so ids stay globally unique. The pre-epoch setup scan (vocab
maps, event split, normalizer stats) persists to a sidecar cache
(`--cache-setup`/`--rebuild-setup-cache`), precomputable ahead of time via
`dwarf warm-cache`.
</details>
<details>
<summary><h2 style="display:inline">CLI reference</h2></summary>
**`giant`** — train, run, and analyze the surrogate:
| Command | Does |
|---|---|
| `new-run` | scaffold a `config.toml` + run directory from flags |
| `train DATA` | train the two-stage model |
| `model summary` | build-only parameter counts, without training |
| `predict DATA --checkpoint …` | per-step predictions from a checkpoint |
| `rollout DATA --checkpoint … --geometry …` | full autoregressive shower rollout |
| `analyze prep/submit` | build a run dir; `submit` also queues HTCondor compute jobs |
| `analyze compute-one` / `merge-one` | one plot × chunk reduction / merge (what a condor job runs) |
| `analyze render <run_dir> --gallery` | merge chunks → styled PDFs + HTML gallery (local, needs LaTeX) |
| `analyze metrics <train_run_dir>` | training-progress plots from `metrics.csv` |
| `analyze list` | every catalog plot id |
**`dwarf`** — dataset/tooling CLI:
| Command | Does |
|---|---|
| `convert` | ROOT Steps tree → parquet (`--jobs N` fans out) |
| `migrate` | one-time move into the raw/processed/pools/derived layout |
| `bump-gen` / `bump-schema` / `status` | dataset versioning |
| `update-manifest` / `create-manifest` | point/build a manifest of parquet files |
| `make-root` | generate new ROOT shards via a minicalosim executable |
| `build-geometry-oracle` | fit position → (material, layer_id) for rollout |
| `warm-cache` | precompute `giant train`'s setup-stage sidecar |
| `hparam-scan` | grid-scan dropout × n_blocks × hidden_dim |
Worth knowing on `giant train` (full surface behind `--help`):
`--mode {flow,ddpm,wgan}` / `--stage1-generator` / `--stage2-generator`,
`--stage2-decoder {autoregressive,one_shot}`, `--conditioning
{physical,embedding,onehot}`, `--router` / `--router-type` / `--n-experts` /
`--router-axis`, `--stage{1,2}-init-from` + `--stage{1,2}-freeze` (retrain one
stage against a fixed other one), `--precision {fp32,bf16}`, `--wandb`.
</details>
<details>
<summary><h2 style="display:inline">Rollout &amp; analysis</h2></summary>
`giant rollout` seeds showers from each event's highest-energy entry step, then
autoregressively steps the model to completion — advancing all active tracks
breadth-first, batched — pushing secondaries as new tracks and looking up
`material`/`layer_id` from the geometry oracle each step. Tracks terminate on
one of six reasons (energy cutoff, max steps, detector escape, natural end,
unknown pdg, max tracks); every reason but escape deposits the remaining energy
locally, so showers conserve energy by construction — only `escaped` counts as
leakage.
`giant analyze` compares one or more rollouts against a single held-out
reference: `prep` resolves shared bin edges/groups once, `submit`/`compute-one`
run each (plot, `event_id`-disjoint chunk) pair as a polars/numpy-only HTCondor
job, `render` merges the chunks and produces the styled PDFs + HTML gallery
locally (the only step that needs LaTeX). Each rollout gets its own colored
series against one shared reference line. `giant analyze metrics` is a separate
entry point — training-progress plots straight from a run's `metrics.csv`.
</details>
<details>
<summary><h2 style="display:inline">Install</h2></summary>
| Extra | Adds | For |
|---|---|---|
| `cpu` **or** `cuda` | torch 2.3.x | required — mutually exclusive, pick one |
| `geometry` | scikit-learn | `dwarf build-geometry-oracle`, rollout |
| `analysis` | matplotlib, plotstyle | `giant analyze render` |
| `convert` | uproot, awkward | `dwarf convert` |
| `wandb` | wandb | `giant train --wandb` |
| `dev` | pytest, ruff, ty, + all of the above | development |
```bash ```bash
uv sync --extra cpu # CPU-only torch (use --extra cuda for CUDA 11.8 instead) uv sync --extra cpu --extra dev # everything needed to develop
uv sync --extra cpu --extra dev # add dev tools (pytest, ruff, ty)
uv sync --extra cpu --extra geometry # add scikit-learn, for `dwarf build-geometry-oracle` / rollout
uv sync --extra cpu --extra analysis # matplotlib/polars/plotstyle, for `giant analyze render`
uv sync --extra cpu --extra convert # uproot/awkward/polars, for `dwarf convert`
uv sync --extra cpu --extra wandb # W&B logging (`giant train --wandb`)
``` ```
The `dev` extra pulls in `convert`, `analysis`, `geometry` and `wandb` as well. Plain `uv sync` with no extra installs **no torch at all** — always include
`--extra cpu` or `--extra cuda`.
`cpu` and `cuda` are mutually exclusive — pick one to select the torch build (pinned to 2.3.x). Plain `uv sync` installs no torch at all. See `CLAUDE.md` for details. </details>
## Training, prediction, rollout <details>
<summary><h2 style="display:inline">Development</h2></summary>
```bash ```bash
giant new-run --hidden-dim 512 --lr 3e-4 --comment "..." # scaffold a config.toml + run dir uv run pytest # 964 tests
giant train path/to/steps.parquet # train (flow stage 1 + wgan stage 2, default)
giant predict path/to/steps.parquet --checkpoint checkpoints/.../best.pt
dwarf build-geometry-oracle path/to/steps.parquet --out oracle.pkl # position → material/layer_id
giant rollout path/to/steps.parquet --checkpoint checkpoints/.../best.pt --geometry oracle.pkl
```
Useful flags on `giant train`:
- `--mode {flow,ddpm,wgan}` sets both stages' objective at once; `--stage1-generator`/`--stage2-generator` override per stage
- `--stage2-decoder {autoregressive,one_shot}` — Stage 2 decoding strategy (see Architecture)
- `--conditioning {physical,embedding,onehot}` — conditioning representation
- `--router` / `--router-type` / `--n-experts` / `--router-axis` — MoE routing
- `--stage2-stage1-context {truth,sampled}` — feed Stage 2 the ground-truth or the model's own sampled Stage-1 outcome (annealable via `stage2_model.ctx_p_start`/`ctx_p_end`)
- `--precision {fp32,bf16}` — bf16 autocast in the training loop
- `--wandb` — log per-epoch metrics to Weights & Biases (needs `uv sync --extra wandb`); metric names are `<stage>/<split>/<metric>` plus an unprefixed run-level tail, all derived from `giant/training/trainers.py` `MetricSpec`s
- `--no-cache-setup` / `--rebuild-setup-cache` — control the setup-stage sidecar cache (vocab maps, event split, normalizer stats); `dwarf warm-cache` precomputes it
- `--stage1-init-from`/`--stage2-init-from` (checkpoint `.pt`) + `--stage1-freeze`/`--stage2-freeze` — load a stage's weights from another checkpoint and never update them, so the other stage can be retrained alone against a fixed, known-good one while still producing a complete, rollout-capable checkpoint
Config-file-only knobs (no CLI flag — use `--config config.toml`): `stage2_model.autoregressive.teacher_forcing`/`.history`, `stage2_model.particle_type.target`/`.class_weighting`, `stage2_model.n_sec.mode`/`.owner`, `conditioning.share_stages`, `stage*_model.trunk.*` and the finer `router` knobs (`lambda_balance`, `gumbel`, `learn_width`, …). `configs/` holds kept reference configs. v0.2 flat-schema configs and checkpoints load fine (auto-migrated).
`giant rollout` seeds showers from each event's highest-energy entry step, then autoregressively steps the model to completion, pushing secondaries as new tracks and looking up `material`/`layer_id` from the geometry oracle each step. Tracks terminate on energy cutoff, max steps, detector escape, or natural end; energy is deposited locally on every stop except escape, so showers conserve energy by construction.
## Validation and analysis
- `giant.validate.validate_marginals` — step-level marginal + KL-divergence checks during training (`--validate-every`)
- `giant analyze` — deeper rollout-vs-reference diagnostics (marginals by energy/pdg/material, per-event totals, shower profiles, species share, leakage, secondaries):
```bash
giant analyze submit rollout.yaml --accounting-group cms # prep + one HTCondor job per plot × chunk (compute only)
giant analyze submit a.yaml b.yaml --accounting-group cms --label flow --label wgan # N rollouts vs one shared reference
giant analyze render <run_dir> --gallery # local: merge chunks, then styled PDFs + HTML gallery (needs LaTeX)
giant analyze list # every catalog plot id
giant analyze prep rollout.yaml --chunks 8 # just the run directory, no submission
giant analyze compute-one --id marginal_edep --run-dir <run_dir> --chunk 0 # what a condor job runs
giant analyze merge-one --id marginal_edep --run-dir <run_dir> # merge one plot's chunks (debugging)
```
`<run_dir>` defaults to `<cwd>/analysis_runs/analysis_<id>` (`--run-dir` overrides it; `prep`/`submit` print it). Multiple rollout YAMLs must all name the same reference (`dataset`) file; each renders as its own colored series against one reference line/panel. Compute jobs are polars/numpy only; only `render` needs LaTeX, so it always runs locally.
Separately, `giant analyze metrics <train_run_dir>` renders training-progress plots (loss/lr/accuracy/grad-norm/router/wgan/throughput) straight from a training run's `metrics.csv`.
## Development
```bash
uv run pytest # run tests
uv run ruff check . # lint uv run ruff check . # lint
uv run ruff format . # format uv run ruff format . # format
uv run ty check . # type check uv run ty check . # type check
``` ```
Gitea Actions (`.gitea/workflows/ci.yml`) runs lint + format-check + type-check
+ tests on every push and PR; merges to `master` auto-bump the patch version
and regenerate `CHANGELOG.md` — don't hand-edit either.
</details>
## License
Not yet decided — treat this repository as all-rights-reserved until a
`LICENSE` file is added.