12 Commits

Author SHA1 Message Date
gitea-actions ff204732d7 chore: update changelog for v0.3.7 [skip ci] 2026-08-24 09:31:26 +00:00
gitea-actions 02ed4e531c chore: bump version 0.3.6 -> 0.3.7 [skip ci] 2026-08-24 09:31:25 +00:00
lars 1b6c8b33b7 Merge pull request 'Add rollout-quality distance, confusion, containment and router plots (gitea #76)' (#79) from fix/issue-76 into master
CI / Lint (ruff check) (push) Successful in 31s
CI / Format (ruff format) (push) Successful in 34s
CI / Type check (ty) (push) Successful in 37s
CI / Sync project version with tag (push) Has been skipped
CI / Tests (push) Successful in 2m45s
CI / Bump version, tag, and update changelog on merge to master (push) Successful in 42s
Reviewed-on: #79
2026-08-24 11:22:11 +02:00
lars ffb7c0cc2a Add rollout-quality distance, confusion, containment and router plots (gitea #76)
CI / Format (ruff format) (push) Successful in 30s
CI / Lint (ruff check) (push) Successful in 30s
CI / Sync project version with tag (push) Has been skipped
CI / Type check (ty) (push) Successful in 32s
CI / Lint (ruff check) (pull_request) Successful in 33s
CI / Format (ruff format) (pull_request) Successful in 31s
CI / Type check (ty) (pull_request) Successful in 35s
CI / Sync project version with tag (pull_request) Has been skipped
CI / Tests (push) Successful in 5m51s
CI / Tests (pull_request) Successful in 5m5s
CI / Bump version, tag, and update changelog on merge to master (push) Has been skipped
CI / Bump version, tag, and update changelog on merge to master (pull_request) Has been skipped
Picks 4 of the 7 catalog additions the issue proposed (the smaller-lift
ones; 2D joint plots, PIT calibration, and the throughput/accuracy scatter
are left for follow-up issues):

- marginal_distance_summary: a var x grouping-axis KS-statistic heatmap,
  reusing the existing marginal hist1d compute and just adding a finalize —
  a single at-a-glance regression scorecard instead of N overlay plots.
- n_sec_confusion: predicted (rollout) vs true (reference) secondary count
  per event, paired by event_id since a rollout is seeded from the same
  events as its reference file. Needed a new zero-filling primitive
  (reduce.sec_count_by_event) since a plain group_by over secondary rows
  silently drops zero-secondary events.
- shower_containment_depth_{90,95}: per-event depth containing 90%/95% of
  deposited energy, derived from the same per-event depth-bin matrix the
  longitudinal profile already computes.
- router_specialization: max gate weight vs energy per side, summarizing
  router_gating's full stacked area into the one trend line the roadmap's
  MoE writeup describes (the ~60-65% ceiling), to make a future
  lambda_balance>0 retrain's effect on specialization checkable at a glance.

Both new heatmap-shaped plots (distance summary, confusion matrix) share one
new "heatmap" Reduced kind/renderer rather than two near-identical ones.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-24 11:12:16 +02:00
gitea-actions 97f5bbf9f0 chore: update changelog for v0.3.6 [skip ci] 2026-08-24 08:02:49 +00:00
gitea-actions 060353ea4a chore: bump version 0.3.5 -> 0.3.6 [skip ci] 2026-08-24 08:02:48 +00:00
lars b3f28e98af Merge pull request 'Give CriticModel a registry-built trunk and StageModel base (gitea #57)' (#74) from fix/issue-57 into master
CI / Format (ruff format) (push) Successful in 31s
CI / Lint (ruff check) (push) Successful in 31s
CI / Sync project version with tag (push) Has been skipped
CI / Type check (ty) (push) Successful in 33s
CI / Tests (push) Successful in 2m44s
CI / Bump version, tag, and update changelog on merge to master (push) Successful in 41s
Reviewed-on: #74
2026-08-24 09:57:58 +02:00
lars 4b2e0ba98e Give CriticModel a registry-built trunk and StageModel base (gitea #57)
CI / Format (ruff format) (push) Successful in 33s
CI / Lint (ruff check) (push) Successful in 36s
CI / Sync project version with tag (push) Has been skipped
CI / Type check (ty) (push) Successful in 27s
CI / Lint (ruff check) (pull_request) Successful in 27s
CI / Format (ruff format) (pull_request) Successful in 29s
CI / Tests (push) Successful in 3m33s
CI / Sync project version with tag (pull_request) Has been skipped
CI / Bump version, tag, and update changelog on merge to master (push) Has been skipped
CI / Type check (ty) (pull_request) Successful in 32s
CI / Tests (pull_request) Successful in 2m45s
CI / Bump version, tag, and update changelog on merge to master (pull_request) Has been skipped
CriticModel was the one stage-shaped class left out of the trunk-registry
(gitea #33), block-conditioning-registry (gitea #34), and StageModel-base
(gitea #39) refactors: it hand-rolled a plain ResBlock stack, so a
routed/FiLM/AdaLN trunk was available to every generative stage model except
the critic competing against them under WGAN-GP.

CriticModel now subclasses StageModel (reusing its cond_enc construction, and
a stage-2 context-fusion helper factored out of Stage2OneShot onto the base)
and builds its body via build_trunk (output width 1) instead of a bespoke
ResBlock loop, so trunk.type/trunk.block_conditioning now affect the critic
too. Each stage's critic inherits its own generator's trunk config rather
than a new critic_trunk config key, mirroring the existing
critic_hidden_dim/critic_n_res_blocks "0 = inherit from generator" pattern.
Router mixing (MoE) for the critic stays out of scope. Since CriticModel is
training-only and never persisted for inference, and WGAN-GP is still
unbenchmarked, its state_dict shape has no back-compat burden.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-24 09:48:26 +02:00
gitea-actions 1675052ecd chore: update changelog for v0.3.5 [skip ci] 2026-08-24 07:37:23 +00:00
gitea-actions eb9d331bea chore: bump version 0.3.4 -> 0.3.5 [skip ci] 2026-08-24 07:37:22 +00:00
lars 12689cf5b6 Merge pull request 'Add "none" variants for router, history, and trunk (gitea #45)' (#73) from fix/issue-45 into master
CI / Lint (ruff check) (push) Successful in 30s
CI / Format (ruff format) (push) Successful in 29s
CI / Sync project version with tag (push) Has been skipped
CI / Type check (ty) (push) Successful in 34s
CI / Tests (push) Successful in 2m47s
CI / Bump version, tag, and update changelog on merge to master (push) Successful in 42s
Reviewed-on: #73
2026-08-24 09:32:30 +02:00
lars 732d5f1cd2 Add "none" variants for router, history, and trunk (gitea #45)
CI / Format (ruff format) (push) Successful in 30s
CI / Lint (ruff check) (push) Successful in 34s
CI / Sync project version with tag (push) Has been skipped
CI / Type check (ty) (push) Successful in 42s
CI / Lint (ruff check) (pull_request) Successful in 41s
CI / Format (ruff format) (pull_request) Successful in 46s
CI / Sync project version with tag (pull_request) Has been skipped
CI / Type check (ty) (pull_request) Successful in 43s
CI / Tests (push) Successful in 5m12s
CI / Tests (pull_request) Successful in 5m10s
CI / Bump version, tag, and update changelog on merge to master (push) Has been skipped
CI / Bump version, tag, and update changelog on merge to master (pull_request) Has been skipped
Turns "is this component earning its parameters?" into a one-line
config flip for each of the three pluggable network components:

- router.type = "none" (NoneRouter, giant/model/routers.py): still
  builds n_experts expert trunks via RoutedTrunk, but replaces the
  learned gate with a uniform 1/n_experts weight for every row — no
  centers/embeddings/classifier. Distinct from router.enabled=false
  (which drops routing/mixing entirely): this isolates whether the
  *learned routing signal* specifically is earning its parameters,
  holding expert count fixed.

- stage2_model.autoregressive.history = "none" (NoHistory,
  giant/model/history.py): ignores feat/has_prev entirely and always
  returns zeros, ablating whether the AR decoder's history
  conditioning earns its parameters. Already validated for free by
  gitea #35's generic HISTORY_REGISTRY membership check.

- trunk.type = "linear" (LinearTrunk, giant/model/trunks.py): a bare
  nn.Linear(in_dim + cond_dim, out_dim) body, no ResBlock stack. Per
  gitea #33's design, this composes for free with router.enabled=true
  ("mixture of trivial linear experts").

Both blocking issues (#33 trunk registry, #35 pluggable history
encoder) are closed, so this was unblocked.
2026-08-24 09:22:36 +02:00
21 changed files with 776 additions and 55 deletions
+1 -1
View File
@@ -1,5 +1,5 @@
[tool.bumpversion]
current_version = "0.3.4"
current_version = "0.3.7"
parse = "(?P<major>\\d+)\\.(?P<minor>\\d+)\\.(?P<patch>\\d+)"
serialize = ["{major}.{minor}.{patch}"]
search = "{current_version}"
+18
View File
@@ -1,5 +1,23 @@
# Changelog
## [0.3.7] - 2026-08-24
### Added
- Add rollout-quality distance, confusion, containment and router plots [gitea #76](https://git.larsbogner.de/lars/giant/issues/76)
## [0.3.6] - 2026-08-24
### Changed
- Give CriticModel a registry-built trunk and StageModel base [gitea #57](https://git.larsbogner.de/lars/giant/issues/57)
## [0.3.5] - 2026-08-24
### Added
- Add "none" variants for router, history, and trunk [gitea #45](https://git.larsbogner.de/lars/giant/issues/45)
## [0.3.4] - 2026-08-23
### Added
+230
View File
@@ -47,6 +47,7 @@ from giant.analysis.reduce import (
leakage_fraction,
profile_finalize,
profile_partial,
sec_count_by_event,
species_share,
sum_merge,
transverse_expr,
@@ -56,6 +57,7 @@ from giant.analysis.router_gating import (
compute_router_gating,
compute_router_share_by_pdg,
compute_router_share_by_process,
compute_router_specialization,
)
from giant.analysis.sources import Side, open_side, physical_steps, secondaries
from giant.analysis.type_embedding_distance import compute_type_embedding_l1_distance
@@ -176,6 +178,68 @@ def _np_hist_pair(r: np.ndarray, t: np.ndarray, nbins: int) -> tuple[np.ndarray,
return edges, np.histogram(r, edges)[0], np.histogram(t, edges)[0]
def _ks_statistic(r_counts, t_counts) -> float:
"""KS statistic (max |CDF diff|) between two same-edge binned histograms.
``nan`` when neither side has any mass (nothing to compare); 1.0 (maximal
mismatch) when exactly one side is entirely empty and the other isn't —
correctly the worst score rather than an undefined one.
"""
r_counts = np.asarray(r_counts, dtype=np.float64)
t_counts = np.asarray(t_counts, dtype=np.float64)
r_tot, t_tot = r_counts.sum(), t_counts.sum()
if r_tot == 0 and t_tot == 0:
return float("nan")
if r_tot == 0 or t_tot == 0:
return 1.0
r_cdf = np.cumsum(r_counts) / r_tot
t_cdf = np.cumsum(t_counts) / t_tot
return float(np.max(np.abs(r_cdf - t_cdf)))
def _integer_confusion(t: np.ndarray, r: np.ndarray, max_bins: int = 21) -> tuple[list[str], np.ndarray]:
"""Confusion matrix of two paired small-integer arrays (e.g. secondary counts).
Bins are consecutive integers ``0..cap``, with the last bin an overflow
``"cap+"`` bucket, so an occasional pathological count doesn't blow up the
heatmap. Returns ``(labels, matrix)`` with ``matrix[i, j]`` counting pairs
with ``t == i`` and ``r == j`` (both clipped into ``[0, cap]``).
"""
cap = min(max(int(t.max()) if len(t) else 0, int(r.max()) if len(r) else 0, 1), max_bins - 1)
t_c = np.clip(t.astype(np.int64), 0, cap)
r_c = np.clip(r.astype(np.int64), 0, cap)
n = cap + 1
mat = np.zeros((n, n), dtype=np.int64)
np.add.at(mat, (t_c, r_c), 1)
labels = [str(i) for i in range(cap)] + [f"{cap}+"]
return labels, mat
def _containment_depths(mat: np.ndarray, edges: np.ndarray, quantile: float) -> np.ndarray:
"""Per-event depth containing ``quantile`` of that event's deposited energy.
``mat`` is a ``(n_events, n_bins)`` edep-per-depth-bin sum matrix (see
``reduce.profile_partial``); bins are ordered by increasing depth (matching
``edges``, monotonic). Zero-energy events are dropped — containment depth is
undefined for them.
"""
totals = mat.sum(axis=1)
valid = totals > 0
mat, totals = mat[valid], totals[valid]
cum = np.cumsum(mat, axis=1) / totals[:, None]
idx = (cum >= quantile).argmax(axis=1) # first bin whose cumulative fraction reaches quantile
return edges[1:][idx]
def _group_keys(ctx: Context, axis: str) -> list:
"""The group keys ``_marginal_grouped_finalize`` iterates for ``axis``."""
if axis == "pdg":
return list(ctx.top_pdgs)
if axis == "material":
return list(ctx.materials)
return list(range(len(ctx.energy_edges) - 1)) # energy
# Human-readable figure titles per marginal variable (the axis labels carry units;
# these read cleanly as a title without them).
_TITLE_NAMES = {
@@ -299,6 +363,64 @@ def _marginal_grouped_finalize(parts: list[dict], ctx: Context, var: str, axis:
)
# ---------------------------------------------------------------------------
# distance summary: a var x group-axis scorecard, reusing the marginal hists
# ---------------------------------------------------------------------------
def _distance_summary_partial(b: Bundle) -> dict:
out: dict[str, dict] = {}
for var in MARGINAL_VARS:
out[var] = {"overall": _marginal_overall_partial(b, var)}
for axis in GROUPING_AXES:
out[var][axis] = _marginal_grouped_partial(b, var, axis)
return out
def _distance_summary_finalize(parts: list[dict], ctx: Context) -> Reduced:
col_labels = ["overall", *GROUPING_AXES]
matrix: list[list[float]] = []
for var in MARGINAL_VARS:
edges = _marginal_edges(ctx, var)
nb = len(edges) - 1
row: list[float] = []
r = sum_merge([p[var]["overall"]["r"] for p in parts])
t = sum_merge([p[var]["overall"]["t"] for p in parts])
row.append(_ks_statistic(_finalize_counts(r, 0, nb), _finalize_counts(t, 0, nb)))
for axis in GROUPING_AXES:
r = sum_merge([p[var][axis]["r"] for p in parts])
t = sum_merge([p[var][axis]["t"] for p in parts])
dists, weights = [], []
for k in _group_keys(ctx, axis):
rc, tc = _finalize_counts(r, k, nb), _finalize_counts(t, k, nb)
w = sum(rc) + sum(tc)
if w == 0:
continue
dists.append(_ks_statistic(rc, tc))
weights.append(w)
row.append(float(np.average(dists, weights=weights)) if dists else float("nan"))
matrix.append(row)
return Reduced(
id="marginal_distance_summary",
family="quality",
kind="heatmap",
title="Marginal distance summary (KS statistic, rollout vs reference)",
xlabel="grouping axis",
payload={
"matrix": matrix,
"row_labels": [_TITLE_NAMES[v] for v in MARGINAL_VARS],
"col_labels": col_labels,
"ylabel": "marginal variable",
"cbar_label": "KS statistic (0 = identical, 1 = maximal mismatch)",
"vmin": 0.0,
"vmax": 1.0,
},
)
# ---------------------------------------------------------------------------
# per-event scalar observables
# ---------------------------------------------------------------------------
@@ -438,6 +560,41 @@ def _profile_finalize(
)
# ---------------------------------------------------------------------------
# shower containment depth (reuses the longitudinal profile's per-event matrix)
# ---------------------------------------------------------------------------
_CONTAINMENT_QUANTILES: list[tuple[float, str]] = [
(0.90, "shower_containment_depth_90"),
(0.95, "shower_containment_depth_95"),
]
def _containment_finalize(parts: list[dict], ctx: Context, spec_id: str, quantile: float) -> Reduced:
edges = np.asarray(ctx.depth_edges)
nb = len(edges) - 1
_assert_event_disjoint([p["r_ids"] for p in parts], spec_id, "rollout")
_assert_event_disjoint([p["t_ids"] for p in parts], spec_id, "reference")
r_full = np.concatenate([np.asarray(p["r_mat"], dtype=float).reshape(-1, nb) for p in parts], axis=0)
t_full = np.concatenate([np.asarray(p["t_mat"], dtype=float).reshape(-1, nb) for p in parts], axis=0)
r_depth = _containment_depths(r_full, edges, quantile)
t_depth = _containment_depths(t_full, edges, quantile)
hedges, rc, tc = _np_hist_pair(r_depth, t_depth, ctx.n_marginal_bins)
return Reduced(
id=spec_id,
family="shower",
kind="overlay_hist",
title=f"Shower containment depth ({quantile:.0%} of deposited energy)",
xlabel=f"depth containing {quantile:.0%} of deposited energy [mm]",
payload={
"edges": hedges.tolist(),
_ROLL: rc.astype(np.int64).tolist(),
_REF: tc.astype(np.int64).tolist(),
"log_y": False,
},
)
# ---------------------------------------------------------------------------
# species share + leakage
# ---------------------------------------------------------------------------
@@ -627,6 +784,43 @@ def _sec_cos_angle_finalize(parts: list[dict], ctx: Context) -> Reduced:
)
def _n_sec_confusion_partial(b: Bundle) -> dict:
r_sec, t_sec = _sec_frames(b)
r_ids, r_n = sec_count_by_event(b.r_phys, r_sec)
t_ids, t_n = sec_count_by_event(b.t_all, t_sec)
return {"r_ids": r_ids.tolist(), "r_n": r_n.tolist(), "t_ids": t_ids.tolist(), "t_n": t_n.tolist()}
def _n_sec_confusion_finalize(parts: list[dict], ctx: Context) -> Reduced:
r_ids = np.concatenate([np.asarray(p["r_ids"], dtype=np.int64) for p in parts])
r_n = np.concatenate([np.asarray(p["r_n"], dtype=np.int64) for p in parts])
t_ids = np.concatenate([np.asarray(p["t_ids"], dtype=np.int64) for p in parts])
t_n = np.concatenate([np.asarray(p["t_n"], dtype=np.int64) for p in parts])
# event-disjoint chunking (see Bundle.open) means each event_id appears in
# exactly one part on each side, so a plain dict build is a safe merge.
r_map = dict(zip(r_ids.tolist(), r_n.tolist()))
t_map = dict(zip(t_ids.tolist(), t_n.tolist()))
common = sorted(set(r_map) & set(t_map))
true_n = np.array([t_map[e] for e in common], dtype=np.int64)
pred_n = np.array([r_map[e] for e in common], dtype=np.int64)
labels, mat = _integer_confusion(true_n, pred_n)
return Reduced(
id="n_sec_confusion",
family="secondaries",
kind="heatmap",
title="Predicted vs true secondary count per event",
xlabel="predicted secondaries (rollout)",
payload={
"matrix": mat.tolist(),
"row_labels": labels,
"col_labels": labels,
"ylabel": "true secondaries (reference)",
"cbar_label": "event count",
"vmin": 0.0,
},
)
# ---------------------------------------------------------------------------
# router diagnostics (not chunked — already bounded/subsampled)
# ---------------------------------------------------------------------------
@@ -640,6 +834,9 @@ _router_share_pdg_partial, _router_share_pdg_finalize = _unchunkable(
_router_share_process_partial, _router_share_process_finalize = _unchunkable(
lambda b: compute_router_share_by_process(b.checkpoint, b.t_phys)
)
_router_specialization_partial, _router_specialization_finalize = _unchunkable(
lambda b: compute_router_specialization(b.checkpoint, b.r_phys, b.t_phys)
)
_type_embedding_l1_distance_partial, _type_embedding_l1_distance_finalize = _unchunkable(
lambda b: compute_type_embedding_l1_distance(b.type_embedding_l1_dist)
)
@@ -676,6 +873,15 @@ def build_catalog() -> list[PlotSpec]:
)
)
specs.append(
PlotSpec(
"marginal_distance_summary",
"quality",
compute_partial=_distance_summary_partial,
finalize=_distance_summary_finalize,
)
)
specs += [
PlotSpec(
"event_total_edep",
@@ -745,6 +951,17 @@ def build_catalog() -> list[PlotSpec]:
"transverse_edges",
),
),
]
for quantile, spec_id in _CONTAINMENT_QUANTILES:
specs.append(
PlotSpec(
spec_id,
"shower",
compute_partial=lambda b: _profile_partial(b, depth_expr, "depth_edges"),
finalize=lambda parts, ctx, q=quantile, sid=spec_id: _containment_finalize(parts, ctx, sid, q),
)
)
specs += [
PlotSpec(
"species_edep_share",
"species",
@@ -781,6 +998,12 @@ def build_catalog() -> list[PlotSpec]:
compute_partial=_sec_cos_angle_partial,
finalize=_sec_cos_angle_finalize,
),
PlotSpec(
"n_sec_confusion",
"secondaries",
compute_partial=_n_sec_confusion_partial,
finalize=_n_sec_confusion_finalize,
),
PlotSpec(
"router_gating",
"model",
@@ -802,6 +1025,13 @@ def build_catalog() -> list[PlotSpec]:
finalize=_router_share_process_finalize,
chunkable=False,
),
PlotSpec(
"router_specialization",
"model",
compute_partial=_router_specialization_partial,
finalize=_router_specialization_finalize,
chunkable=False,
),
PlotSpec(
"type_embedding_l1_distance",
"model",
+17
View File
@@ -271,3 +271,20 @@ def leakage_fraction(lf: pl.LazyFrame) -> np.ndarray:
escaped = per_event["escaped"].fill_null(0.0).to_numpy()
total = deposited + escaped
return np.where(total > 0, escaped / total, 0.0)
def sec_count_by_event(lf_all: pl.LazyFrame, sec_lf: pl.LazyFrame) -> tuple[np.ndarray, np.ndarray]:
"""Per-event secondary count, zero-filled for events that produced none.
Two bounded per-event ``group_by``s — the full event set (from ``lf_all``)
and the secondary counts (from ``sec_lf``, see ``sources.secondaries``) —
merged in Python via a dict. Both results are event-granularity (not
per-row), so this stays in the same bounded-memory budget as
``event_scalars``; a plain ``group_by`` on ``sec_lf`` alone would silently
drop zero-secondary events instead of zero-filling them.
"""
ev = lf_all.select("event_id").unique().collect(engine="streaming")["event_id"].to_numpy()
cnt_df = sec_lf.group_by("event_id").agg(pl.len().alias("n")).collect(engine="streaming")
cnt = dict(zip(cnt_df["event_id"].to_list(), cnt_df["n"].to_list()))
counts = np.array([cnt.get(int(e), 0) for e in ev], dtype=np.int64)
return ev, counts
+4
View File
@@ -19,6 +19,10 @@ from pathlib import Path
# "single_hist" one series only (e.g. rollout leakage; reference has none)
# "router_gating" stacked mean MoE gate weight vs energy, rollout + reference
# "router_share" stacked bar of MoE top-1 dispatch share by category
# "router_specialization" max gate weight vs energy, rollout + reference (one
# scalar trend line summarizing "router_gating")
# "heatmap" row x col matrix + colorbar (distance scorecard or a
# predicted-vs-true confusion matrix)
# "unavailable" plot not applicable to this run (e.g. non-MoE checkpoint)
+43
View File
@@ -260,6 +260,47 @@ def _render_router_share(r: Reduced, params: dict):
return fig
def _render_router_specialization(r: Reduced, params: dict):
fig, ax = ps.new_figure("thesis-single", title=r.title, params=params)
for key in ("reference", "rollout"):
side = r.payload.get(key)
if side and side["centers"]:
ax.plot(side["centers"], side["score"], label=_SERIES_LABELS[key], marker="o", markersize=3)
chance = r.payload.get("chance_level")
if chance is not None:
ax.axhline(chance, linestyle="--", color="gray", label="chance level (1/n_experts)")
if r.payload.get("log_x"):
ax.set_xscale("log")
ax.set_ylim(0, 1)
ax.set_xlabel(r.xlabel)
ax.set_ylabel("max gate weight")
ps.style_legend(ax, title=f"{r.payload.get('router_type', '')} router")
return fig
def _render_heatmap(r: Reduced, params: dict):
mat = np.asarray(r.payload["matrix"], dtype=float)
row_labels = r.payload["row_labels"]
col_labels = r.payload["col_labels"]
fig, ax = ps.new_figure("thesis-single", title=r.title, params=params)
im = ax.imshow(
mat,
origin="upper",
aspect="auto",
cmap=r.payload.get("cmap", "viridis"),
vmin=r.payload.get("vmin"),
vmax=r.payload.get("vmax"),
)
ax.set_xticks(range(len(col_labels)))
ax.set_xticklabels(col_labels, rotation=45, ha="right")
ax.set_yticks(range(len(row_labels)))
ax.set_yticklabels(row_labels)
ax.set_xlabel(r.xlabel)
ax.set_ylabel(r.payload.get("ylabel", ""))
fig.colorbar(im, ax=ax, label=r.payload.get("cbar_label", "value"))
return fig
def _render_unavailable(r: Reduced, params: dict):
fig, ax = ps.new_figure("thesis-single", title=r.title, params=params)
ax.axis("off")
@@ -284,6 +325,8 @@ _RENDERERS = {
"bar": _render_bar,
"router_gating": _render_router_gating,
"router_share": _render_router_share,
"router_specialization": _render_router_specialization,
"heatmap": _render_heatmap,
"unavailable": _render_unavailable,
}
+49
View File
@@ -203,6 +203,7 @@ _TITLES = {
"router_gating": "Router gating (mixture-of-experts decision boundaries)",
"router_share_by_pdg": "Router expert share by particle species",
"router_share_by_process": "Router expert share by physics process",
"router_specialization": "Router specialization score vs energy (max gate weight)",
}
@@ -250,6 +251,54 @@ def compute_router_gating(
)
def compute_router_specialization(
checkpoint: str | Path | None,
r_phys: pl.LazyFrame,
t_phys: pl.LazyFrame,
seed: int = 0,
) -> Reduced:
"""Scalar specialization trend: max gate weight vs energy, per side.
Summarizes `router_gating`'s full per-expert stacked area into one curve —
the routing plan's own "how sharp is the boundary here" number (1/n_experts
= uniform/no specialization, 1.0 = one expert fully owns that energy). Same
quantile energy bins as `router_gating` (`_quantile_bins`), so this is
directly comparable to that plot's ceiling described in the roadmap's MoE
writeup.
"""
handle = load_router(checkpoint) if checkpoint else None
if handle is None:
return _unavailable("router_specialization")
sides: dict[str, dict] = {}
for name, lf in (("rollout", r_phys), ("reference", t_phys)):
df = _subsample(lf, _SAMPLE_ROWS, seed)
df, gate = _gate_for_df(handle, df)
x = df["pre_E"].to_numpy()
if len(x):
binned = _quantile_bins(x, gate, _N_BINS)
means = np.asarray(binned["means"])
score = means.max(axis=1).tolist() if means.size else []
sides[name] = {"centers": binned["centers"], "score": score}
else:
sides[name] = {"centers": [], "score": []}
return Reduced(
id="router_specialization",
family="model",
kind="router_specialization",
title=_TITLES["router_specialization"],
xlabel="pre-step energy [MeV]",
payload={
"router_type": handle.router_type,
"n_experts": handle.router.n_experts,
"log_x": True,
"chance_level": 1.0 / handle.router.n_experts,
**sides,
},
)
def compute_router_share_by_pdg(
checkpoint: str | Path | None,
r_phys: pl.LazyFrame,
+3 -3
View File
@@ -6,8 +6,8 @@ streaming `group_by` pass(es) over the chunk (see `catalog.py`/`reduce.py`).
`_COST_MODEL` below is ``spec_id -> (intercept_s, seconds_per_row)``.
``n_rows`` is the combined rollout+reference row count of the job's input:
the chunk's row count for `chunkable=True` specs, the whole dataset's for the
three `chunkable=False` router specs (they always run as a single job
regardless of chunk count).
`chunkable=False` router specs in `_ROUTER_IDS` (they always run as a single
job regardless of chunk count).
Calibrated 2026-07-27 from real HTCondor timings (`condor_history`
``RemoteWallClockTime``) of a production run: prediction ``563f5ee3``
@@ -54,7 +54,7 @@ _FIXED_OVERHEAD_S = 60.0
# scan. Calibrated from the 3 real router jobs' observed wall times (119, 66,
# 124s) — max minus _FIXED_OVERHEAD_S, on top of it.
_ROUTER_FIXED_S = 64.0
_ROUTER_IDS = frozenset({"router_gating", "router_share_by_pdg", "router_share_by_process"})
_ROUTER_IDS = frozenset({"router_gating", "router_share_by_pdg", "router_share_by_process", "router_specialization"})
# Conservative fallback for any catalog id not in _COST_MODEL (e.g. a plot
# added after the last calibration run) — the most expensive fitted per-row
+4
View File
@@ -202,6 +202,8 @@ def build_critics(model_config: dict) -> dict[str, nn.Module | None]:
cond_out_dim=cond_out_dim,
dropout=s1_spec.dropout,
stage="stage1",
trunk_type=s1_spec.trunk.type,
block_conditioning=s1_spec.trunk.block_conditioning,
)
if s2_spec.active and build_objective(s2_spec.generator).is_adversarial:
@@ -225,6 +227,8 @@ def build_critics(model_config: dict) -> dict[str, nn.Module | None]:
dropout=s2_spec.dropout,
stage="stage2",
context_dim=s2_spec.context_dim,
trunk_type=s2_spec.trunk.type,
block_conditioning=s2_spec.trunk.block_conditioning,
)
return result
+17
View File
@@ -63,6 +63,23 @@ def build_history(name: str, in_dim: int, out_dim: int, **kwargs) -> HistoryEnco
return cls(in_dim, out_dim, **filtered)
@register_history("none")
class NoHistory(HistoryEncoder):
"""No history signal at all — ignores feat/has_prev entirely and always
returns zeros. Ablates whether the AR decoder's history conditioning is
earning its parameters. `init_cache`/`step` use the base class's O(1)
defaults unmodified (this encoder's own `forward` is already O(1) per
call regardless of prefix length)."""
def __init__(self, in_dim: int, out_dim: int) -> None:
super().__init__()
self.out_dim = out_dim
def forward(self, feat: torch.Tensor, has_prev: torch.Tensor) -> torch.Tensor:
B, K, _ = feat.shape
return torch.zeros(B, K, self.out_dim, device=feat.device, dtype=feat.dtype)
@register_history("markov")
class MarkovHistory(HistoryEncoder):
"""Summarizes the previous secondary's own `(energy_fraction, direction,
+57 -32
View File
@@ -8,7 +8,7 @@ from giant.config import ConditioningAxisConfig, HeadConfig, ParticleTypeConfig
from giant.constants import CONT_SLOT_DIM, K_MAX, PARTICLE_PHYS_DIM, SEC_DIM, SEC_SLOT_DIM, X_DIM
from giant.model.encoders import ConditionEncoder
from giant.model.history import HistoryEncoder, build_history
from giant.model.layers import ContextAdapter, ResBlock, SinusoidalEmbedding, build_mlp_head
from giant.model.layers import ContextAdapter, SinusoidalEmbedding, build_mlp_head
from giant.model.objectives import build_objective
from giant.model.routers import Router
from giant.model.trunks import build_trunk
@@ -188,6 +188,26 @@ class StageModel(nn.Module):
hidden = max(1, round(hidden_dim * head_cfg.hidden_ratio))
self.stop_head = build_mlp_head(cond_out_dim, 1, hidden, head_cfg.depth)
def _build_context_fusion(self, x_dim: int, context_dim: int, cond_out_dim: int) -> None:
"""Builds `self.context_adapter`/`self.fuse` — the stage-2-style
context-fusion pattern (project the previous stage's outcome down to
`context_dim` via `ContextAdapter`, concat onto the base conditioning,
project back to `cond_out_dim`) shared by `Stage2OneShot` and a
`stage="stage2"` `CriticModel` (gitea #57). Call from a subclass's
`__init__` before using `_cond_embed`."""
self.context_adapter = ContextAdapter(x_dim, context_dim)
self.fuse = nn.Sequential(
nn.Linear(cond_out_dim + context_dim, cond_out_dim),
nn.SiLU(),
)
def _cond_embed(self, cond_cont: torch.Tensor, cond_cat: torch.Tensor, stage1_out: torch.Tensor) -> torch.Tensor:
"""Fuses base conditioning with the previous stage's outcome — pairs
with `_build_context_fusion`."""
base = self.cond_enc(cond_cont, cond_cat)
ctx = self.context_adapter(stage1_out)
return self.fuse(torch.cat([base, ctx], dim=-1))
def _require_n_sec_head(self) -> None:
if self.n_sec_head is None:
raise RuntimeError(
@@ -363,11 +383,7 @@ class Stage2OneShot(StageModel):
particle_type_cfg=particle_type_cfg,
cond_enc=cond_enc,
)
self.context_adapter = ContextAdapter(x_dim, context_dim)
self.fuse = nn.Sequential(
nn.Linear(cond_out_dim + context_dim, cond_out_dim),
nn.SiLU(),
)
self._build_context_fusion(x_dim, context_dim, cond_out_dim)
target = self.particle_type_cfg.target
type_head_out_dim = None if target == "physical" else k_max * self.type_dim
self._build_trunk_and_heads(
@@ -386,11 +402,6 @@ class Stage2OneShot(StageModel):
type_head_cfg=type_head_cfg,
)
def _cond_embed(self, cond_cont: torch.Tensor, cond_cat: torch.Tensor, stage1_out: torch.Tensor) -> torch.Tensor:
base = self.cond_enc(cond_cont, cond_cat)
ctx = self.context_adapter(stage1_out)
return self.fuse(torch.cat([base, ctx], dim=-1))
def forward(
self,
x_t: torch.Tensor,
@@ -702,11 +713,25 @@ class Stage2Autoregressive(StageModel):
return self.stop_head(c_emb.reshape(B * K, -1)).view(B, K)
class CriticModel(nn.Module):
class CriticModel(StageModel):
"""Generator-agnostic WGAN-GP critic body: a scalar realism score, for
either stage (`stage="stage1"` mirrors v0.2 `Critic`; `stage="stage2"`
mirrors v0.2 `SecondaryCritic`, adding the same context-fusion path as
`Stage2OneShot`). Used only when that stage's `generator == "wgan"`."""
`Stage2OneShot`, via `StageModel._build_context_fusion`/`_cond_embed`).
Used only when that stage's `generator == "wgan"`.
Subclasses `StageModel` for the `cond_enc` construction and (stage 2)
context-fusion scaffolding only its trunk is built directly via
`build_trunk` (output width 1) rather than through
`_build_trunk_and_heads`, since that helper is shaped around a
generator's `Objective`/time-embedding/flow-matching concerns
(`forward`'s `(x_t, cond) -> vector` shape) that don't apply to a critic's
`(x, cond) -> scalar` (gitea #57). `generator="wgan"` is passed to the
base purely because that's factually when a critic exists; nothing here
ever calls `_build_trunk_and_heads`, so no head/time-embedding machinery
is built from it. Never routed (MoE) that's a separate, unrequested
axis of scope; see gitea #57's proposal, which covers only the trunk/
block registries."""
def __init__(
self,
@@ -722,22 +747,26 @@ class CriticModel(nn.Module):
stage: str = "stage1",
context_dim: int = 64,
context_in_dim: int = X_DIM,
trunk_type: str = "resmlp",
block_conditioning: str = "add",
) -> None:
super().__init__()
super().__init__(
pdg_vocab,
mat_vocab,
particle_cfg,
material_cfg,
cond_out_dim=cond_out_dim,
generator="wgan",
noise_dim=0,
)
if stage not in ("stage1", "stage2"):
raise ValueError(f"stage must be 'stage1' or 'stage2', got {stage!r}")
self.stage = stage
self.cond_enc = ConditionEncoder(pdg_vocab, mat_vocab, particle_cfg, material_cfg, out_dim=cond_out_dim)
if stage == "stage2":
self.context_adapter = ContextAdapter(context_in_dim, context_dim)
self.fuse = nn.Sequential(
nn.Linear(cond_out_dim + context_dim, cond_out_dim),
nn.SiLU(),
)
self.input_proj = nn.Linear(in_dim, hidden_dim)
self.blocks = nn.ModuleList([ResBlock(hidden_dim, cond_out_dim, dropout=dropout) for _ in range(n_res_blocks)])
self.out_norm = nn.LayerNorm(hidden_dim)
self.out_proj = nn.Linear(hidden_dim, 1)
self._build_context_fusion(context_in_dim, context_dim, cond_out_dim)
self.trunk = build_trunk(
None, trunk_type, in_dim, 1, hidden_dim, n_res_blocks, cond_out_dim, dropout, block_conditioning
)
def forward(
self,
@@ -746,13 +775,9 @@ class CriticModel(nn.Module):
cond_cat: torch.Tensor,
stage1_out: torch.Tensor | None = None,
) -> torch.Tensor:
base = self.cond_enc(cond_cont, cond_cat)
if self.stage == "stage2":
ctx = self.context_adapter(stage1_out)
cond = self.fuse(torch.cat([base, ctx], dim=-1))
assert stage1_out is not None, "stage='stage2' CriticModel requires stage1_out"
cond = self._cond_embed(cond_cont, cond_cat, stage1_out)
else:
cond = base
h = self.input_proj(x)
for block in self.blocks:
h = block(h, cond)
return self.out_proj(self.out_norm(h)).squeeze(-1)
cond = self.cond_enc(cond_cont, cond_cat)
return self.trunk(x, cond, cond_cont, cond_cat).squeeze(-1)
+6
View File
@@ -15,6 +15,7 @@ from giant.model.history import (
AttentionHistory,
HistoryEncoder,
MarkovHistory,
NoHistory,
_CausalAttnBlock,
build_history,
register_history,
@@ -54,6 +55,7 @@ from giant.model.routers import (
ROUTER_REGISTRY,
ComposedRouter,
EnergyRouter,
NoneRouter,
PdgRouter,
ProcessRouter,
Router,
@@ -67,6 +69,7 @@ from giant.model.routers import (
from giant.model.trunks import (
TRUNK_REGISTRY,
ExpertTrunk,
LinearTrunk,
RoutedTrunk,
Trunk,
_route_forward,
@@ -90,7 +93,10 @@ __all__ = [
"FlowObjective",
"HISTORY_REGISTRY",
"HistoryEncoder",
"LinearTrunk",
"MarkovHistory",
"NoHistory",
"NoneRouter",
"OBJECTIVE_REGISTRY",
"Objective",
"PdgRouter",
+18
View File
@@ -144,6 +144,24 @@ def _inverse_bounded_interp(value: float, lo: float, hi: float) -> float:
return math.log(p / (1 - p))
@register_router("none")
class NoneRouter(Router):
"""Uniform 1/n_experts gate — no learned routing signal at all.
Still builds n_experts expert trunks via RoutedTrunk (same parameter
budget as a real router), but every row gets an identical weight
regardless of conditioning. Ablates whether the *learned routing
signal* as opposed to simply having multiple experts is earning
its parameters. `top1()` (the base class default) always dispatches to
expert 0 (argmax of a uniform vector), which still exercises
RoutedTrunk's real per-expert grouped-dispatch code path at eval time.
"""
def gate(self, cond_cont: torch.Tensor, cond_cat: torch.Tensor) -> torch.Tensor:
B = cond_cont.shape[0]
return torch.full((B, self.n_experts), 1.0 / self.n_experts, device=cond_cont.device)
@register_router("energy")
class EnergyRouter(Router):
"""Soft turn-on gate over normalized pre-step log-energy.
+43
View File
@@ -94,6 +94,49 @@ class ExpertTrunk(nn.Module):
return self.out_proj(x)
@register_trunk("linear")
class LinearTrunk(nn.Module):
"""`nn.Linear(in_dim + cond_dim, out_dim)` over `concat([x, cond])` —
the trivial trunk body: no hidden layer, no ResBlock stack, no
nonlinearity. Ablates whether trunk depth/nonlinearity is earning its
parameters, holding everything else (heads, ConditionEncoder,
generator, ...) fixed. Composes for free with `router.enabled = true`
(gitea #33): a RoutedTrunk of n_experts linear bodies is "mixture of
trivial linear experts". `hidden_dim`/`n_blocks`/`dropout`/
`block_conditioning` are accepted and ignored, matching
`build_expert_body`'s shared factory signature.
`x` the trunk's own input (e.g. the noised primary vector for flow
matching) does not already carry conditioning; that's fused in
per-body via `cond`. So this concatenates `x` and `cond` itself to
remain a valid, conditioning-dependent model.
"""
def __init__(
self,
in_dim: int,
out_dim: int,
hidden_dim: int,
n_blocks: int,
cond_dim: int,
dropout: float = 0.0,
block_conditioning: str = "add",
) -> None:
super().__init__()
self.in_dim = in_dim
self.out_dim = out_dim
self.linear = nn.Linear(in_dim + cond_dim, out_dim)
def forward(
self,
x: torch.Tensor,
cond: torch.Tensor,
cond_cont: torch.Tensor | None = None,
cond_cat: torch.Tensor | None = None,
) -> torch.Tensor:
return self.linear(torch.cat([x, cond], dim=-1))
def _route_forward(
experts: nn.ModuleList,
router: Router,
+1 -1
View File
@@ -1,6 +1,6 @@
[project]
name = "giant"
version = "0.3.4"
version = "0.3.7"
description = "Geant4 step-function surrogate via conditional flow matching"
readme = "README.md"
requires-python = ">=3.12"
+14
View File
@@ -159,6 +159,20 @@ def test_secondaries_rollout_vs_reference_align():
assert t["pdg"].to_list() == [22, 22]
def test_sec_count_by_event_zero_fills_events_with_no_secondaries():
r_phys = physical_steps(_rollout_frame(), Side.rollout)
r_sec = secondaries(_rollout_frame(), Side.rollout)
ev, n = R.sec_count_by_event(r_phys, r_sec)
# event 1 has one secondary track; event 2 has none and must still appear (as 0),
# not silently drop out of a plain group_by on the secondaries frame alone.
assert dict(zip(ev.tolist(), n.tolist())) == {1: 1, 2: 0}
t_all = _reference_frame()
t_sec = secondaries(t_all, Side.reference)
ev, n = R.sec_count_by_event(t_all, t_sec)
assert dict(zip(ev.tolist(), n.tolist())) == {1: 1, 2: 1}
def test_leakage_fraction():
frac = R.leakage_fraction(_rollout_frame())
# event 1: escaped pre_E=30, deposited=90 -> 30/120 = 0.25; event 2: 0
+66 -2
View File
@@ -6,7 +6,13 @@ import numpy as np
import pytest
from giant.analysis import build_catalog, catalog_ids, get_spec
from giant.analysis.catalog import Bundle, PlotSpec
from giant.analysis.catalog import (
Bundle,
PlotSpec,
_containment_depths,
_integer_confusion,
_ks_statistic,
)
from giant.analysis.context import Context, build_context
from tests.test_analysis_reduce import _reference_frame, _rollout_frame
@@ -53,6 +59,8 @@ def test_every_spec_computes_valid_reduced(bundle: Bundle):
"single_hist",
"router_gating",
"router_share",
"router_specialization",
"heatmap",
"unavailable",
}
assert r.title and r.xlabel
@@ -88,6 +96,14 @@ def _validate_payload(r) -> None:
for side in ("rollout", "reference"):
if side in p:
assert cat in p[side]
elif r.kind == "router_specialization":
for side in ("rollout", "reference"):
if side in p:
assert len(p[side]["centers"]) == len(p[side]["score"])
elif r.kind == "heatmap":
assert len(p["matrix"]) == len(p["row_labels"])
for row in p["matrix"]:
assert len(row) == len(p["col_labels"])
# ---------------------------------------------------------------------------
@@ -98,7 +114,10 @@ def _validate_payload(r) -> None:
# sec_count_per_species via pdg-keyed sums), concat-then-finalize with
# data-dependent edges (event_total_edep), concat-then-mean/std (shower_
# longitudinal), concat-then-max-edge (leakage_fraction), pdg-keyed sum with a
# ratio (species_edep_share), and a chunkable=False passthrough (router_gating).
# ratio (species_edep_share), a chunkable=False passthrough (router_gating),
# nested sum-merge into a scorecard (marginal_distance_summary), concat-then-
# event-id-join (n_sec_confusion), and concat-then-per-event-derived-quantity
# (shower_containment_depth_90, reusing the profile matrix's own merge shape).
_CHUNK_EQUIVALENCE_IDS = [
"marginal_edep",
"species_edep_share",
@@ -107,6 +126,9 @@ _CHUNK_EQUIVALENCE_IDS = [
"leakage_fraction",
"sec_count_per_species",
"router_gating",
"marginal_distance_summary",
"n_sec_confusion",
"shower_containment_depth_90",
]
@@ -146,3 +168,45 @@ def test_chunked_matches_unchunked(ctx: Context, spec_id: str):
assert chunked.id == unchunked.id
assert chunked.kind == unchunked.kind
_assert_payload_close(unchunked.payload, chunked.payload)
# ---------------------------------------------------------------------------
# new (gitea #76) reductions: KS distance, confusion matrix, containment depth
# ---------------------------------------------------------------------------
def test_ks_statistic():
assert _ks_statistic([10, 10], [10, 10]) == 0.0 # identical shape -> 0
assert _ks_statistic([10, 0], [0, 10]) == 1.0 # fully disjoint -> 1
assert _ks_statistic([0, 0], [0, 0]) != _ks_statistic([0, 0], [0, 0]) # nan (no data either side)
assert _ks_statistic([10, 0], [0, 0]) == 1.0 # one side empty, other isn't -> maximal mismatch
def test_integer_confusion_matches_event_pairing():
# true (reference) n_sec = [1, 1]; predicted (rollout) n_sec = [1, 0]
labels, mat = _integer_confusion(np.array([1, 1]), np.array([1, 0]))
assert labels == ["0", "1+"]
assert mat.tolist() == [[0, 0], [1, 1]] # row=true, col=pred
def test_integer_confusion_caps_pathological_outliers():
labels, mat = _integer_confusion(np.array([0, 500]), np.array([0, 0]), max_bins=5)
assert labels[-1] == "4+"
assert mat.shape == (5, 5)
assert mat.sum() == 2
def test_containment_depths_simple_ramp():
# one event, edep concentrated in the first bin -> 90%/95% containment
# depth is the first bin's right edge; a zero-energy event is dropped.
mat = np.array([[9.0, 1.0, 0.0], [0.0, 0.0, 0.0]])
edges = np.array([0.0, 1.0, 2.0, 3.0])
depths = _containment_depths(mat, edges, 0.90)
assert depths.tolist() == [1.0]
def test_n_sec_confusion_spec(bundle):
spec = get_spec("n_sec_confusion")
r = spec.finalize([spec.compute_partial(bundle)], bundle.ctx)
assert r.payload["row_labels"] == r.payload["col_labels"] == ["0", "1+"]
assert r.payload["matrix"] == [[0, 0], [1, 1]]
+85 -14
View File
@@ -8,8 +8,12 @@ from giant.model.network import (
HISTORY_REGISTRY,
AttentionHistory,
ConditionEncoder,
CriticModel,
FilmResBlock,
HistoryEncoder,
LinearTrunk,
MarkovHistory,
NoHistory,
SinusoidalEmbedding,
Stage1Model,
Stage2Autoregressive,
@@ -465,16 +469,52 @@ def test_attention_history_step_matches_forward():
assert torch.allclose(stepped, expected, atol=1e-5)
# --- NoHistory (gitea #45) ----------------------------------------------------
def test_no_history_shape():
hist = NoHistory(in_dim=7, out_dim=12)
B, K = 3, 5
feat = torch.randn(B, K, 7)
has_prev = (torch.arange(K) >= 1).unsqueeze(0).expand(B, -1)
out = hist(feat, has_prev)
assert out.shape == (B, K, 12)
def test_no_history_ignores_feat_and_has_prev():
hist = NoHistory(in_dim=4, out_dim=6)
B, K = 2, 3
has_prev_a = (torch.arange(K) >= 1).unsqueeze(0).expand(B, -1)
has_prev_b = torch.zeros(B, K, dtype=torch.bool)
feat_a = torch.randn(B, K, 4)
feat_b = torch.randn(B, K, 4) * 100
out_a = hist(feat_a, has_prev_a)
out_b = hist(feat_b, has_prev_b)
assert torch.equal(out_a, torch.zeros(B, K, 6))
assert torch.equal(out_a, out_b)
def test_no_history_uses_base_class_o1_defaults():
hist = NoHistory(in_dim=4, out_dim=6)
assert hist.init_cache() is None
feat = torch.randn(2, 1, 4)
has_prev = torch.ones(2, 1, dtype=torch.bool)
out, cache = hist.step(feat, has_prev, "unused-cache")
assert torch.equal(out, torch.zeros(2, 1, 6))
assert cache == "unused-cache"
# --- HISTORY_REGISTRY / build_history (gitea #35) ----------------------------
def test_history_registry_has_exactly_the_two_known_histories():
assert set(HISTORY_REGISTRY) == {"markov", "attention"}
def test_history_registry_has_exactly_the_known_histories():
assert set(HISTORY_REGISTRY) == {"markov", "attention", "none"}
def test_build_history_returns_correct_concrete_type():
assert isinstance(build_history("markov", 4, 6), MarkovHistory)
assert isinstance(build_history("attention", 4, 8), AttentionHistory)
assert isinstance(build_history("none", 4, 6), NoHistory)
def test_build_history_unknown_name_raises():
@@ -605,7 +645,7 @@ def test_stage2_autoregressive_n_sec_head_and_type_head_cfg_control_hidden_width
@pytest.mark.parametrize("target", ["physical", "onehot", "embedding"])
@pytest.mark.parametrize("generator", ["wgan", "flow"])
@pytest.mark.parametrize("history", ["markov", "attention"])
@pytest.mark.parametrize("history", ["markov", "attention", "none"])
def test_stage2_autoregressive_forward_shape(target, generator, history):
B, K, emb_dim = 4, 5, 6
model = _build_stage2_ar(target, generator, emb_dim=emb_dim, k_max=K, history=history)
@@ -907,7 +947,7 @@ def test_build_critics_particle_type_n_classes_overrides_conditioning_emb_dim():
assert wider_critic is not None
# k_max=3 slots, each CONT_SLOT_DIM + n_classes wide under wgan folding —
# widening n_classes alone (emb_dim stays 4) must widen the critic input.
assert wider_critic.input_proj.in_features > default_n_classes_critic.input_proj.in_features
assert wider_critic.trunk.input_proj.in_features > default_n_classes_critic.trunk.input_proj.in_features
# ── build_models/build_critics: DEFAULT_CONFIG fallback drift (issues.md #1) ─
@@ -984,12 +1024,12 @@ def test_build_critics_omitted_particle_type_matches_default_config():
cfg["stage2_model"]["generator"] = "wgan"
onehot_critic = build_critics(cfg)["stage2"]
assert onehot_critic is not None
onehot_in_dim = onehot_critic.input_proj.in_features
onehot_in_dim = onehot_critic.trunk.input_proj.in_features
cfg["stage2_model"]["particle_type"] = {"target": "physical"}
physical_critic = build_critics(cfg)["stage2"]
assert physical_critic is not None
physical_in_dim = physical_critic.input_proj.in_features
physical_in_dim = physical_critic.trunk.input_proj.in_features
# onehot's per-slot type width is emb_dim classes vs. physical's fixed
# (log-mass, charge) pair — different unless emb_dim happens to be 2, so
@@ -1009,15 +1049,15 @@ def test_build_critics_stage1_critic_hidden_dim_and_n_res_blocks_override_genera
inherited = build_critics(cfg)["stage1"]
assert inherited is not None
assert inherited.input_proj.out_features == 8
assert len(inherited.blocks) == 1
assert inherited.trunk.input_proj.out_features == 8
assert len(inherited.trunk.blocks) == 1
cfg["stage1_model"]["wgan"]["critic_hidden_dim"] = 16
cfg["stage1_model"]["wgan"]["critic_n_res_blocks"] = 3
overridden = build_critics(cfg)["stage1"]
assert overridden is not None
assert overridden.input_proj.out_features == 16
assert len(overridden.blocks) == 3
assert overridden.trunk.input_proj.out_features == 16
assert len(overridden.trunk.blocks) == 3
def test_build_critics_stage2_critic_hidden_dim_and_n_res_blocks_override_generator_size():
@@ -1028,15 +1068,15 @@ def test_build_critics_stage2_critic_hidden_dim_and_n_res_blocks_override_genera
inherited = build_critics(cfg)["stage2"]
assert inherited is not None
assert inherited.input_proj.out_features == 8
assert len(inherited.blocks) == 1
assert inherited.trunk.input_proj.out_features == 8
assert len(inherited.trunk.blocks) == 1
cfg["stage2_model"]["wgan"]["critic_hidden_dim"] = 16
cfg["stage2_model"]["wgan"]["critic_n_res_blocks"] = 3
overridden = build_critics(cfg)["stage2"]
assert overridden is not None
assert overridden.input_proj.out_features == 16
assert len(overridden.blocks) == 3
assert overridden.trunk.input_proj.out_features == 16
assert len(overridden.trunk.blocks) == 3
# ── StageModel base (gitea #39): Stage1Model/Stage2OneShot/Stage2Autoregressive
@@ -1197,3 +1237,34 @@ def test_stagemodel_time_emb_matches_objective_needs_time(cls, generator):
assert model.generator_kind == generator
assert model.noise_dim == 8
assert (model.time_emb is not None) == build_objective(generator).needs_time
# ── CriticModel uses the trunk/block registries + StageModel base (gitea #57) ─
def test_critic_model_is_stagemodel_subclass():
assert issubclass(CriticModel, StageModel)
@pytest.mark.parametrize("stage", ["stage1", "stage2"])
def test_build_critics_threads_trunk_type_from_generator_config(stage):
cfg = _minimal_model_config(share_stages=False)
cfg["stage1_model"]["generator"] = "wgan"
cfg["stage2_model"]["generator"] = "wgan"
cfg[f"{stage}_model"]["trunk"] = {"type": "linear"}
critic = build_critics(cfg)[stage]
assert critic is not None
assert isinstance(critic.trunk, LinearTrunk)
@pytest.mark.parametrize("stage", ["stage1", "stage2"])
def test_build_critics_threads_block_conditioning_from_generator_config(stage):
cfg = _minimal_model_config(share_stages=False)
cfg["stage1_model"]["generator"] = "wgan"
cfg["stage2_model"]["generator"] = "wgan"
cfg[f"{stage}_model"]["trunk"] = {"block_conditioning": "film"}
critic = build_critics(cfg)[stage]
assert critic is not None
assert all(isinstance(block, FilmResBlock) for block in critic.trunk.blocks)
+71
View File
@@ -13,6 +13,8 @@ from giant.model.network import (
EnergyRouter,
ExpertTrunk,
FilmResBlock,
LinearTrunk,
NoneRouter,
PdgRouter,
ProcessRouter,
ROUTER_REGISTRY,
@@ -73,6 +75,34 @@ def test_energy_router_registered():
assert ROUTER_REGISTRY["energy"] is EnergyRouter
# ── NoneRouter (gitea #45) ──────────────────────────────────────────────────
def test_none_router_registered():
assert ROUTER_REGISTRY["none"] is NoneRouter
def test_none_router_gate_is_uniform():
router = NoneRouter(n_experts=4)
cond_cont, cond_cat = _cond(16)
g = router.gate(cond_cont, cond_cat)
assert g.shape == (16, 4)
torch.testing.assert_close(g, torch.full((16, 4), 0.25))
def test_none_router_gate_ignores_conditioning():
router = NoneRouter(n_experts=3)
cond_cont_a, cond_cat_a = _cond(8)
cond_cont_b, cond_cat_b = _cond(8)
torch.testing.assert_close(router.gate(cond_cont_a, cond_cat_a), router.gate(cond_cont_b, cond_cat_b))
def test_none_router_top1_always_expert_zero():
router = NoneRouter(n_experts=4)
cond_cont, cond_cat = _cond(16)
assert torch.equal(router.top1(cond_cont, cond_cat), torch.zeros(16, dtype=torch.long))
def test_energy_router_gate_partition_of_unity():
router = EnergyRouter(n_experts=4)
cond_cont, cond_cat = _cond(16)
@@ -168,6 +198,47 @@ def test_build_expert_body_unknown_type_raises():
raise AssertionError("expected ValueError for unknown trunk type")
# ── LinearTrunk (gitea #45) ──────────────────────────────────────────────────
def test_trunk_registry_has_linear():
assert "linear" in TRUNK_REGISTRY
assert TRUNK_REGISTRY["linear"] is LinearTrunk
def test_linear_trunk_forward_shape():
trunk = build_expert_body("linear", in_dim=9, out_dim=9, hidden_dim=64, n_blocks=6, cond_dim=12)
assert trunk.in_dim == 9
assert trunk.out_dim == 9
x = torch.randn(5, 9)
cond = torch.randn(5, 12)
out = trunk(x, cond)
assert out.shape == (5, 9)
def test_linear_trunk_depends_on_x_and_cond():
trunk = build_expert_body("linear", in_dim=9, out_dim=9, hidden_dim=64, n_blocks=6, cond_dim=12)
x = torch.randn(5, 9)
cond_a = torch.randn(5, 12)
cond_b = torch.randn(5, 12)
assert not torch.allclose(trunk(x, cond_a), trunk(x, cond_b))
def test_routed_linear_trunk_is_mixture_of_trivial_experts():
"""trunk.type = 'linear' composes for free with router.enabled = true
(gitea #33's comment on this issue) — a RoutedTrunk of n_experts linear
bodies."""
router = build_router("energy", n_experts=3)
trunk = RoutedTrunk(router, "linear", in_dim=9, out_dim=9, hidden_dim=64, n_res_blocks=6, cond_dim=12)
assert len(trunk.experts) == 3
assert all(isinstance(e, LinearTrunk) for e in trunk.experts)
x = torch.randn(5, 9)
cond = torch.randn(5, 12)
cond_cont, cond_cat = _cond(5)
out = trunk(x, cond, cond_cont, cond_cat)
assert out.shape == (5, 9)
# ── BLOCK_REGISTRY / build_block (gitea #34) ────────────────────────────────
+28 -1
View File
@@ -2,7 +2,7 @@ import torch
from giant.config import ConditioningAxisConfig
from giant.constants import COND_DIM, K_MAX, SEC_DIM, SEC_SLOT_DIM, X_DIM
from giant.model.network import CriticModel, Stage1Model, Stage2OneShot
from giant.model.network import CriticModel, LinearTrunk, Stage1Model, Stage2OneShot
from giant.model.wgan import critic_loss, generator_loss, gradient_penalty
from giant.sample import sample_secondaries_wgan, sample_wgan
@@ -115,6 +115,33 @@ def test_critic_output_shape():
assert out.shape == (B,)
def test_critic_model_honours_trunk_type_and_block_conditioning():
"""gitea #57: CriticModel routes its body through build_trunk/build_block
like every generator stage model, instead of hand-rolling a plain
ResBlock stack."""
B = 8
critic = CriticModel(
pdg_vocab=3,
mat_vocab=2,
particle_cfg=PARTICLE_CFG,
material_cfg=MATERIAL_CFG,
in_dim=X_DIM,
hidden_dim=32,
n_res_blocks=2,
stage="stage1",
trunk_type="linear",
block_conditioning="adaln",
)
assert isinstance(critic.trunk, LinearTrunk)
cond_cont, cond_cat = _cond(B)
real = torch.randn(B, X_DIM)
fake = torch.randn(B, X_DIM)
loss = critic_loss(lambda x: critic(x, cond_cont, cond_cat), real, fake.detach(), gp_weight=10.0)
loss.backward()
for name, p in critic.named_parameters():
assert p.grad is not None, f"no grad for {name}"
def test_sample_wgan_shape():
B = 6
model = _small_generator()
Generated
+1 -1
View File
@@ -675,7 +675,7 @@ wheels = [
[[package]]
name = "giant"
version = "0.3.4"
version = "0.3.7"
source = { editable = "." }
dependencies = [
{ name = "numpy" },