Deduplicate n_sec_head/type_head MLPs into build_mlp_head (gitea #36) #53

Merged
lars merged 1 commits from fix/issue-36 into master 2026-08-14 11:05:00 +02:00
Owner

The same two-layer classifier head (Linear(cond_out_dim, hidden_dim // 2)
-> SiLU -> Linear(hidden_dim // 2, out_dim)) was hand-rolled five times in
giant/model/models.py: Stage1Model.n_sec_head, Stage2OneShot.n_sec_head/
.type_head, and Stage2Autoregressive.n_sec_head/.type_head. The // 2
ratio and fixed 2-layer depth were undocumented magic numbers, and both
n_sec accuracy and secondary-species accuracy are known weak spots that
were untunable independently of the trunk they hang off.

Adds build_mlp_head(in_dim, out_dim, hidden, depth, act) to
giant/model/layers.py (depth=1 is a bare Linear; depth>=2 matches the old
hardcoded shape exactly), and a new HeadConfig (hidden_ratio, depth)
dataclass in giant/config.py, wired in as stage1_model.heads.n_sec and
stage2_model.heads.{n_sec,type} — split per head type (not one shared
block per stage) since n_sec and species prediction are called out as
separate weak spots that may want independent capacity. Defaults
(hidden_ratio=0.5, depth=2) reproduce the old hardcoded architecture
bit-for-bit, so every existing config.toml and migrated v0.2 checkpoint
is unaffected; no changes were needed to migrate_config or the legacy
migration surfaces. No new CLI flags, matching how other nested
sub-config (router., trunk.) is set via config.toml rather than
per-field flags.

Co-Authored-By: Claude Opus 5 noreply@anthropic.com

The same two-layer classifier head (Linear(cond_out_dim, hidden_dim // 2) -> SiLU -> Linear(hidden_dim // 2, out_dim)) was hand-rolled five times in giant/model/models.py: Stage1Model.n_sec_head, Stage2OneShot.n_sec_head/ .type_head, and Stage2Autoregressive.n_sec_head/.type_head. The `// 2` ratio and fixed 2-layer depth were undocumented magic numbers, and both n_sec accuracy and secondary-species accuracy are known weak spots that were untunable independently of the trunk they hang off. Adds `build_mlp_head(in_dim, out_dim, hidden, depth, act)` to giant/model/layers.py (depth=1 is a bare Linear; depth>=2 matches the old hardcoded shape exactly), and a new `HeadConfig` (hidden_ratio, depth) dataclass in giant/config.py, wired in as `stage1_model.heads.n_sec` and `stage2_model.heads.{n_sec,type}` — split per head type (not one shared block per stage) since n_sec and species prediction are called out as separate weak spots that may want independent capacity. Defaults (hidden_ratio=0.5, depth=2) reproduce the old hardcoded architecture bit-for-bit, so every existing config.toml and migrated v0.2 checkpoint is unaffected; no changes were needed to migrate_config or the legacy migration surfaces. No new CLI flags, matching how other nested sub-config (router.*, trunk.*) is set via config.toml rather than per-field flags. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lars added 1 commit 2026-08-14 10:58:23 +02:00
Deduplicate n_sec_head/type_head MLPs into build_mlp_head (gitea #36)
CI / Format (ruff format) (push) Successful in 28s
CI / Lint (ruff check) (push) Successful in 29s
CI / Sync project version with tag (push) Has been skipped
CI / Lint (ruff check) (pull_request) Successful in 34s
CI / Type check (ty) (push) Successful in 37s
CI / Format (ruff format) (pull_request) Successful in 44s
CI / Type check (ty) (pull_request) Successful in 46s
CI / Sync project version with tag (pull_request) Has been skipped
CI / Tests (pull_request) Successful in 4m5s
CI / Tests (push) Successful in 4m6s
593c5f4d34
The same two-layer classifier head (Linear(cond_out_dim, hidden_dim // 2)
-> SiLU -> Linear(hidden_dim // 2, out_dim)) was hand-rolled five times in
giant/model/models.py: Stage1Model.n_sec_head, Stage2OneShot.n_sec_head/
.type_head, and Stage2Autoregressive.n_sec_head/.type_head. The `// 2`
ratio and fixed 2-layer depth were undocumented magic numbers, and both
n_sec accuracy and secondary-species accuracy are known weak spots that
were untunable independently of the trunk they hang off.

Adds `build_mlp_head(in_dim, out_dim, hidden, depth, act)` to
giant/model/layers.py (depth=1 is a bare Linear; depth>=2 matches the old
hardcoded shape exactly), and a new `HeadConfig` (hidden_ratio, depth)
dataclass in giant/config.py, wired in as `stage1_model.heads.n_sec` and
`stage2_model.heads.{n_sec,type}` — split per head type (not one shared
block per stage) since n_sec and species prediction are called out as
separate weak spots that may want independent capacity. Defaults
(hidden_ratio=0.5, depth=2) reproduce the old hardcoded architecture
bit-for-bit, so every existing config.toml and migrated v0.2 checkpoint
is unaffected; no changes were needed to migrate_config or the legacy
migration surfaces. No new CLI flags, matching how other nested
sub-config (router.*, trunk.*) is set via config.toml rather than
per-field flags.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lars merged commit b63edcb8f9 into master 2026-08-14 11:05:00 +02:00
lars deleted branch fix/issue-36 2026-08-14 11:05:00 +02:00
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: lars/giant#53