One workflow TOML now parameterises a whole experiment and `giant workflow
run <spec.toml>` turns it into a b2luigi DAG whose targets are files on
/ceph: nothing already produced is recomputed, every step waits for its
inputs, and HTCondor submission/polling is b2luigi's job.
- spec.py: workflow TOML -> frozen dataclasses with name-uniqueness and
cross-reference validation, unknown keys rejected the way giant.config
rejects them, and a short spec_hash per task that folds in its transitive
parents — so an edited spec re-runs exactly the affected subtree.
- htcondor.py: the CPU/GPU submit settings. The GPU requirement strings
(ProvidesEtpCeph + optional device/memory pins) are ported from the
condor-gpu-train-rollout branch rather than rewritten.
- tasks.py: DatasetTask, WarmCacheTask, GeometryOracleTask, TrainEpochTask
(one short GPU job per epoch, chained via --resume, which the training
loop already supports unchanged), TrainTask (publishes best.pt/last.pt and
a concatenated metrics.csv so downstream never sees the epoch fan-out),
RolloutTask, AnalysisPrepTask, AnalysisComputeTask (one job per plot x
chunk, walltime sized from run_meta.json at submit time), AnalysisRenderTask
(always local — the only step importing plotstyle/LaTeX), WorkflowTask.
Task bodies call the existing entry points; none of them reimplement
anything.
- run.py + `giant workflow run`: settings wiring and the script b2luigi
re-executes on workers. add_filename_to_cmd is off because b2luigi passes
only the script's basename, and --spec is forwarded via
task_cmd_additional_args so a worker resolves the identical task graph.
configs/workflow_example.toml is the documented starting point.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>