`scripts` was published as a top-level distribution package, colliding
with one of the most generic names in the Python ecosystem and
shadowable by a stray scripts/ dir on the portal machines' shared
/work/lbogner. Move it under the giant namespace; the dwarf command
name is unchanged, only the Python import path and file location move.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The prior _COST_MODEL/_FIXED_OVERHEAD_S were fit only against local
synthetic benchmarks (up to 2M rows/side), which can't see docker
pull or real /ceph read latency and wildly overestimated real jobs
(~1200-1800s predicted vs 50-320s median actual, from condor_history
on production run 563f5ee3, --chunks 4, ~254M total rows).
Refit each spec's per-row rate through the origin against its median
real wall-clock time (not max, to avoid baking a few /ceph-contention
spikes into a rate that would then wrongly scale with dataset size),
and raised RUNTIME_SAFETY_MARGIN to compensate for that same
contention risk instead.
Each condor job's +RequestWalltime used to be one flat 3600s default
for every (plot, chunk), regardless of how much data it actually
streams over. `prep` now records each chunk's rollout+reference row
count, and `giant/analysis/runtime_estimate.py` turns that into a
per-job estimate: a per-spec (intercept, seconds/row) cost model fit
by `scripts/profile_analysis_costs.py` against synthetic mock data on
this machine, plus a fixed overhead placeholder (docker/uv/shared-fs
startup — unmeasurable here, no /ceph access) and a single
RUNTIME_SAFETY_MARGIN multiplier. jobs.txt gains a walltime column and
the submit description references it via $(walltime) instead of a
constant.