`_row_subsample` hash-filters on `pre_E`, but `build_context`'s
prediction branch called it on a frame already projected down to
`pred_<var>`/`true_<var>` columns, so any run whose prediction file
exceeds `sample_rows` (the default is 1M; real predict outputs can be
100M+ rows) failed with `ColumnNotFoundError: pre_E`.
Subsample the full paired frame first, then project — matching every
other _row_subsample call site — and restructure the loop to subsample
once per prediction side instead of once per paired variable, cutting
6 streaming passes over the prediction file down to 1.
Add a regression test with sample_rows below the fixture's row count
so the hash-filter branch is actually exercised (the existing test
never took it).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qb7xBAa6aAR94AzimgpPxq