Add context-cutoff dataset materialization

This commit is contained in:
jiachun
2026-08-18 19:28:20 +08:00
parent fea85047f1
commit c7943843cc
6 changed files with 414 additions and 6 deletions
+21 -5
View File
@@ -14,20 +14,22 @@ The local dataset snapshot contains 207,489 trajectories over 22,320 issues:
| `0` | Externally marked failed | 95,487 |
| `-1` | Outcome unknown | 46,758 |
`resolved` is never changed. Only `resolved=1` is eligible for successful SFT;
the profiler can still describe failed and unknown traces for analysis. Final
decisions label `resolved=0` as `EXCLUDE_FAILED_OUTCOME` and `resolved=-1` as
`HOLD_UNVERIFIED_OUTCOME` rather than recommending them for training.
`resolved` is never changed. The published context splits retain successful,
failed, and unknown outcomes because non-successful traces can still contain
useful repository exploration, tool use, and recovery behavior. Downstream
training can use the preserved label to choose its own mixture.
## Design
The pipeline has three commands:
The pipeline has four commands:
1. `profile` streams the original JSONL or Parquet dataset and writes one
deterministic metrics row per trajectory.
2. `summarize` computes dataset quantiles and writes transparent decisions.
3. `sample` selects full trajectories from every rule and length bucket for
human validation.
4. `materialize` writes two unbalanced, training-ready Parquet splits at the
131,072-token and 81,920-token cutoffs.
No command edits, truncates, repairs, or invents trajectory turns. Source data
and generated files are joined by `trajectory_id`/`sample_id`.
@@ -117,8 +119,22 @@ swe-qc sample \
--output qc_outputs/deterministic_v3/review_sample.jsonl \
--per-group 20 \
--seed 20260818
swe-qc materialize \
--decisions qc_outputs/deterministic_v3/decisions.jsonl \
--output artifacts/modelscope_openswe_traces_repurpose \
--workers 16 \
--resume
```
`materialize` excludes hard rejects and review flags, but does not balance or
rewrite outcome categories. It preserves every source column and appends six
`qc_*` provenance columns. In the current full run, `context_131072` contains
186,665 trajectories (60,276 successful, 84,949 failed, and 41,440 unknown),
while `context_81920` contains 114,437 trajectories (42,596 successful, 48,145
failed, and 23,696 unknown). The shorter split is an exact subset of the longer
split.
For a Parquet directory and `--workers > 1`, each worker is a separate process
that reads one shard at a time and writes an atomic shard part. This bypasses
the Python GIL, bounds live trajectory memory to roughly one small Parquet batch