Add context-cutoff dataset materialization
This commit is contained in:
@@ -14,20 +14,22 @@ The local dataset snapshot contains 207,489 trajectories over 22,320 issues:
|
||||
| `0` | Externally marked failed | 95,487 |
|
||||
| `-1` | Outcome unknown | 46,758 |
|
||||
|
||||
`resolved` is never changed. Only `resolved=1` is eligible for successful SFT;
|
||||
the profiler can still describe failed and unknown traces for analysis. Final
|
||||
decisions label `resolved=0` as `EXCLUDE_FAILED_OUTCOME` and `resolved=-1` as
|
||||
`HOLD_UNVERIFIED_OUTCOME` rather than recommending them for training.
|
||||
`resolved` is never changed. The published context splits retain successful,
|
||||
failed, and unknown outcomes because non-successful traces can still contain
|
||||
useful repository exploration, tool use, and recovery behavior. Downstream
|
||||
training can use the preserved label to choose its own mixture.
|
||||
|
||||
## Design
|
||||
|
||||
The pipeline has three commands:
|
||||
The pipeline has four commands:
|
||||
|
||||
1. `profile` streams the original JSONL or Parquet dataset and writes one
|
||||
deterministic metrics row per trajectory.
|
||||
2. `summarize` computes dataset quantiles and writes transparent decisions.
|
||||
3. `sample` selects full trajectories from every rule and length bucket for
|
||||
human validation.
|
||||
4. `materialize` writes two unbalanced, training-ready Parquet splits at the
|
||||
131,072-token and 81,920-token cutoffs.
|
||||
|
||||
No command edits, truncates, repairs, or invents trajectory turns. Source data
|
||||
and generated files are joined by `trajectory_id`/`sample_id`.
|
||||
@@ -117,8 +119,22 @@ swe-qc sample \
|
||||
--output qc_outputs/deterministic_v3/review_sample.jsonl \
|
||||
--per-group 20 \
|
||||
--seed 20260818
|
||||
|
||||
swe-qc materialize \
|
||||
--decisions qc_outputs/deterministic_v3/decisions.jsonl \
|
||||
--output artifacts/modelscope_openswe_traces_repurpose \
|
||||
--workers 16 \
|
||||
--resume
|
||||
```
|
||||
|
||||
`materialize` excludes hard rejects and review flags, but does not balance or
|
||||
rewrite outcome categories. It preserves every source column and appends six
|
||||
`qc_*` provenance columns. In the current full run, `context_131072` contains
|
||||
186,665 trajectories (60,276 successful, 84,949 failed, and 41,440 unknown),
|
||||
while `context_81920` contains 114,437 trajectories (42,596 successful, 48,145
|
||||
failed, and 23,696 unknown). The shorter split is an exact subset of the longer
|
||||
split.
|
||||
|
||||
For a Parquet directory and `--workers > 1`, each worker is a separate process
|
||||
that reads one shard at a time and writes an atomic shard part. This bypasses
|
||||
the Python GIL, bounds live trajectory memory to roughly one small Parquet batch
|
||||
|
||||
Reference in New Issue
Block a user