Files
OpenSWETraces_cleanup/README.md
T

159 lines
6.3 KiB
Markdown

# Open-SWE-Traces Deterministic Profiler
This project profiles `nvidia/Open-SWE-Traces` for coding-agent SFT without an
LLM judge, repository sandbox, or trajectory rewriting. It never claims that a
patch is correct. It records facts that can be reproduced from the dataset and
uses a small set of explicit heuristics to remove traces that are definitely
broken or unsuitable for a chosen training context.
The local dataset snapshot contains 207,489 trajectories over 22,320 issues:
| `resolved` | Meaning | Count |
|---:|---|---:|
| `1` | Externally marked successful | 65,244 |
| `0` | Externally marked failed | 95,487 |
| `-1` | Outcome unknown | 46,758 |
`resolved` is never changed. The published context splits retain successful,
failed, and unknown outcomes because non-successful traces can still contain
useful repository exploration, tool use, and recovery behavior. Downstream
training can use the preserved label to choose its own mixture.
## Design
The pipeline has four commands:
1. `profile` streams the original JSONL or Parquet dataset and writes one
deterministic metrics row per trajectory.
2. `summarize` computes dataset quantiles and writes transparent decisions.
3. `sample` selects full trajectories from every rule and length bucket for
human validation.
4. `materialize` writes two unbalanced, training-ready Parquet splits at the
131,072-token and 81,920-token cutoffs.
No command edits, truncates, repairs, or invents trajectory turns. Source data
and generated files are joined by `trajectory_id`/`sample_id`.
## Metrics
Each profile row contains:
- turn and role counts;
- canonical serialized characters and UTF-8 bytes;
- exact token count when a `tokenizer.json` is supplied;
- tool-call count and count by tool name;
- failed tool-call count and rate;
- observable error type and tool distributions;
- normalized error positions across five equal trajectory regions;
- early-error count and fraction;
- longest consecutive failed-tool-call run;
- malformed calls, unknown tools, missing results, and orphan results.
The canonical token stream is compact, sorted JSON containing only `tools` and
`trajectory`. This is reproducible but is not presented as a universal chat
template. If training uses another serializer or tokenizer, rerun `profile`
with that exact tokenizer instead of converting characters to fake token counts.
## Heuristics
The following conditions are hard rejections because they represent broken
training structure or an unambiguous repeated-failure pattern:
- malformed/unknown tool calls or broken assistant/tool pairing;
- at least five consecutive failed tool calls;
- at least five failed calls and a failure rate of at least 50%.
Three weaker patterns are review flags, not automatic rejection:
- error count at or above the dataset-wide p99 threshold;
- at least eight failed calls, a failure rate of at least 20%, and errors in at
least four of five trajectory regions;
- at least three errors with at least 60% of all errors in the first 20% of tool
calls.
Length is separated from quality. Tokenized records are assigned to:
- `LE_81920`: fast 80K training subset;
- `81921_TO_131072`: 128K training subset;
- `131073_TO_262144`: deferred long-context subset;
- `GT_262144`: excluded from the default small-model training run.
The summary reports how much of the total token mass is contributed by the
longest 0.1%, 1%, and 5% of samples. This makes long-tail removal a measurable
dataset decision rather than a guess.
## Installation
```bash
cd /home/kxqandccx/kxq/tomlu/OpenSWETraces_cleanup
./.venv/bin/pip install -e '.[dev,tokens]'
```
`tokenizers` is optional. Without it, profiling still produces turn, character,
byte, and tool-error metrics, but no token bucket decision.
## Full workflow
Use an explicit tokenizer already present on the machine:
```bash
TOKENIZER_JSON=/path/to/tokenizer.json
mkdir -p qc_outputs/deterministic_v3
swe-qc profile \
--input raw/Open-SWE-Traces \
--output qc_outputs/deterministic_v3/metrics.jsonl \
--errors qc_outputs/deterministic_v3/profile.errors.jsonl \
--tokenizer-json "$TOKENIZER_JSON" \
--workers 24 \
--resume
swe-qc summarize \
--metrics qc_outputs/deterministic_v3/metrics.jsonl \
--summary qc_outputs/deterministic_v3/summary.json \
--decisions qc_outputs/deterministic_v3/decisions.jsonl
swe-qc sample \
--input raw/Open-SWE-Traces \
--decisions qc_outputs/deterministic_v3/decisions.jsonl \
--output qc_outputs/deterministic_v3/review_sample.jsonl \
--per-group 20 \
--seed 20260818
swe-qc materialize \
--decisions qc_outputs/deterministic_v3/decisions.jsonl \
--output artifacts/modelscope_openswe_traces_repurpose \
--workers 16 \
--resume
```
`materialize` excludes hard rejects and review flags, but does not balance or
rewrite outcome categories. It preserves every source column and appends six
`qc_*` provenance columns. In the current full run, `context_131072` contains
186,665 trajectories (60,276 successful, 84,949 failed, and 41,440 unknown),
while `context_81920` contains 114,437 trajectories (42,596 successful, 48,145
failed, and 23,696 unknown). The shorter split is an exact subset of the longer
split.
For a Parquet directory and `--workers > 1`, each worker is a separate process
that reads one shard at a time and writes an atomic shard part. This bypasses
the Python GIL, bounds live trajectory memory to roughly one small Parquet batch
per process, and lets `profile --resume` reuse completed shard parts after an
interruption. The final metrics and error manifests are concatenated atomically
in stable shard order. For JSONL or a limited pilot, `--resume` instead skips
sample IDs already present in the append-only output. `summarize` and `sample`
also replace their outputs atomically.
Profile rows and decisions record the absolute source Parquet shard. `sample`
uses that provenance to read only shards containing selected IDs instead of
scanning the complete dataset.
## Error detection boundary
The detector intentionally favors precision over recall. It recognizes explicit
non-zero exit codes, test/build failure summaries, timeouts, permission errors,
missing commands/files, and tool exceptions. It does not treat the mere word
"error" as a failure, because tools often print source code or logs containing
that word. Human samples should be used to refine patterns only when the raw
tool observation provides an unambiguous signal.