Files
OpenSWETraces_cleanup/README.md
T

140 lines
5.4 KiB
Markdown

# Open-SWE-Traces Deterministic Profiler
This project profiles `nvidia/Open-SWE-Traces` for coding-agent SFT without an
LLM judge, repository sandbox, or trajectory rewriting. It never claims that a
patch is correct. It records facts that can be reproduced from the dataset and
uses a small set of explicit heuristics to remove traces that are definitely
broken or unsuitable for a chosen training context.
The local dataset snapshot contains 207,489 trajectories over 22,320 issues:
| `resolved` | Meaning | Count |
|---:|---|---:|
| `1` | Externally marked successful | 65,244 |
| `0` | Externally marked failed | 95,487 |
| `-1` | Outcome unknown | 46,758 |
`resolved` is never changed. Only `resolved=1` is eligible for successful SFT;
the profiler can still describe failed and unknown traces for analysis. Final
decisions label `resolved=0` as `EXCLUDE_FAILED_OUTCOME` and `resolved=-1` as
`HOLD_UNVERIFIED_OUTCOME` rather than recommending them for training.
## Design
The pipeline has three commands:
1. `profile` streams the original JSONL or Parquet dataset and writes one
deterministic metrics row per trajectory.
2. `summarize` computes dataset quantiles and writes transparent decisions.
3. `sample` selects full trajectories from every rule and length bucket for
human validation.
No command edits, truncates, repairs, or invents trajectory turns. Source data
and generated files are joined by `trajectory_id`/`sample_id`.
## Metrics
Each profile row contains:
- turn and role counts;
- canonical serialized characters and UTF-8 bytes;
- exact token count when a `tokenizer.json` is supplied;
- tool-call count and count by tool name;
- failed tool-call count and rate;
- observable error type and tool distributions;
- normalized error positions across five equal trajectory regions;
- early-error count and fraction;
- longest consecutive failed-tool-call run;
- malformed calls, unknown tools, missing results, and orphan results.
The canonical token stream is compact, sorted JSON containing only `tools` and
`trajectory`. This is reproducible but is not presented as a universal chat
template. If training uses another serializer or tokenizer, rerun `profile`
with that exact tokenizer instead of converting characters to fake token counts.
## Heuristics
The following conditions are hard rejections because they represent broken
training structure or an unambiguous repeated-failure pattern:
- malformed/unknown tool calls or broken assistant/tool pairing;
- at least five consecutive failed tool calls;
- at least five failed calls and a failure rate of at least 50%.
Three weaker patterns are review flags, not automatic rejection:
- error count at or above the dataset-wide p99 threshold;
- at least eight failed calls, a failure rate of at least 20%, and errors in at
least four of five trajectory regions;
- at least three errors with at least 60% of all errors in the first 20% of tool
calls.
Length is separated from quality. Tokenized records are assigned to:
- `LE_81920`: fast 80K training subset;
- `81921_TO_131072`: 128K training subset;
- `131073_TO_262144`: deferred long-context subset;
- `GT_262144`: excluded from the default small-model training run.
The summary reports how much of the total token mass is contributed by the
longest 0.1%, 1%, and 5% of samples. This makes long-tail removal a measurable
dataset decision rather than a guess.
## Installation
```bash
cd /home/kxqandccx/kxq/tomlu/OpenSWETraces_cleanup
./.venv/bin/pip install -e '.[dev,tokens]'
```
`tokenizers` is optional. Without it, profiling still produces turn, character,
byte, and tool-error metrics, but no token bucket decision.
## Full workflow
Use an explicit tokenizer already present on the machine:
```bash
TOKENIZER_JSON=/path/to/tokenizer.json
mkdir -p qc_outputs/deterministic_v3
swe-qc profile \
--input raw/Open-SWE-Traces \
--output qc_outputs/deterministic_v3/metrics.jsonl \
--errors qc_outputs/deterministic_v3/profile.errors.jsonl \
--tokenizer-json "$TOKENIZER_JSON" \
--workers 8 \
--resume
swe-qc summarize \
--metrics qc_outputs/deterministic_v3/metrics.jsonl \
--summary qc_outputs/deterministic_v3/summary.json \
--decisions qc_outputs/deterministic_v3/decisions.jsonl
swe-qc sample \
--input raw/Open-SWE-Traces \
--decisions qc_outputs/deterministic_v3/decisions.jsonl \
--output qc_outputs/deterministic_v3/review_sample.jsonl \
--per-group 20 \
--seed 20260818
```
`profile --resume` skips sample IDs already present in the metrics file. Writes
are append-only and flushed per record. `summarize` and `sample` atomically
replace their outputs so a partial file is never mistaken for a complete run.
Any dataset subdirectory containing Parquet shards is also a valid `--input`,
so the four trajectory families can be profiled in separate CPU processes and
their metrics concatenated before `summarize`.
Profile rows and decisions record the absolute source Parquet shard. `sample`
uses that provenance to read only shards containing selected IDs instead of
scanning the complete dataset.
## Error detection boundary
The detector intentionally favors precision over recall. It recognizes explicit
non-zero exit codes, test/build failure summaries, timeouts, permission errors,
missing commands/files, and tool exceptions. It does not treat the mere word
"error" as a failure, because tools often print source code or logs containing
that word. Human samples should be used to refine patterns only when the raw
tool observation provides an unambiguous signal.