Open-SWE-Traces Deterministic Profiler
This project profiles nvidia/Open-SWE-Traces for coding-agent SFT without an
LLM judge, repository sandbox, or trajectory rewriting. It never claims that a
patch is correct. It records facts that can be reproduced from the dataset and
uses a small set of explicit heuristics to remove traces that are definitely
broken or unsuitable for a chosen training context.
The local dataset snapshot contains 207,489 trajectories over 22,320 issues:
resolved |
Meaning | Count |
|---|---|---|
1 |
Externally marked successful | 65,244 |
0 |
Externally marked failed | 95,487 |
-1 |
Outcome unknown | 46,758 |
resolved is never changed. Only resolved=1 is eligible for successful SFT;
the profiler can still describe failed and unknown traces for analysis. Final
decisions label resolved=0 as EXCLUDE_FAILED_OUTCOME and resolved=-1 as
HOLD_UNVERIFIED_OUTCOME rather than recommending them for training.
Design
The pipeline has three commands:
profilestreams the original JSONL or Parquet dataset and writes one deterministic metrics row per trajectory.summarizecomputes dataset quantiles and writes transparent decisions.sampleselects full trajectories from every rule and length bucket for human validation.
No command edits, truncates, repairs, or invents trajectory turns. Source data
and generated files are joined by trajectory_id/sample_id.
Metrics
Each profile row contains:
- turn and role counts;
- canonical serialized characters and UTF-8 bytes;
- exact token count when a
tokenizer.jsonis supplied; - tool-call count and count by tool name;
- failed tool-call count and rate;
- observable error type and tool distributions;
- normalized error positions across five equal trajectory regions;
- early-error count and fraction;
- longest consecutive failed-tool-call run;
- malformed calls, unknown tools, missing results, and orphan results.
The canonical token stream is compact, sorted JSON containing only tools and
trajectory. This is reproducible but is not presented as a universal chat
template. If training uses another serializer or tokenizer, rerun profile
with that exact tokenizer instead of converting characters to fake token counts.
Heuristics
The following conditions are hard rejections because they represent broken training structure or an unambiguous repeated-failure pattern:
- malformed/unknown tool calls or broken assistant/tool pairing;
- at least five consecutive failed tool calls;
- at least five failed calls and a failure rate of at least 50%.
Three weaker patterns are review flags, not automatic rejection:
- error count at or above the dataset-wide p99 threshold;
- at least eight failed calls, a failure rate of at least 20%, and errors in at least four of five trajectory regions;
- at least three errors with at least 60% of all errors in the first 20% of tool calls.
Length is separated from quality. Tokenized records are assigned to:
LE_81920: fast 80K training subset;81921_TO_131072: 128K training subset;131073_TO_262144: deferred long-context subset;GT_262144: excluded from the default small-model training run.
The summary reports how much of the total token mass is contributed by the longest 0.1%, 1%, and 5% of samples. This makes long-tail removal a measurable dataset decision rather than a guess.
Installation
cd /home/kxqandccx/kxq/tomlu/OpenSWETraces_cleanup
./.venv/bin/pip install -e '.[dev,tokens]'
tokenizers is optional. Without it, profiling still produces turn, character,
byte, and tool-error metrics, but no token bucket decision.
Full workflow
Use an explicit tokenizer already present on the machine:
TOKENIZER_JSON=/path/to/tokenizer.json
mkdir -p qc_outputs/deterministic_v3
swe-qc profile \
--input raw/Open-SWE-Traces \
--output qc_outputs/deterministic_v3/metrics.jsonl \
--errors qc_outputs/deterministic_v3/profile.errors.jsonl \
--tokenizer-json "$TOKENIZER_JSON" \
--workers 8 \
--resume
swe-qc summarize \
--metrics qc_outputs/deterministic_v3/metrics.jsonl \
--summary qc_outputs/deterministic_v3/summary.json \
--decisions qc_outputs/deterministic_v3/decisions.jsonl
swe-qc sample \
--input raw/Open-SWE-Traces \
--decisions qc_outputs/deterministic_v3/decisions.jsonl \
--output qc_outputs/deterministic_v3/review_sample.jsonl \
--per-group 20 \
--seed 20260818
profile --resume skips sample IDs already present in the metrics file. Writes
are append-only and flushed per record. summarize and sample atomically
replace their outputs so a partial file is never mistaken for a complete run.
Any dataset subdirectory containing Parquet shards is also a valid --input,
so the four trajectory families can be profiled in separate CPU processes and
their metrics concatenated before summarize.
Profile rows and decisions record the absolute source Parquet shard. sample
uses that provenance to read only shards containing selected IDs instead of
scanning the complete dataset.
Error detection boundary
The detector intentionally favors precision over recall. It recognizes explicit non-zero exit codes, test/build failure summaries, timeouts, permission errors, missing commands/files, and tool exceptions. It does not treat the mere word "error" as a failure, because tools often print source code or logs containing that word. Human samples should be used to refine patterns only when the raw tool observation provides an unambiguous signal.