Files
OpenSWETraces_cleanup/README.md
T

5.4 KiB

Open-SWE-Traces Deterministic Profiler

This project profiles nvidia/Open-SWE-Traces for coding-agent SFT without an LLM judge, repository sandbox, or trajectory rewriting. It never claims that a patch is correct. It records facts that can be reproduced from the dataset and uses a small set of explicit heuristics to remove traces that are definitely broken or unsuitable for a chosen training context.

The local dataset snapshot contains 207,489 trajectories over 22,320 issues:

resolved Meaning Count
1 Externally marked successful 65,244
0 Externally marked failed 95,487
-1 Outcome unknown 46,758

resolved is never changed. Only resolved=1 is eligible for successful SFT; the profiler can still describe failed and unknown traces for analysis. Final decisions label resolved=0 as EXCLUDE_FAILED_OUTCOME and resolved=-1 as HOLD_UNVERIFIED_OUTCOME rather than recommending them for training.

Design

The pipeline has three commands:

  1. profile streams the original JSONL or Parquet dataset and writes one deterministic metrics row per trajectory.
  2. summarize computes dataset quantiles and writes transparent decisions.
  3. sample selects full trajectories from every rule and length bucket for human validation.

No command edits, truncates, repairs, or invents trajectory turns. Source data and generated files are joined by trajectory_id/sample_id.

Metrics

Each profile row contains:

  • turn and role counts;
  • canonical serialized characters and UTF-8 bytes;
  • exact token count when a tokenizer.json is supplied;
  • tool-call count and count by tool name;
  • failed tool-call count and rate;
  • observable error type and tool distributions;
  • normalized error positions across five equal trajectory regions;
  • early-error count and fraction;
  • longest consecutive failed-tool-call run;
  • malformed calls, unknown tools, missing results, and orphan results.

The canonical token stream is compact, sorted JSON containing only tools and trajectory. This is reproducible but is not presented as a universal chat template. If training uses another serializer or tokenizer, rerun profile with that exact tokenizer instead of converting characters to fake token counts.

Heuristics

The following conditions are hard rejections because they represent broken training structure or an unambiguous repeated-failure pattern:

  • malformed/unknown tool calls or broken assistant/tool pairing;
  • at least five consecutive failed tool calls;
  • at least five failed calls and a failure rate of at least 50%.

Three weaker patterns are review flags, not automatic rejection:

  • error count at or above the dataset-wide p99 threshold;
  • at least eight failed calls, a failure rate of at least 20%, and errors in at least four of five trajectory regions;
  • at least three errors with at least 60% of all errors in the first 20% of tool calls.

Length is separated from quality. Tokenized records are assigned to:

  • LE_81920: fast 80K training subset;
  • 81921_TO_131072: 128K training subset;
  • 131073_TO_262144: deferred long-context subset;
  • GT_262144: excluded from the default small-model training run.

The summary reports how much of the total token mass is contributed by the longest 0.1%, 1%, and 5% of samples. This makes long-tail removal a measurable dataset decision rather than a guess.

Installation

cd /home/kxqandccx/kxq/tomlu/OpenSWETraces_cleanup
./.venv/bin/pip install -e '.[dev,tokens]'

tokenizers is optional. Without it, profiling still produces turn, character, byte, and tool-error metrics, but no token bucket decision.

Full workflow

Use an explicit tokenizer already present on the machine:

TOKENIZER_JSON=/path/to/tokenizer.json
mkdir -p qc_outputs/deterministic_v3

swe-qc profile \
  --input raw/Open-SWE-Traces \
  --output qc_outputs/deterministic_v3/metrics.jsonl \
  --errors qc_outputs/deterministic_v3/profile.errors.jsonl \
  --tokenizer-json "$TOKENIZER_JSON" \
  --workers 8 \
  --resume

swe-qc summarize \
  --metrics qc_outputs/deterministic_v3/metrics.jsonl \
  --summary qc_outputs/deterministic_v3/summary.json \
  --decisions qc_outputs/deterministic_v3/decisions.jsonl

swe-qc sample \
  --input raw/Open-SWE-Traces \
  --decisions qc_outputs/deterministic_v3/decisions.jsonl \
  --output qc_outputs/deterministic_v3/review_sample.jsonl \
  --per-group 20 \
  --seed 20260818

profile --resume skips sample IDs already present in the metrics file. Writes are append-only and flushed per record. summarize and sample atomically replace their outputs so a partial file is never mistaken for a complete run. Any dataset subdirectory containing Parquet shards is also a valid --input, so the four trajectory families can be profiled in separate CPU processes and their metrics concatenated before summarize. Profile rows and decisions record the absolute source Parquet shard. sample uses that provenance to read only shards containing selected IDs instead of scanning the complete dataset.

Error detection boundary

The detector intentionally favors precision over recall. It recognizes explicit non-zero exit codes, test/build failure summaries, timeouts, permission errors, missing commands/files, and tool exceptions. It does not treat the mere word "error" as a failure, because tools often print source code or logs containing that word. Human samples should be used to refine patterns only when the raw tool observation provides an unambiguous signal.