Initial Open-SWE-Traces cleanup pipeline

This commit is contained in:
2026-08-06 22:53:47 +08:00
commit 044bd03f0e
35 changed files with 3638 additions and 0 deletions
+179
View File
@@ -0,0 +1,179 @@
# GLM-5.2 Sample-20 Classification and Repair Audit
Date: 2026-08-05
Dataset: `nvidia/Open-SWE-Traces`
Sample: uniform random sample of 20 trajectories, seed `20260805`
## Executive result
GLM-5.2 classified all 20 records. The direct full-SFT decision exactly matched
the existing manual review: samples 1, 12, 13, 15, 17, and 18 were accepted,
for a total of 6/20. There were no false rejects relative to the manual
keep-candidate set and no additional false-positive keeps.
Truncation cannot turn any of the other 14 records into a complete successful
trajectory because it cannot create missing task-resolution evidence. The
count of additional `SFT_FULL` records obtainable by static truncation is
therefore 0/20.
For step-level SFT, the current GLM repair planner has poor recall. Both the
original repair rubric and the revised explicit step-salvage rubric produced
zero `CREATE_STEP_EXAMPLE` proposals. A conservative human audit identified
four high-confidence step-only candidates and four borderline candidates.
## Classification distribution
| Classification | Count |
|---|---:|
| `ACCEPT_SILVER_POSITIVE / SFT_FULL` | 6 |
| `ACCEPT_NEGATIVE / DPO_REJECTED` | 7 |
| `HOLD_UNVERIFIED / HOLD` | 5 |
| `REJECT / DROP` | 2 |
The source outcomes in the sample were six `resolved=1`, seven `resolved=0`,
and seven `resolved=-1`. All six `resolved=1` records were accepted and all
other records were prevented from becoming full-SFT positives.
## Comparison with the prior manual review
| Manual group | Samples | GLM result |
|---|---|---|
| Keep candidate | 1, 12, 13, 15, 17, 18 | All six accepted as `SFT_FULL` |
| Repair/replay | 7, 9, 10, 14 | Three held/rejected and one dropped; none repaired |
| Reject as positive | 2, 3, 4, 5, 6, 8, 11, 16, 19, 20 | None accepted as positive |
The classification threshold is not too strict for full positive trajectories
on this sample: keep-set agreement is 20/20 and keep precision/recall against
the prior manual labels are both 6/6. The pipeline is, however, too strict at
recovering training value from repair/replay records.
## Repair proposal v1 audit
The first repair pass returned 20 schema-valid plans:
| Repair decision | Count |
|---|---:|
| `NO_CHANGE` | 8 |
| `REQUIRES_EXECUTION` | 8 |
| `APPLY_STATIC_REPAIR` | 2 |
| `DROP` | 2 |
It proposed four `REWRITE_FINAL_SUMMARY` operations and no truncations.
The proposal set was not reliably executable:
1. Sample 3 had a syntactically executable summary rewrite, but it would alter
an authentic rejected response. It adds no positive-SFT value and weakens
the original DPO negative signal.
2. Sample 4 used an object as the replacement for `REWRITE_FINAL_SUMMARY`,
while the deterministic applier requires a string. Application failed with
`PolicyViolation`.
3. Samples 6 and 14 used `REQUIRES_EXECUTION` while also including mutation
operations. The applier does not authorize mutation for that decision.
Only one of the two `APPLY_STATIC_REPAIR` plans executed, and that successful
mutation remained `DPO_REJECTED`. Consequently v1 produced zero additional SFT
records.
## Repair proposal v2 audit
The rubric was revised to define strict step salvage and local policy was
strengthened to require:
- no operations for `NO_CHANGE`, `REQUIRES_EXECUTION`, or `DROP`;
- no rewriting of authentic `DPO_REJECTED` trajectories;
- a complete assistant replacement plus truncation for every step example;
- operation-specific replacement types;
- `SFT_STEP_ONLY` as the maximum use for truncated examples.
Seventeen policy tests plus Ruff checks passed after the change.
GLM returned 18 policy-valid v2 plans and two locally rejected plans. Samples 3
and 8 again tried to rewrite authentic DPO negatives and were rejected by the
new deterministic guard. The 18 valid plans contained 11 `NO_CHANGE`, five
`REQUIRES_EXECUTION`, and two `DROP` decisions. They still contained zero
truncation proposals.
This is safer than v1, but it confirms that prompt wording alone did not give
GLM adequate recall for step-only salvage.
## Truncation eligibility
### Complete trajectory SFT
Additional records repairable to `SFT_FULL` through truncation: **0/20**.
A truncated failed or unverified trajectory lacks a verified terminal solution.
Treating it as full-SFT would convert absence of evidence into a success label.
### Conservative step-only SFT
High-confidence human-audited candidates: **4/20**.
| Sample | Proposed cutoff region | Corrected next-step target | Why it is statically defensible |
|---:|---|---|---|
| 4 | After the observed no-match edge-case failure around turns 186-190 | Continue debugging the `undefined` result instead of discarding the failing case and claiming completion | The failure is present in the tool output before the corrected action |
| 6 | Before the unsupported completion claim at turn 181 | Inspect and implement the explicitly requested missing `MockRequest` scope, or state that the task remains incomplete | The missing scope appears in the user request, not only in the reference patch |
| 7 | Before the first prohibited `wrap_test.go` edit at turn 171 | Preserve the test file and inspect/refactor the production wrapper API so existing tests compile | The user explicitly prohibited test changes and the failing compiler output is already visible |
| 9 | Before the first prohibited `tests/integration/library.rs` edit at turn 95 | Keep tests unchanged and repair the production API/constructor path first | The no-test-edit constraint is explicit and visible before the bad action |
These candidates may be used only as prefix-plus-corrected-next-action examples.
They must end at the corrected assistant/tool-call turn and contain no invented
tool result.
Borderline candidates requiring human confirmation: samples 2, 3, 8, and 10.
- Sample 2 created an issue reproduction but did not integrate it into the
rust-analyzer test harness before moving on. A useful next step exists, but
the standalone Rust file itself does not exercise the analyzer panic.
- Samples 3 and 8 have explicit failing test output and can teach failure
recovery, but several following turns are partially reasonable, so the exact
first irrecoverable assistant turn requires closer annotation.
- Sample 10 can be cut before the prohibited test edit at turn 85, but the
production patch is already semantically wrong for the target macOS branch,
making the retained context questionable for SFT.
Under a conservative policy, use the confirmed count of four. Under a broader
error-recovery curriculum, the upper candidate count is eight, but the four
borderline cases should not be admitted automatically.
## Model errors and reliability findings
1. Sample 16 is `resolved=-1`, but GLM emitted the hard-fail code
`RESOLVED_ZERO_FOR_POSITIVE`. Sample 5 is `resolved=0`, but one pass omitted
that code. Local outcome mapping prevented an incorrect training upgrade,
but hard-fail code semantics need deterministic validation.
2. Classification quality was materially better than repair-generation
quality. The keep/reject boundary matched the manual review, while repair
plans contained type errors, decision/operation contradictions, and zero
truncation recall.
3. Long reasoning caused repeated gateway 502 responses. Turn-preserving
evidence compaction plus a low-latency retry completed all 20
classifications. API transport failures must remain separate from QC
decisions.
4. A JSON-schema-valid plan is not sufficient. Operation-specific local policy
and deterministic dry-run are required before any mutation.
## Recommended production decision
- Admit the six accepted records to the silver `SFT_FULL` pool.
- Preserve the seven explicit negatives unchanged for DPO/error analysis.
- Drop samples 10 and 16 from positive/step training in their current form.
- Keep five unverified records in the execution-required pool.
- Create step-level candidates only from the four confirmed truncation cases,
then independently review their exact replacement messages and dry-run the
deterministic applier.
- Do not use GLM repair proposals without local policy enforcement and a second
reviewer.
## Output files
All files are under `outputs/glm52_sample20_20260805/`:
- `classifications.jsonl`: 20 completed classifications
- `classification_errors_attempt1.jsonl`: archived gateway failures
- `repair_plans.jsonl`: 20 v1 repair plans
- `repair_plans_v2.jsonl`: 18 policy-valid v2 plans
- `repair_plan_v2_errors.jsonl`: two policy-rejected v2 proposals
- `repaired_static.jsonl`: one v1 dry-run mutation
- `apply_repair_errors.jsonl`: one v1 application failure