Files
OpenSWETraces_cleanup/reports/glm52_sample20_qc_audit_20260805.md

8.5 KiB

GLM-5.2 Sample-20 Classification and Repair Audit

Date: 2026-08-05
Dataset: nvidia/Open-SWE-Traces
Sample: uniform random sample of 20 trajectories, seed 20260805

Executive result

GLM-5.2 classified all 20 records. The direct full-SFT decision exactly matched the existing manual review: samples 1, 12, 13, 15, 17, and 18 were accepted, for a total of 6/20. There were no false rejects relative to the manual keep-candidate set and no additional false-positive keeps.

Truncation cannot turn any of the other 14 records into a complete successful trajectory because it cannot create missing task-resolution evidence. The count of additional SFT_FULL records obtainable by static truncation is therefore 0/20.

For step-level SFT, the current GLM repair planner has poor recall. Both the original repair rubric and the revised explicit step-salvage rubric produced zero CREATE_STEP_EXAMPLE proposals. A conservative human audit identified four high-confidence step-only candidates and four borderline candidates.

Classification distribution

Classification Count
ACCEPT_SILVER_POSITIVE / SFT_FULL 6
ACCEPT_NEGATIVE / DPO_REJECTED 7
HOLD_UNVERIFIED / HOLD 5
REJECT / DROP 2

The source outcomes in the sample were six resolved=1, seven resolved=0, and seven resolved=-1. All six resolved=1 records were accepted and all other records were prevented from becoming full-SFT positives.

Comparison with the prior manual review

Manual group Samples GLM result
Keep candidate 1, 12, 13, 15, 17, 18 All six accepted as SFT_FULL
Repair/replay 7, 9, 10, 14 Three held/rejected and one dropped; none repaired
Reject as positive 2, 3, 4, 5, 6, 8, 11, 16, 19, 20 None accepted as positive

The classification threshold is not too strict for full positive trajectories on this sample: keep-set agreement is 20/20 and keep precision/recall against the prior manual labels are both 6/6. The pipeline is, however, too strict at recovering training value from repair/replay records.

Repair proposal v1 audit

The first repair pass returned 20 schema-valid plans:

Repair decision Count
NO_CHANGE 8
REQUIRES_EXECUTION 8
APPLY_STATIC_REPAIR 2
DROP 2

It proposed four REWRITE_FINAL_SUMMARY operations and no truncations.

The proposal set was not reliably executable:

  1. Sample 3 had a syntactically executable summary rewrite, but it would alter an authentic rejected response. It adds no positive-SFT value and weakens the original DPO negative signal.
  2. Sample 4 used an object as the replacement for REWRITE_FINAL_SUMMARY, while the deterministic applier requires a string. Application failed with PolicyViolation.
  3. Samples 6 and 14 used REQUIRES_EXECUTION while also including mutation operations. The applier does not authorize mutation for that decision.

Only one of the two APPLY_STATIC_REPAIR plans executed, and that successful mutation remained DPO_REJECTED. Consequently v1 produced zero additional SFT records.

Repair proposal v2 audit

The rubric was revised to define strict step salvage and local policy was strengthened to require:

  • no operations for NO_CHANGE, REQUIRES_EXECUTION, or DROP;
  • no rewriting of authentic DPO_REJECTED trajectories;
  • a complete assistant replacement plus truncation for every step example;
  • operation-specific replacement types;
  • SFT_STEP_ONLY as the maximum use for truncated examples.

Seventeen policy tests plus Ruff checks passed after the change.

GLM returned 18 policy-valid v2 plans and two locally rejected plans. Samples 3 and 8 again tried to rewrite authentic DPO negatives and were rejected by the new deterministic guard. The 18 valid plans contained 11 NO_CHANGE, five REQUIRES_EXECUTION, and two DROP decisions. They still contained zero truncation proposals.

This is safer than v1, but it confirms that prompt wording alone did not give GLM adequate recall for step-only salvage.

Truncation eligibility

Complete trajectory SFT

Additional records repairable to SFT_FULL through truncation: 0/20.

A truncated failed or unverified trajectory lacks a verified terminal solution. Treating it as full-SFT would convert absence of evidence into a success label.

Conservative step-only SFT

High-confidence human-audited candidates: 4/20.

Sample Proposed cutoff region Corrected next-step target Why it is statically defensible
4 After the observed no-match edge-case failure around turns 186-190 Continue debugging the undefined result instead of discarding the failing case and claiming completion The failure is present in the tool output before the corrected action
6 Before the unsupported completion claim at turn 181 Inspect and implement the explicitly requested missing MockRequest scope, or state that the task remains incomplete The missing scope appears in the user request, not only in the reference patch
7 Before the first prohibited wrap_test.go edit at turn 171 Preserve the test file and inspect/refactor the production wrapper API so existing tests compile The user explicitly prohibited test changes and the failing compiler output is already visible
9 Before the first prohibited tests/integration/library.rs edit at turn 95 Keep tests unchanged and repair the production API/constructor path first The no-test-edit constraint is explicit and visible before the bad action

These candidates may be used only as prefix-plus-corrected-next-action examples. They must end at the corrected assistant/tool-call turn and contain no invented tool result.

Borderline candidates requiring human confirmation: samples 2, 3, 8, and 10.

  • Sample 2 created an issue reproduction but did not integrate it into the rust-analyzer test harness before moving on. A useful next step exists, but the standalone Rust file itself does not exercise the analyzer panic.
  • Samples 3 and 8 have explicit failing test output and can teach failure recovery, but several following turns are partially reasonable, so the exact first irrecoverable assistant turn requires closer annotation.
  • Sample 10 can be cut before the prohibited test edit at turn 85, but the production patch is already semantically wrong for the target macOS branch, making the retained context questionable for SFT.

Under a conservative policy, use the confirmed count of four. Under a broader error-recovery curriculum, the upper candidate count is eight, but the four borderline cases should not be admitted automatically.

Model errors and reliability findings

  1. Sample 16 is resolved=-1, but GLM emitted the hard-fail code RESOLVED_ZERO_FOR_POSITIVE. Sample 5 is resolved=0, but one pass omitted that code. Local outcome mapping prevented an incorrect training upgrade, but hard-fail code semantics need deterministic validation.
  2. Classification quality was materially better than repair-generation quality. The keep/reject boundary matched the manual review, while repair plans contained type errors, decision/operation contradictions, and zero truncation recall.
  3. Long reasoning caused repeated gateway 502 responses. Turn-preserving evidence compaction plus a low-latency retry completed all 20 classifications. API transport failures must remain separate from QC decisions.
  4. A JSON-schema-valid plan is not sufficient. Operation-specific local policy and deterministic dry-run are required before any mutation.
  • Admit the six accepted records to the silver SFT_FULL pool.
  • Preserve the seven explicit negatives unchanged for DPO/error analysis.
  • Drop samples 10 and 16 from positive/step training in their current form.
  • Keep five unverified records in the execution-required pool.
  • Create step-level candidates only from the four confirmed truncation cases, then independently review their exact replacement messages and dry-run the deterministic applier.
  • Do not use GLM repair proposals without local policy enforcement and a second reviewer.

Output files

All files are under outputs/glm52_sample20_20260805/:

  • classifications.jsonl: 20 completed classifications
  • classification_errors_attempt1.jsonl: archived gateway failures
  • repair_plans.jsonl: 20 v1 repair plans
  • repair_plans_v2.jsonl: 18 policy-valid v2 plans
  • repair_plan_v2_errors.jsonl: two policy-rejected v2 proposals
  • repaired_static.jsonl: one v1 dry-run mutation
  • apply_repair_errors.jsonl: one v1 application failure