Files
OpenSWETraces_cleanup/README.md
T

365 lines
13 KiB
Markdown

# SWE Data Processing
`swe-data-processing` is a conservative, auditable Python pipeline for cleaning
[`nvidia/Open-SWE-Traces`](https://huggingface.co/datasets/nvidia/Open-SWE-Traces)
before supervised fine-tuning of smaller coding agents.
The project is designed for a restricted environment in which repository
containers and agent sandboxes are unavailable. It combines deterministic local
checks with GLM-5.2 API judgments. It does **not** claim that static inspection can
prove code correctness. Instead, it separates internally consistent silver
positives, useful negatives, safely repairable formatting issues, unverified
records, and corrupted records.
## Dataset snapshot
The current local snapshot contains 207,489 trajectories over 22,320 unique
issues. The immutable dataset outcome field is named `resolved`:
| `resolved` | Meaning | Count | Percentage |
|---:|---|---:|---:|
| `1` | Successful candidate | 65,244 | 31.44% |
| `0` | Explicit failure | 95,487 | 46.02% |
| `-1` | Unknown outcome | 46,758 | 22.54% |
Records must be split by `instance_id`, not by trajectory, to prevent the same
issue from leaking across train and evaluation sets.
## Safety model
Static cleanup may improve representation quality, but it may not create new
execution facts. The implementation enforces these invariants locally after
every model response:
1. `resolved` is immutable.
2. A `resolved=0` record can never become a full successful SFT example.
3. A `resolved=-1` record cannot become `SFT_FULL` without execution.
4. Existing tool observations cannot be rewritten or fabricated.
5. Model and reference patches are immutable during static repair.
6. Truncated records have a maximum use of `SFT_STEP_ONLY`.
7. A silver positive requires reliable post-edit verification, no hard failures,
and a passing value for every QC dimension.
8. Every repair requires a separate review API call.
The API model proposes decisions and repair plans. Deterministic Python code
validates schemas, enforces outcome policy, applies only allowlisted edits, and
records provenance.
Long trajectories are sent to the API through a turn-preserving evidence view.
Test observations, state-changing turns, malformed calls, final turns, patches,
and stable turn IDs receive priority. Shortened values include their original
character count and SHA-256 hash. This compaction affects only the API prompt;
classification never rewrites source JSONL or Parquet records.
## API endpoint
The configured gateway is OpenAI chat-completions compatible:
```text
POST https://llm-api.cowin.run/v1/chat/completions
```
`/v1/text-completion` and `/v1/text-completions` resolve to the gateway's web
console rather than an inference API, so they are not used.
The default model is `glm-5.2`. The API key is read only from `GLM_API_KEY`. Do
not write a key into source code, command history, output manifests, or this
README.
## Project layout
```text
swe_data_processing/
├── pyproject.toml
├── README.md
├── changelog.md # Versioned strategy changes
├── .env.example
├── src/swe_data_processing/
│ ├── cli.py # Command-line entry points
│ ├── client.py # GLM API client and JSON validation
│ ├── config.py # Environment-only runtime settings
│ ├── evidence.py # Prompt-only evidence compaction
│ ├── features.py # Deterministic static evidence extraction
│ ├── io.py # Streaming Parquet and JSONL readers
│ ├── policy.py # Immutable local policy guards
│ ├── repair.py # Allowlisted deterministic repair application
│ ├── resources.py # Packaged prompt/schema loading
│ ├── workflow.py # Classify, plan repair, and review stages
│ ├── prompts/ # Version-controlled GLM system prompts
│ └── schemas/ # JSON Schemas for every API stage
├── tests/ # Offline unit tests; no network calls
├── scripts/ # Existing sampling and profiling utilities
├── raw/Open-SWE-Traces/ # Downloaded dataset; ignored by Git
├── samples/ # Human-review samples
├── reports/ # Dataset reports and QC rubrics
└── qc_outputs/ # Generated manifests; ignored by Git
```
## Installation
The existing remote virtual environment can install the package in editable
mode:
```bash
cd /mnt/beegfs/yi/swe_data_processing
./.venv/bin/pip install -e '.[dev]'
```
For a fresh environment:
```bash
python3 -m venv .venv
./.venv/bin/pip install --upgrade pip
./.venv/bin/pip install -e '.[dev]'
```
## Configuration
Export credentials in the shell that launches the pipeline:
```bash
export GLM_API_KEY='your-runtime-secret'
export GLM_API_BASE='https://llm-api.cowin.run'
export GLM_API_PATH='/v1/chat/completions'
export GLM_MODEL='glm-5.2'
```
Optional settings and their defaults:
```bash
export GLM_TIMEOUT_SECONDS=300
export GLM_MAX_RETRIES=5
export GLM_MAX_TOKENS=8192
export GLM_TEMPERATURE=0.0
export GLM_REASONING_EFFORT=high
export GLM_THINKING_ENABLED=true
```
`GLM_MAX_RETRIES=5` means one initial request plus at most five retries. The
client retries timeouts, connection failures, HTTP 429/5xx responses, malformed
JSON, and schema-invalid model output with bounded exponential backoff. HTTP
401/403 authentication failures are never retried.
If the gateway rejects GLM-specific `thinking` or `reasoning_effort` fields, the
client automatically retries with the portable OpenAI-compatible request subset.
## Commands
### 1. Verify API authentication and structured output
```bash
swe-qc smoke-test
```
The command prints the endpoint, model, request ID, usage, and a tiny validated
JSON response. It never prints the API key.
### 2. Extract deterministic evidence without API calls
From the 20-record review sample:
```bash
mkdir -p qc_outputs
swe-qc features \
--input samples/sample_20_seed_20260805.jsonl \
--output qc_outputs/sample20.features.jsonl \
--errors qc_outputs/sample20.features.errors.jsonl
```
From all Parquet shards:
```bash
swe-qc features \
--input raw/Open-SWE-Traces \
--output qc_outputs/all.features.jsonl \
--errors qc_outputs/all.features.errors.jsonl \
--resume
```
Features include malformed tool arguments, unknown tools, role alternation,
state-changing turns, post-edit test evidence, masked shell pipelines, patch file
sets, patch size ratios, and explicit user constraints.
### 3. Classify trajectories through GLM-5.2
Run a small pilot first:
```bash
swe-qc classify \
--input samples/sample_20_seed_20260805.jsonl \
--output qc_outputs/sample20.classifications.jsonl \
--errors qc_outputs/sample20.classification.errors.jsonl \
--limit 20 \
--resume
```
The classifier returns one of:
- `ACCEPT_SILVER_POSITIVE`
- `ACCEPT_NEGATIVE`
- `STATIC_REPAIR`
- `HOLD_UNVERIFIED`
- `REJECT`
Training use is tracked separately as `SFT_FULL`, `SFT_STEP_ONLY`,
`DPO_REJECTED`, `ERROR_ANALYSIS`, `HOLD`, or `DROP`.
### 4. Locate a safe prefix and score it independently
`audit` uses two isolated GLM calls for failed and unknown trajectories:
1. The boundary call receives the trajectory as roughly 32-turn blocks without
splitting an assistant action from its immediate tool result. It cannot see
`resolved` or reference patches, but it does see the trajectory's own final
model patch to check task coverage and diff pollution. It selects the earliest
unrepaired major/critical assistant turn, keeps a safe full process trace,
or returns `HOLD` when neither decision is supported.
2. Python slices the exact messages before that assistant turn. The quality call
receives only this materialized prefix, so suffix behavior cannot affect its
score.
For a complete trajectory, the quality call also sees its final model patch. A
truncated prefix never receives that suffix-derived patch. State-changing turns
are recorded as audit telemetry but are not automatic boundaries: writing code
is not itself an error, and recovered experiments are useful process data. The
effective boundary is the earliest unrecovered major or critical behavior found
by the boundary stage. The isolated quality stage rejects a retained prefix that
still contains an unrecovered severe problem.
Successful trajectories skip boundary selection and are scored as complete
trajectories. The command remains simple:
```bash
swe-qc audit \
--input samples/sample_20_seed_20260805.jsonl \
--output qc_outputs/sample20.audits.jsonl \
--errors qc_outputs/sample20.audit.errors.jsonl \
--workers 5 \
--resume
```
`--workers` bounds the number of records processed concurrently. JSONL writes
remain serialized in the main thread, so each completed record is appended
atomically even when API requests run in parallel. Output order follows request
completion order; `sample_id` remains the stable join key.
The gateway previously returned frequent HTTP 429 responses at 20 workers, so
five workers is the practical default for long runs unless the service limit is
raised.
The quality call lists concrete erroneous and inefficient assistant actions and
scores five 0-20 dimensions: planning, investigation, tool use and observation,
progress, and clarity/efficiency. Python recomputes issue counts, sums the five
dimensions, and assigns:
- `HIGH`: 80-100
- `MEDIUM`: 60-79
- `LOW`: 40-59
- `REJECT`: below 40, or an unrepaired major/critical problem remains
Incomplete prefixes can still be useful. Normal failed experiments are retained
when the agent later diagnoses, repairs, and verifies them. A prefix ending in a
complete tool result is valid process-SFT structure.
### 5. Create static repair plans
Only classifications that explicitly return `STATIC_REPAIR` are processed by
default:
```bash
swe-qc repair-plan \
--input samples/sample_20_seed_20260805.jsonl \
--classifications qc_outputs/sample20.classifications.jsonl \
--output qc_outputs/sample20.repair-plans.jsonl \
--errors qc_outputs/sample20.repair-plan.errors.jsonl \
--resume
```
The model may propose only operations from the allowlist. Code changes, test
result synthesis, patch replacement, and outcome upgrades are prohibited.
### 6. Apply plans deterministically
```bash
swe-qc apply-repair \
--input samples/sample_20_seed_20260805.jsonl \
--plans qc_outputs/sample20.repair-plans.jsonl \
--output qc_outputs/sample20.repaired.jsonl \
--errors qc_outputs/sample20.apply.errors.jsonl \
--resume
```
The applier checks target turns, read-only pair deletion, immutable patch fields,
and retained tool-output hashes. It emits a structured diff with input and output
SHA-256 hashes.
### 7. Independently review repairs
```bash
swe-qc review \
--input samples/sample_20_seed_20260805.jsonl \
--classifications qc_outputs/sample20.classifications.jsonl \
--plans qc_outputs/sample20.repair-plans.jsonl \
--repaired qc_outputs/sample20.repaired.jsonl \
--output qc_outputs/sample20.reviews.jsonl \
--errors qc_outputs/sample20.review.errors.jsonl \
--resume
```
Approval means that a static repair preserved evidence and structure. It does not
mean that the code patch was executed or proved correct.
## Output and resume behavior
Pipeline manifests are append-only JSONL. `--resume` reads completed sample IDs
from the output and skips them. Each API response includes non-sensitive
provenance:
- model and endpoint;
- gateway request ID;
- token usage when available;
- input SHA-256;
- compatibility-fallback flag;
- UTC creation time.
Errors are written as compact records containing sample ID, stage, exception
type, and a bounded message. Raw prompts and credentials are not copied into
error logs.
## Recommended rollout
1. Run offline tests.
2. Run `smoke-test` once.
3. Run classification on the 20 manually reviewed records.
4. Compare GLM decisions with the human labels.
5. Build a 2,000-record stratified calibration set.
6. Require at least 95% precision for `ACCEPT_SILVER_POSITIVE` before scaling.
7. Keep `resolved=0` as negative data rather than attempting to turn it into
positive trajectories.
8. Keep `resolved=-1` in an unverified manifest unless it contains explicit,
reliable failure evidence.
## Development
Run all offline tests:
```bash
./.venv/bin/pytest
```
Run lint checks:
```bash
./.venv/bin/ruff check src tests
```
Tests use an `httpx.MockTransport` and never contact the GLM endpoint.
## Credential handling
- `.env` files are ignored by Git.
- The provided API key is not stored anywhere in this project.
- Use a secret manager or a protected runtime environment variable for batch
jobs.
- Rotate the key if it has been copied into a public log, issue, or repository.