158 lines
7.7 KiB
Markdown
158 lines
7.7 KiB
Markdown
# Changelog
|
|
|
|
This file records strategy changes that materially affect dataset decisions or
|
|
training-data semantics. Generated audit manifests are not treated as stable API
|
|
contracts.
|
|
|
|
## 3.1.0 - 2026-08-18
|
|
|
|
### Training-data materialization
|
|
|
|
- Added `materialize`, which streams the original Parquet shards and writes
|
|
two ModelScope/Hugging Face compatible splits at 131,072 and 81,920 tokens.
|
|
- Removed category balancing from dataset construction. Successful, failed,
|
|
and unknown outcomes retain their natural post-filter distribution.
|
|
- Both splits exclude deterministic hard rejects and review flags without
|
|
truncating, repairing, or otherwise rewriting trajectories.
|
|
- Preserved every source field and appended six `qc_*` columns for filter and
|
|
metric provenance.
|
|
- Materialization runs one source shard per process, writes atomic Zstandard
|
|
Parquet files, supports resume, and checks final row counts.
|
|
- The verified full run produced 186,665 rows in `context_131072` and 114,437
|
|
rows in `context_81920`; the shorter split is an exact subset of the longer.
|
|
|
|
## 3.0.1 - 2026-08-18
|
|
|
|
### CPU profiling
|
|
|
|
- Replaced directory-level Python thread concurrency with Parquet shard-level
|
|
processes so CPU-bound tokenization and metric extraction bypass the GIL.
|
|
- Each process streams batches of 16 records to an atomic shard part; the parent
|
|
merges completed parts in source order without loading them into memory.
|
|
- `profile --resume` now reuses completed shard parts, while interrupted
|
|
temporary parts are safely recomputed.
|
|
- Review sampling uses stored Parquet provenance and reads only shards that
|
|
contain selected trajectories.
|
|
|
|
## 3.0.0 - 2026-08-18
|
|
|
|
### Strategy replacement
|
|
|
|
- Removed the GLM client, prompts, schemas, model-based classification, boundary
|
|
selection, quality scoring, repair planning, and repair review.
|
|
- Replaced model judgments with deterministic length, token, tool-failure,
|
|
failure-position, structure, and streak metrics.
|
|
- Added explicit 80K, 128K, 256K, and over-256K context buckets.
|
|
- Added dataset-level quantiles, token-tail contribution statistics, p99 tool
|
|
error outlier detection, and stratified human-review sampling.
|
|
- Reserved automatic rejection for broken tool structure, five consecutive
|
|
failures, or a failure rate of at least 50%. Distributed failures, relative
|
|
error outliers, and early failure clusters require review.
|
|
- Token counts now require an explicit local `tokenizer.json`; the code never
|
|
substitutes a character-based approximation.
|
|
- Reduced the public CLI to `profile`, `summarize`, and `sample`.
|
|
|
|
### Motivation
|
|
|
|
- Multi-call LLM auditing was operationally unstable and could hallucinate safe
|
|
boundaries or miss incorrect final patches.
|
|
- Model-generated trajectory repair had no deterministic invariant capable of
|
|
proving that the rewritten trace remained correct.
|
|
- The new pipeline is reproducible, inspectable, CPU-only, and does not mutate
|
|
source trajectories.
|
|
|
|
## 2.1.0 - 2026-08-09
|
|
|
|
### Strategy correction
|
|
|
|
- Removed the unconditional cutoff before the first state-changing turn for
|
|
failed and unverified trajectories. A write or edit is not evidence of an
|
|
error by itself, and this cap discarded normal implementation, recovery, and
|
|
verification work.
|
|
- The semantic boundary now remains the earliest unrecovered major or critical
|
|
assistant behavior selected by the boundary stage. The isolated prefix-quality
|
|
stage remains the safety gate. The first stateful turn is retained as telemetry
|
|
for paired evaluation against the former policy.
|
|
|
|
## 2.0.0 - 2026-08-07
|
|
|
|
### Strategy
|
|
|
|
- Replaced the single-call trajectory audit with two isolated model calls:
|
|
boundary selection followed by prefix-only quality scoring.
|
|
- The boundary call receives an outcome-blind and reference-patch-blind view.
|
|
It does receive the trajectory's own final model patch to check coverage and
|
|
final diff pollution.
|
|
Long trajectories are presented as ordered blocks targeting 32 turns, with
|
|
absolute turn IDs and immediate assistant/tool-result pairs kept together.
|
|
- Python materializes the exact prefix before the selected assistant turn. The
|
|
quality call is then built again from that prefix, so it cannot see the
|
|
removed suffix, external outcome, model patch, reference patch, or boundary
|
|
rationale.
|
|
- Simplified both prompts and added concrete positive, recovery, and late-claim
|
|
examples.
|
|
- Replaced the weighted six-dimension score and issue penalties with five
|
|
equally weighted 0-20 dimensions. Python computes the final 0-100 score and
|
|
tier.
|
|
- Distinguished a safe full process trajectory (`KEEP_FULL`) from an uncertain
|
|
trajectory (`HOLD`) so the boundary stage does not reject useful data merely
|
|
because no truncation is needed.
|
|
- Kept an explicit list of erroneous and inefficient assistant actions as a
|
|
quality surrogate; Python recomputes its counts locally.
|
|
- Each listed issue records whether it was recovered. Python rejects a model
|
|
response that claims a prefix is safe while also listing an unrecovered
|
|
major/critical issue.
|
|
|
|
### Validation
|
|
|
|
- A boundary must be an assistant turn inside the selected block.
|
|
- Boundary evidence must include that assistant turn ID; Python attaches the
|
|
exact immutable source excerpt rather than trusting model-copied quotes.
|
|
- The model does not repeat a block ID or prefix-safe boolean. Python derives
|
|
the containing block from the validated boundary turn, while the isolated
|
|
second call judges the materialized prefix itself.
|
|
- `KEEP_FULL` now requires explicit review of task coverage, exact interfaces,
|
|
final repository state, user scope constraints, and claim/tool consistency;
|
|
passing tests alone are not treated as proof.
|
|
- Complete-trajectory scoring receives the final model patch, while truncated
|
|
prefix scoring remains suffix-isolated.
|
|
- Failed and unverified (`resolved=0/-1`) records are capped before the first
|
|
statically detected state-changing assistant turn. The earlier of that cap
|
|
and the GLM boundary is materialized, preferring shorter safe investigation
|
|
traces over late cuts.
|
|
- Removed redundant model-written prefix-validity booleans. Python derives
|
|
safety directly from unrecovered major/critical behavior issues.
|
|
- A prefix scorer cannot cite a turn outside the materialized prefix.
|
|
- `HIGH`, `MEDIUM`, `LOW`, and `REJECT` are computed locally rather than trusted
|
|
from model output.
|
|
|
|
### Stress pilot
|
|
|
|
- A difficulty-weighted 15-record set was used to expose late-boundary errors;
|
|
it is not an estimate of the full dataset acceptance rate.
|
|
- Raw semantic boundaries were often too late or incorrectly returned
|
|
`KEEP_FULL`. After applying the non-success stateful cap, two independent
|
|
human reviewers found all 15 effective prefixes safe and no late boundary.
|
|
- Retained-turn ratios ranged from 11.14% to 39.74%, with an 18.73% median.
|
|
The final quality distribution was 10 `HIGH` and 5 `MEDIUM`.
|
|
- The final low-reasoning, five-worker run completed 15/15 records without an
|
|
error in 90.05 seconds. High-reasoning mode was substantially slower and had
|
|
gateway timeouts, so it is not the practical batch default.
|
|
|
|
### Motivation
|
|
|
|
- In the previous 500-record pilot, 360 of 423 prefix candidates included
|
|
suffix behavior in the score despite prompt instructions to score only the
|
|
retained prefix.
|
|
- Manual stress review found useful conservative cutoffs, but also late cutoffs
|
|
that retained failed patches and normal experiments incorrectly treated as
|
|
permanent errors.
|
|
- Prompt wording alone could not guarantee suffix isolation, so the new design
|
|
enforces it in the data passed to each call.
|
|
|
|
## 1.0.0 - 2026-08-06
|
|
|
|
- Initial static classification, one-call trajectory audit, allowlisted repair
|
|
planning, deterministic repair application, and independent repair review.
|
|
- Added API retries and concurrent JSONL processing.
|