Add API retries and concurrent batch processing

This commit is contained in:
2026-08-06 23:21:46 +08:00
parent 044bd03f0e
commit f1a090e8c4
8 changed files with 183 additions and 20 deletions
+12 -1
View File
@@ -127,13 +127,18 @@ Optional settings and their defaults:
```bash
export GLM_TIMEOUT_SECONDS=300
export GLM_MAX_RETRIES=3
export GLM_MAX_RETRIES=5
export GLM_MAX_TOKENS=8192
export GLM_TEMPERATURE=0.0
export GLM_REASONING_EFFORT=high
export GLM_THINKING_ENABLED=true
```
`GLM_MAX_RETRIES=5` means one initial request plus at most five retries. The
client retries timeouts, connection failures, HTTP 429/5xx responses, malformed
JSON, and schema-invalid model output with bounded exponential backoff. HTTP
401/403 authentication failures are never retried.
If the gateway rejects GLM-specific `thinking` or `reasoning_effort` fields, the
client automatically retries with the portable OpenAI-compatible request subset.
@@ -209,9 +214,15 @@ swe-qc audit \
--input samples/sample_20_seed_20260805.jsonl \
--output qc_outputs/sample20.audits.jsonl \
--errors qc_outputs/sample20.audit.errors.jsonl \
--workers 20 \
--resume
```
`--workers` bounds the number of records processed concurrently. JSONL writes
remain serialized in the main thread, so each completed record is appended
atomically even when API requests run in parallel. Output order follows request
completion order; `sample_id` remains the stable join key.
The deterministic score combines weighted process dimensions with penalties
for minor, major, and critical behavior issues. Failed exploratory calls are
not penalized when the agent interprets them correctly and recovers.