Add API retries and concurrent batch processing
This commit is contained in:
@@ -127,13 +127,18 @@ Optional settings and their defaults:
|
||||
|
||||
```bash
|
||||
export GLM_TIMEOUT_SECONDS=300
|
||||
export GLM_MAX_RETRIES=3
|
||||
export GLM_MAX_RETRIES=5
|
||||
export GLM_MAX_TOKENS=8192
|
||||
export GLM_TEMPERATURE=0.0
|
||||
export GLM_REASONING_EFFORT=high
|
||||
export GLM_THINKING_ENABLED=true
|
||||
```
|
||||
|
||||
`GLM_MAX_RETRIES=5` means one initial request plus at most five retries. The
|
||||
client retries timeouts, connection failures, HTTP 429/5xx responses, malformed
|
||||
JSON, and schema-invalid model output with bounded exponential backoff. HTTP
|
||||
401/403 authentication failures are never retried.
|
||||
|
||||
If the gateway rejects GLM-specific `thinking` or `reasoning_effort` fields, the
|
||||
client automatically retries with the portable OpenAI-compatible request subset.
|
||||
|
||||
@@ -209,9 +214,15 @@ swe-qc audit \
|
||||
--input samples/sample_20_seed_20260805.jsonl \
|
||||
--output qc_outputs/sample20.audits.jsonl \
|
||||
--errors qc_outputs/sample20.audit.errors.jsonl \
|
||||
--workers 20 \
|
||||
--resume
|
||||
```
|
||||
|
||||
`--workers` bounds the number of records processed concurrently. JSONL writes
|
||||
remain serialized in the main thread, so each completed record is appended
|
||||
atomically even when API requests run in parallel. Output order follows request
|
||||
completion order; `sample_id` remains the stable join key.
|
||||
|
||||
The deterministic score combines weighted process dimensions with penalties
|
||||
for minor, major, and critical behavior issues. Failed exploratory calls are
|
||||
not penalized when the agent interprets them correctly and recovers.
|
||||
|
||||
Reference in New Issue
Block a user