Dashboardopencode-loreEvaluation launch, judging, and failure…

Evaluation launch, judging, and failure handling

Category: preference
Confidence: 1.00
ID: 01a0aeaf-5427-7f89-9221-7f2c3abb6ae2
Project ID: 6f4be9ff-ed84-4cca-a9e7-732a0b0b8677
Cross-project: No
Recalled in other projects: 0
Source session: 1YUpdJpSmTRNO1Oy9
Created: 2026-08-04 05:42:27
Updated: 2026-09-17 09:25:21

Content

Curate-log parsing must never abort an evaluation. Parse /lore:curate JSONL one non-empty line at a time, surface the first type:error record, ignore malformed lines, and catch outer failures. Determine score validity from primary artifacts such as result.json and session traces, not this auxiliary log. When JUDGE_API_KEY is configured, use the real external Anthropic provider with Sonnet as the default judge; never let a weak or non-Anthropic answering model grade its own output.

Move to: