Dashboardopencode-loreLocalProvider embedding capacity, batch…

LocalProvider embedding capacity, batching, admission, and lifecycle

Category: architecture
Confidence: 1.00
ID: 01a0af08-17e5-76d3-9fa0-9a31601c32dc
Project ID: 6f4be9ff-ed84-4cca-a9e7-732a0b0b8677
Cross-project: No
Recalled in other projects: 0
Source session: 0IiwozJmFR9A6QyFo
Created: 2026-08-25 21:49:22
Updated: 2026-09-17 11:02:18

Content

LocalProvider uses a persistent, memory-gated ONNX pool; each worker costs about 680 MB. Chose lazy, demand-driven growth over eager workers because host freemem() overstates container capacity. Use one immutable tagged memory snapshot per admission; explicit unconstrained state may use host memory, while unknown headroom admits no extra worker and floors the primary worker’s cap. Count active, constructing, retiring, and replacement generations against capacity, reserving a full model budget until a generation completes a real embed. Worker 0 always remains primary; extra workers require concurrent single-query demand. Never replace the last slot before confirmed retirement. Release admission in finally; provider loss, OOM, timeout, rejection, or cancellation must not wedge future work.

Move to: