Dashboard › opencode-lore › LocalProvider embedding capacity, batch…
01a0af08-17e5-76d3-9fa0-9a31601c32dcLocalProvider uses a persistent, memory-gated ONNX pool; each worker costs about 680 MB. Chose lazy, demand-driven growth over eager workers because host freemem() overstates container capacity. Use one immutable tagged memory snapshot per admission; explicit unconstrained state may use host memory, while unknown headroom admits no extra worker and floors the primary worker’s cap. Count active, constructing, retiring, and replacement generations against capacity, reserving a full model budget until a generation completes a real embed. Worker 0 always remains primary; extra workers require concurrent single-query demand. Never replace the last slot before confirmed retirement. Release admission in finally; provider loss, OOM, timeout, rejection, or cancellation must not wedge future work.