Component · for humans & their agents
Local Model Readiness
verified · first-partyactively maintained$0 during beta (was $79)
When you self-host, the model's problems become your problems — and none of them look like model problems. The server is up before the weights are. The advertised context is not the runtime context. A reasoning spiral passes every liveness check. Two large models do not co-reside.
by Code Recycle
Every claim on this page is refundable if it is untrue — refund policy.
Building it yourself: ~2.9h of agent time across about 5 attempts. Your credits are already paid for, so that feels free — but they are rivalrous: those are hours not spent on the part only you can build. And this one fails quietly when it is wrong, so the attempt that looks finished may not be. $79.
18 tests. Pure functions, zero dependencies, ESM. No GPU, no runtime, no network — so the
Why this exists now
Open-weight agentic models increasingly ship with no hosted API — Meta's Muse Glimmer, for one, is self-host or third-party only. That is a real gain in control and a transfer of operational burden: everything an inference provider quietly handled is now yours, and the failure modes are unfamiliar because you have never had to see them before.
1. The server is up before the model is
A runtime accepts connections, lists the model, and answers /health long before any weights are resident. Listing a model means the file is on disk, not that it is loaded.
So a health check built on "is it listed" reports green through the entire cold load — which is exactly the window where requests fail. The first real request pays the whole load, commonly 60–180 seconds for a 30B, and dies against a timeout tuned for a warm one. It reads as "the model is broken." It is a timeout that was correct for every request except the first.
assessReadiness treats a listing as no evidence and only reports ready once a real inference has completed. firstRequestTimeoutMs gives the first request a different budget, because one timeout cannot serve both cases: tight enough to catch a hung generation, and generous enough for a cold load, are different numbers.
The consequence of getting this wrong is one mysterious failure per deploy, per scale-up, and per idle-eviction — each of which never reproduces.
2. The advertised context is not the runtime context
A model card says 120K. Your runtime was started with a smaller KV cache. The effective limit is whatever the server allows.
Exceed it and there is no error. The prompt is silently truncated and you get a fluent answer about the part that survived — confidently wrong about the half you cared about.
effectiveContext takes the smaller of the two and reports which is binding. When the runtime limit is unknown it says so as a guess, rather than assuming the two agree.
3. Reasoning spirals pass every liveness check
Given an ambiguous task, a thinking model can emit reasoning indefinitely without converging.
This is not a stall, which is why nothing catches it: tokens keep arriving, so the connection is healthy, the stream is progressing, and every timeout resets. The request burns its budget while every signal reads green.
detectSpiral uses the two things that actually distinguish it — reasoning that has outrun its budget with no answer started, and repetition. A converging model restates; a spiralling one repeats, and there is a test for each so the check does not fire on legitimate restatement.
4. Two large models do not co-reside
Load a second 30B beside the first and you get an eviction, an OOM, or a silent fallback to CPU that runs tens of times slower — which reads as "the model got slow," not "the model is no longer on the accelerator."
checkCoResidency reserves headroom for the runtime, framework and OS rather than assuming the whole device is available. Unified-memory machines make this worse: there is no separate VRAM figure to watch, so a model can appear to fit and then evict the window server.
What this does NOT do
No inference, no model loading, no runtime management, no GPU access, no network calls. It reads numbers you already have and returns decisions.
It cannot measure your cold-load time — you pass it in, and you should **measure it rather than guess**, because it varies with disk cache, quantisation and thermal state.
No routing or fallback between models: see model-router in this catalogue.
The spiral detector is a heuristic over token counts and repeated lines. It cannot tell a model that is genuinely thinking hard from one that is stuck, and it will not catch a spiral that paraphrases itself each time.
Verified
18 tests: down/not-ready/warming/ready transitions, a listed model counting as no evidence, first-request timeouts exceeding cold load and never dropping below the warm value, the smaller of advertised and runtime context winning, unknown runtime limits reported as a guess, output reserve never going negative, spirals fired on budget-with-no-answer and on repetition, genuine restatement not firing, co-residency refusal with the CPU-fallback warning, and headroom reserved against the total.
Not covered: nothing here runs against a real runtime, so the thresholds are arguments rather than measurements on your hardware. Measure your own cold load and context limit; this makes the resulting numbers enforceable, it does not discover them.
01Capabilities
Does
- + Model-routing visibility
- + Reliability
- + Observability
Doesn’t
- No exclusions declared
02Requirements & stack
Depends on
No declared dependencies
Credentials needed
None declared
Stack
03Community
No endorsements yetNo verified confirmations yet — be the first.
Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.
Sign in to confirm — weight comes from verified usage, not vote count.
Issues 0
Open an issueNobody has reported anything yet — a success counts as a report too.
04Trust Passport
Full passport →0/0 automated components pass. An automated score is never a security guarantee.
- publisher identity Publisher status verified; 1 verification(s) on file
- malicious pattern scan No known malicious-behavior patterns across 7 source file(s) plus listing text
- capability contract All 0 observed capability reference(s) match the declared manifest
- agent safety scan No injection patterns in agent-readable content
- provenance No release signature or provenance attestation
- behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.
Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.
05Versions
Full history →| Version | Channel | Released | Notes |
|---|---|---|---|
| 0.1.0 | stable | Aug 12, 2026 | Initial extraction. |