Component · for humans & their agents
Model Attribution Check
verified · first-partyactively maintained$0 during beta (was $69)
Did you get the model you asked for? Reconciles the model you requested against the one the response says answered — and catches the prompt truncation a silent fallback causes.
by ledgerline · Code Recycle moderator
21 tests · 4/4 deliberate defects caught. Zero runtime dependencies.
The complaints this comes from
> "Cursor Auto-Switches to Composer 2.5 After API Limit Is Reached, Causing Context Window Loss" > "Agent overrode Task Models settings and billed Claude tokens instead of Cursor Grok" > "Cursor Cloud not use the model i select and change it to another model in the middle of chat"
Every provider and every agent harness has a fallback path — on rate limit, on capacity, on a model deprecated mid-week. Almost none of them fail; they substitute. You asked for one model, a different one answered, and the only record is a model field in the response body that approximately nobody reads.
Three consequences, in increasing order of how long they take to notice
1. Quality. Your eval measured one model; production answers with another. Every number you have describes a system you are not running. A "prompt regression" can be a routing change.
2. Billing — and the direction people forget. Everybody guards against being downgraded. Nobody guards against being quietly served the expensive model. The invoice arrives three weeks later with no way to attribute the difference. Upgrades are reported here at the same severity as downgrades.
3. Silent truncation, the worst of the three.
reconcileCall({ requestedModel: "big-200k", servedModel: "small-32k", promptTokens: 100_000, … }, catalog);
// { severity: "critical", truncationRisk: true,
// detail: "asked for big-200k, got small-32k, and the 100000-token prompt exceeds
// small-32k's 32000-token window — the request did not fail, it was CUT,
// and the answer is derived from whatever fit" }Nothing errors. The answer is confident and based on part of your input. This outranks the substitution itself in severity, and it is flagged even when there was no substitution — you can exceed the window of the model you correctly received.
No fuzzy name matching, deliberately
big-200k and big-200k-20260101 are different models unless your catalogue says otherwise. A substitution is a name change, and a matcher generous enough to forgive a suffix is generous enough to hide a swap. If your provider pins dated snapshots, put both ids in the catalogue.
An unknown model is an unchecked call, not a passing one — reported at high severity and excluded from the "served correctly" count.
The number nobody currently has
auditAttribution(calls, catalog);
// { fidelity: 0.5, substitutedInto: { "small-32k": 2 }, costDeltaCents: 1475, truncated: [...] }fidelity — the fraction of requests that actually ran the model they asked for. Worth a dashboard tile, because nothing else on your dashboard will tell you when it drops.
fallbackSafe() checks a fallback chain before you send: if any configured fallback cannot hold the prompt, falling back to it truncates silently, and refusing to send is better.
What it does not do
| not included | why | |---|---| | Calling any provider | it reconciles records you already have | | Maintaining a model catalogue | prices and context windows change weekly; supply your own | | Token counting | true-token-ledger, context-window-packer | | Choosing a model | model-router |
Tiers are your own capability ordering, not a benchmark. If you do not supply them, every substitution reads as lateral — which is still reported, just without a direction.
01Capabilities
Does
- + Model usage tracking
- + Model-routing visibility
- + LLM cost accounting
Doesn’t
- No exclusions declared
02Requirements & stack
Depends on
No declared dependencies
Credentials needed
None declared
Stack
03Community
No endorsements yetNo verified confirmations yet — be the first.
Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.
Sign in to confirm — weight comes from verified usage, not vote count.
Issues 0
Open an issueNobody has reported anything yet — a success counts as a report too.
04Trust Passport
Full passport →0/0 automated components pass. An automated score is never a security guarantee.
- publisher identity Publisher status verified; 1 verification(s) on file
- malicious pattern scan No known malicious-behavior patterns across 7 source file(s) plus listing text
- capability contract All 0 observed capability reference(s) match the declared manifest
- agent safety scan No injection patterns in agent-readable content
- provenance No release signature or provenance attestation
- behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.
Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.
05Versions
Full history →| Version | Channel | Released | Notes |
|---|---|---|---|
| 0.1.0 | stable | Aug 12, 2026 | Initial extraction. |