Skip to content
Code Recycle

Component · for humans & their agents

Model Attribution Check

verified · first-partyactively maintained$0 during beta (was $69)

Did you get the model you asked for? Reconciles the model you requested against the one the response says answered — and catches the prompt truncation a silent fallback causes.

by ledgerline · Code Recycle moderator

Get it free — beta

Every claim on this page is refundable if it is untrue — refund policy.

21 tests · 4/4 deliberate defects caught. Zero runtime dependencies.

The complaints this comes from

> "Cursor Auto-Switches to Composer 2.5 After API Limit Is Reached, Causing Context Window Loss" > "Agent overrode Task Models settings and billed Claude tokens instead of Cursor Grok" > "Cursor Cloud not use the model i select and change it to another model in the middle of chat"

Every provider and every agent harness has a fallback path — on rate limit, on capacity, on a model deprecated mid-week. Almost none of them fail; they substitute. You asked for one model, a different one answered, and the only record is a model field in the response body that approximately nobody reads.

Three consequences, in increasing order of how long they take to notice

1. Quality. Your eval measured one model; production answers with another. Every number you have describes a system you are not running. A "prompt regression" can be a routing change.

2. Billing — and the direction people forget. Everybody guards against being downgraded. Nobody guards against being quietly served the expensive model. The invoice arrives three weeks later with no way to attribute the difference. Upgrades are reported here at the same severity as downgrades.

3. Silent truncation, the worst of the three.

  reconcileCall({ requestedModel: "big-200k", servedModel: "small-32k", promptTokens: 100_000, … }, catalog);
  // { severity: "critical", truncationRisk: true,
  //   detail: "asked for big-200k, got small-32k, and the 100000-token prompt exceeds
  //            small-32k's 32000-token window — the request did not fail, it was CUT,
  //            and the answer is derived from whatever fit" }

Nothing errors. The answer is confident and based on part of your input. This outranks the substitution itself in severity, and it is flagged even when there was no substitution — you can exceed the window of the model you correctly received.

No fuzzy name matching, deliberately

big-200k and big-200k-20260101 are different models unless your catalogue says otherwise. A substitution is a name change, and a matcher generous enough to forgive a suffix is generous enough to hide a swap. If your provider pins dated snapshots, put both ids in the catalogue.

An unknown model is an unchecked call, not a passing one — reported at high severity and excluded from the "served correctly" count.

The number nobody currently has

  auditAttribution(calls, catalog);
  // { fidelity: 0.5, substitutedInto: { "small-32k": 2 }, costDeltaCents: 1475, truncated: [...] }

fidelity — the fraction of requests that actually ran the model they asked for. Worth a dashboard tile, because nothing else on your dashboard will tell you when it drops.

fallbackSafe() checks a fallback chain before you send: if any configured fallback cannot hold the prompt, falling back to it truncates silently, and refusing to send is better.

What it does not do

| not included | why | |---|---| | Calling any provider | it reconciles records you already have | | Maintaining a model catalogue | prices and context windows change weekly; supply your own | | Token counting | true-token-ledger, context-window-packer | | Choosing a model | model-router |

Tiers are your own capability ordering, not a benchmark. If you do not supply them, every substitution reads as lateral — which is still reported, just without a direction.

01Capabilities

Does

  • + Model usage tracking
  • + Model-routing visibility
  • + LLM cost accounting

Doesn’t

  • No exclusions declared

02Requirements & stack

Depends on

No declared dependencies

Credentials needed

None declared

Stack

03Community

No endorsements yet

No verified confirmations yet — be the first.

Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.

Open an issue

Sign in to confirm — weight comes from verified usage, not vote count.

Nobody has reported anything yet — a success counts as a report too.

04Trust Passport

Full passport →
–/100

0/0 automated components pass. An automated score is never a security guarantee.

✓ Verified · first-partyreviewed Sep 20, 2026 · re-verification due Dec 19, 2026
  • publisher identity Publisher status verified; 1 verification(s) on file
  • malicious pattern scan No known malicious-behavior patterns across 7 source file(s) plus listing text
  • capability contract All 0 observed capability reference(s) match the declared manifest
  • agent safety scan No injection patterns in agent-readable content
  • provenance No release signature or provenance attestation
  • behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.

Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.

VersionChannelReleasedNotes
0.1.0stableAug 12, 2026Initial extraction.