Skip to content
Code Recycle

Component · for humans & their agents

Eval Design Auditor

verified · first-partyactively maintained$0 during beta (was $59)

Your eval suite reports 94% and the number is meaningless — not wrong, meaningless. If 95 of 100 cases expect the same answer, a grader that always says it scores 95%.

by Code Recycle

Get it free — beta

Every claim on this page is refundable if it is untrue — refund policy.

Verified: 61 tests

Audits an eval suite's DESIGN before you trust what it scores. Zero runtime dependencies. Never runs a model, never grades anything, never performs I/O.

Audits an eval suite's DESIGN before you trust what it scores. Zero runtime dependencies. Never runs a model, never grades anything, never performs I/O.

THE SILENT FAILURE. A team writes an eval suite, it reports 94%, and the number is MEANINGLESS -- not wrong, meaningless. Nothing errors. The suite runs green and the model ships.

CLASS IMBALANCE. If 95 of 100 cases expect the same answer, a grader that always returns that answer scores 95%. The number looks excellent and measures nothing. This is arithmetic, not opinion, and it is invisible in the headline figure -- which is the only figure anyone reads.

The suite also cannot tell you it is broken, because a broken suite and an excellent model produce the SAME output: a high score. That symmetry is the whole problem. Every other check in a pipeline surfaces failure as failure; an eval surfaces its own failure as success.

So this audits the design rather than the score: the distribution of expected answers, ambiguous wording that makes a case ungradeable, cases whose expected answer does not follow from their input, and the gap between what the suite covers and what it claims to.

It reports what would have to be true for the score to mean what it appears to mean -- so a 94% is either earned or explained, rather than simply accepted.

VERIFIED: 61 tests, measured by running the suite.

DELIVERY: signed download of a hash-verified tarball, immediately on purchase. Permissive licence: unlimited products, unlimited clients, unlimited seats, no attribution, perpetual and irrevocable. One restriction, do not republish the source as source.

Interface

What you call, and what comes back. Types and signatures only — the implementation ships with the source.

  export function auditEvalSuite(raw: unknown, options: AuditOptions = {}): AuditReport;
  export function tryCanonicalize(value: unknown): CanonicalizeOk | CanonicalizeFail;
  export function toJSONSafe(value: unknown): JSONValue;
  export function describeValue(value: unknown, maxLen = 80): string;
  export function normalizeText(s: string): string;
  export function detectClassImbalance( cases: ParsedCase[], opts: Pick<ResolvedAuditOptions, "minCasesForImbalanceCheck" | "imbalanceThreshold">, ): ClassImbalanceResult;
  export function detectFixtureContamination(cases: ParsedCase[], examples: ParsedFewShot[]): Finding[];
  export function detectOneSidedCoverage( cases: ParsedCase[], opts: Pick<ResolvedAuditOptions, "minCasesForCoverageCheck" | "coverageTagMinFraction">, ): OneSidedCoverageResult;
  export function resolveOptions(o: AuditOptions): ResolvedAuditOptions;
  export function locationOfCase(c: ParsedCase): string;
  export function parseSuite(raw: unknown): ParseResult;
  export function detectRubricAmbiguity(entries: RubricLine[], terms: string[]): Finding[];
  export type CaseBehavior = "positive" | "refusal" | "abstain" | "error";
  export type Rubric = string | RubricEntry[];
  export type Confidence = "PROVEN" | "SUSPECTED";
  export type Severity = "error" | "warning" | "info";
  export type AuditReport = AuditReportOk | AuditReportParseFailure;

01Capabilities

Does

  • + Experiment and decision gates
  • + Developer tooling
  • + Reliability

Doesn’t

  • No exclusions declared

02Requirements & stack

Depends on

No declared dependencies

Credentials needed

None declared

Stack

typescript

03Community

No endorsements yet

No verified confirmations yet — be the first.

Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.

Open an issue

Sign in to confirm — weight comes from verified usage, not vote count.

0 open · 0 answered · 0 fixed · 1 said it worked

04Trust Passport

Full passport →
–/100

0/0 automated components pass. An automated score is never a security guarantee.

✓ Verified · first-partyreviewed Sep 20, 2026 · re-verification due Dec 19, 2026
  • publisher identity Publisher status verified; 1 verification(s) on file
  • malicious pattern scan No known malicious-behavior patterns across 27 source file(s) plus listing text
  • capability contract All 0 observed capability reference(s) match the declared manifest
  • agent safety scan No injection patterns in agent-readable content
  • provenance No release signature or provenance attestation
  • behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.

Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.

VersionChannelReleasedNotes
1.0.0stableAug 6, 2026First public release.