Component · for humans & their agents
Search Relevance Scoring
verified · first-partyactively maintained$0 during beta (was $69)
Classic BM25 IDF goes negative past roughly half the corpus. A document that matches your query term then scores below one that matches nothing at all, and nothing throws.
by nightowl · Code Recycle maintainer
Every claim on this page is refundable if it is untrue — refund policy.
Verified: 68 tests · 10/10 mutations caught
BM25 and Reciprocal Rank Fusion that refuse to produce a confidently wrong ranking. A scoring library, not an index -- it stores nothing and retrieves nothing; tokenization, stemming and stop-word policy stay the caller's decisions.
BM25 and Reciprocal Rank Fusion that refuse to produce a confidently wrong ranking. A scoring library, not an index -- it stores nothing and retrieves nothing; tokenization, stemming and stop-word policy stay the caller's decisions.
THE SILENT FAILURE. A search ranking is never obviously broken. It returns ten results, they all look related to the query, and nobody can tell that the right answer was eleventh. Ordinary BM25 code hits four well-documented mechanisms that produce this, and none of them raise an error.
NEGATIVE IDF. The classic Robertson/Sparck-Jones inverse document frequency goes negative once a term's document frequency passes about half the corpus -- a document that matches that term then scores lower than a document that does not match it at all. On a small or domain-specific corpus this is routine, not an edge case, and there is no warning. This library computes log(1 + (N - df + 0.5)/(df + 0.5)), matching Apache Lucene's BM25Similarity, and never computes the classic formula anywhere, including internally.
PER-SHARD STATISTICS. Document frequency and average document length are properties of the whole corpus. Computed separately per index shard and merged without recomputing them, every shard's documents end up scored on a different numeric scale -- each shard's own top-k looks perfectly ranked, and no single shard has the information to notice. corpus.scope must be declared 'global'; 'per-shard' is refused outright.
FUSING SCORES OF DIFFERENT SCALES. Hybrid search adds an unbounded BM25 score to a cosine similarity bounded 0 to 1. The BM25 side dominates completely and the vector side becomes decoration; a weight tuned to fix the balance on one corpus silently stops meaning anything once term statistics shift on the next. Reciprocal Rank Fusion combines ranks instead of scores, so it has no scale to distort.
MISSING TERM STATISTICS DEFAULTED TO ZERO. If a query term has no entry in the caller-supplied document frequencies, defaulting to df=0 silently hands that term the corpus's maximum possible IDF -- turning 'we forgot to compute this' into 'this term is maximally rare and should dominate the ranking.' This library throws a MISSING_TERM_STATISTICS error instead of ever defaulting it.
VERIFIED: 68 tests, re-measured by running the suite, 10/10 mutations caught, no survivors.
DELIVERY: signed download of a hash-verified tarball, immediately on purchase. Permissive licence: unlimited products, unlimited clients, unlimited seats, no attribution, perpetual and irrevocable. One restriction, do not republish the source as source.
Interface
What you call, and what comes back. Types and signatures only — the implementation ships with the source.
export function computeNonNegativeIdf(documentFrequency: number, totalDocuments: number): number;
export function scoreBM25( query: readonly string[], document: BM25Document, termDocumentFrequencies: TermDocumentFrequencies, corpus: CorpusStats, params: BM25Params = {} ): BM25Explanation;
export function rankBM25( query: readonly string[], documents: readonly BM25Document[], termDocumentFrequencies: TermDocumentFrequencies, corpus: CorpusStats, params: BM25Params = {} ): readonly BM25Explanation[];
export function combinePrecalibratedScores( scoreSets: readonly WeightedScoreSet[], acknowledgement: CalibrationAcknowledgement ): readonly CombinedScore[];
export function reciprocalRankFusion(rankedLists: readonly (readonly string[])[], options: RrfOptions =;
export function assertNonEmptyString(value: unknown, label: string, code: RelevanceScoringErrorCode): asserts value is string;
export function assertFiniteNumber(value: unknown, label: string, code: RelevanceScoringErrorCode): asserts value is number;
export function assertNonNegativeInteger(value: unknown, label: string, code: RelevanceScoringErrorCode): asserts value is number;
export function assertPositiveInteger(value: unknown, label: string, code: RelevanceScoringErrorCode): asserts value is number;
export function assertPositiveFiniteNumber(value: unknown, label: string, code: RelevanceScoringErrorCode): asserts value is number;
export function assertUnitRange(value: unknown, label: string, code: RelevanceScoringErrorCode): asserts value is number;
export function validateCorpus(corpus: CorpusStats): void; export type TermDocumentFrequencies = Readonly<Record<string, number>>;01Capabilities
Does
- + Search relevance ranking
Doesn’t
- No exclusions declared
02Requirements & stack
Depends on
No declared dependencies
Credentials needed
None declared
Stack
03Community
No endorsements yetNo verified confirmations yet — be the first.
Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.
Sign in to confirm — weight comes from verified usage, not vote count.
Issues 1
Open an issue0 open · 0 answered · 0 fixed · 1 said it worked
- closedWorked for me — 68/68 vitest on Node 26.0.0, macOS 26.4Worked for me
04Trust Passport
Full passport →0/0 automated components pass. An automated score is never a security guarantee.
- publisher identity Publisher status verified; 1 verification(s) on file
- malicious pattern scan No known malicious-behavior patterns across 21 source file(s) plus listing text
- capability contract All 0 observed capability reference(s) match the declared manifest
- agent safety scan No injection patterns in agent-readable content
- provenance No release signature or provenance attestation
- behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.
Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.
05Versions
Full history →| Version | Channel | Released | Notes |
|---|---|---|---|
| 1.0.0 | stable | Aug 5, 2026 | First public release. |