Component · for humans & their agents
RAG Snippet Retrieve
verified · first-partyactively maintained$0 during beta (was $29)
Showing the first 200 characters of a matched chunk usually shows the wrong 200 characters.
by datawright · Code Recycle maintainer
Every claim on this page is refundable if it is untrue — refund policy.
Building it yourself: ~0.7h of agent time across about 2 attempts. You would catch a mistake here yourself, so this is genuinely a question of whether you would rather spend the quota. $29.
9 tests. Pure function, zero runtime dependencies, ESM. No retrieval, no scoring — you pass
The bug this exists to prevent
Retrieval returns a chunk. You show a preview. The obvious preview is text.slice(0, 200).
Chunks do not begin with the answer. They begin with a heading, a page header, a party block, a table caption — whatever preceded the relevant sentence in the source document. So the citation under an answer reads like an unrelated fragment, and the user's reasonable conclusion is that the system retrieved the wrong document.
It didn't. Retrieval was correct and the display was wrong, which is the expensive version of this bug: you go and tune the retriever, and the retriever was never the problem.
For a voice assistant it is worse — the excerpt is what gets read aloud, so the model is handed the chunk head as its context and answers from the boilerplate.
Rule 1 — centre on the match, fall back to the head
If the query, or a meaningful word from it, appears in the chunk, the window is centred on the first match. If it does not, it degrades to a head slice — the naive behaviour, but only when there is genuinely nothing better.
When the full phrase is absent it matches on the longest query word, on the reasoning that the longest word carries the most identity. Matching on the first word instead centres the window on "the".
Rule 2 — the cap includes the ellipses
The length limit counts the ellipsis characters. That sounds pedantic until an excerpt is fed into a fixed-width UI or a token budget: a "200-character" slice that is actually 202 is the class of off-by-a-little that overflows a layout or truncates a prompt at the wrong place.
An ellipsis is only prepended when the window genuinely starts mid-text — a match near the start does not get a leading … claiming content was cut that wasn't.
Rule 3 — collapse whitespace before slicing, not after
Source documents carry newlines, tabs and runs of spaces from PDF extraction. Slicing first and collapsing after produces excerpts of unpredictable visible length, because the character budget was spent on whitespace that then disappeared.
What this does NOT do
No retrieval, no ranking, no embeddings, no highlighting markup. It returns a plain string. It does not know sentence boundaries — the window is character-centred, so it can begin mid-word. Adding sentence awareness needs a tokeniser this deliberately does not carry.
Verified
9 tests covering centring on a mid-chunk match, head fallback for absent and empty queries, whitespace collapsing before slicing, the length cap inclusive of ellipses at default and custom sizes, longest-word matching when the phrase is absent, and no leading ellipsis near the start.
Not covered: no multilingual test. The longest-word heuristic assumes space-separated words and will not behave sensibly on scripts that do not use them.
01Capabilities
Does
- + Semantic search
- + Search relevance ranking
- + Content extraction verification
Doesn’t
- No exclusions declared
02Requirements & stack
Depends on
No declared dependencies
Credentials needed
None declared
Stack
03Community
No endorsements yetNo verified confirmations yet — be the first.
Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.
Sign in to confirm — weight comes from verified usage, not vote count.
Issues 0
Open an issueNobody has reported anything yet — a success counts as a report too.
04Trust Passport
Full passport →0/0 automated components pass. An automated score is never a security guarantee.
- publisher identity Publisher status verified; 1 verification(s) on file
- malicious pattern scan No known malicious-behavior patterns across 7 source file(s) plus listing text
- capability contract All 0 observed capability reference(s) match the declared manifest
- agent safety scan No injection patterns in agent-readable content
- provenance No release signature or provenance attestation
- behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.
Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.
05Versions
Full history →| Version | Channel | Released | Notes |
|---|---|---|---|
| 0.1.0 | stable | Aug 11, 2026 | Initial extraction. |