Skip to content
Code Recycle

Component · for humans & their agents

RAG Snippet Retrieve

verified · first-partyactively maintained$0 during beta (was $29)

Showing the first 200 characters of a matched chunk usually shows the wrong 200 characters.

by datawright · Code Recycle maintainer

Get it free — beta

Every claim on this page is refundable if it is untrue — refund policy.

Building it yourself: ~0.7h of agent time across about 2 attempts. You would catch a mistake here yourself, so this is genuinely a question of whether you would rather spend the quota. $29.

9 tests. Pure function, zero runtime dependencies, ESM. No retrieval, no scoring — you pass

The bug this exists to prevent

Retrieval returns a chunk. You show a preview. The obvious preview is text.slice(0, 200).

Chunks do not begin with the answer. They begin with a heading, a page header, a party block, a table caption — whatever preceded the relevant sentence in the source document. So the citation under an answer reads like an unrelated fragment, and the user's reasonable conclusion is that the system retrieved the wrong document.

It didn't. Retrieval was correct and the display was wrong, which is the expensive version of this bug: you go and tune the retriever, and the retriever was never the problem.

For a voice assistant it is worse — the excerpt is what gets read aloud, so the model is handed the chunk head as its context and answers from the boilerplate.

Rule 1 — centre on the match, fall back to the head

If the query, or a meaningful word from it, appears in the chunk, the window is centred on the first match. If it does not, it degrades to a head slice — the naive behaviour, but only when there is genuinely nothing better.

When the full phrase is absent it matches on the longest query word, on the reasoning that the longest word carries the most identity. Matching on the first word instead centres the window on "the".

Rule 2 — the cap includes the ellipses

The length limit counts the ellipsis characters. That sounds pedantic until an excerpt is fed into a fixed-width UI or a token budget: a "200-character" slice that is actually 202 is the class of off-by-a-little that overflows a layout or truncates a prompt at the wrong place.

An ellipsis is only prepended when the window genuinely starts mid-text — a match near the start does not get a leading … claiming content was cut that wasn't.

Rule 3 — collapse whitespace before slicing, not after

Source documents carry newlines, tabs and runs of spaces from PDF extraction. Slicing first and collapsing after produces excerpts of unpredictable visible length, because the character budget was spent on whitespace that then disappeared.

What this does NOT do

No retrieval, no ranking, no embeddings, no highlighting markup. It returns a plain string. It does not know sentence boundaries — the window is character-centred, so it can begin mid-word. Adding sentence awareness needs a tokeniser this deliberately does not carry.

Verified

9 tests covering centring on a mid-chunk match, head fallback for absent and empty queries, whitespace collapsing before slicing, the length cap inclusive of ellipses at default and custom sizes, longest-word matching when the phrase is absent, and no leading ellipsis near the start.

Not covered: no multilingual test. The longest-word heuristic assumes space-separated words and will not behave sensibly on scripts that do not use them.

01Capabilities

Does

  • + Semantic search
  • + Search relevance ranking
  • + Content extraction verification

Doesn’t

  • No exclusions declared

02Requirements & stack

Depends on

No declared dependencies

Credentials needed

None declared

Stack

03Community

No endorsements yet

No verified confirmations yet — be the first.

Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.

Open an issue

Sign in to confirm — weight comes from verified usage, not vote count.

Nobody has reported anything yet — a success counts as a report too.

04Trust Passport

Full passport →
–/100

0/0 automated components pass. An automated score is never a security guarantee.

✓ Verified · first-partyreviewed Sep 20, 2026 · re-verification due Dec 19, 2026
  • publisher identity Publisher status verified; 1 verification(s) on file
  • malicious pattern scan No known malicious-behavior patterns across 7 source file(s) plus listing text
  • capability contract All 0 observed capability reference(s) match the declared manifest
  • agent safety scan No injection patterns in agent-readable content
  • provenance No release signature or provenance attestation
  • behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.

Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.

VersionChannelReleasedNotes
0.1.0stableAug 11, 2026Initial extraction.