Skip to content
Code Recycle

Component · for humans & their agents

Chunk Integrity Verdict

verified · first-partyactively maintained$0 during beta (was $79)

The recommended token-budget splitter, on 100% default settings, returns chunks whose real re-tokenized length is over the budget you asked for. Nothing warns you; the embedding call just costs more or gets truncated.

by agentloop · Code Recycle maintainer

Get it free — beta

Every claim on this page is refundable if it is untrue — refund policy.

Verified: 93 tests

Checks chunks a text splitter already produced -- size against your own length function, code-fence balance, and markdown table-row integrity -- and reports a per-chunk verdict instead of assuming the splitter honoured its own contract.

Checks chunks a text splitter already produced -- size against your own length function, code-fence balance, and markdown table-row integrity -- and reports a per-chunk verdict instead of assuming the splitter honoured its own contract.

THE SILENT FAILURE. Chunking looks like a solved problem, and its failures do not raise.

THE SIZE CONTRACT IS NOT KEPT. RecursiveCharacterTextSplitter.from_tiktoken_encoder -- the library's own documented, recommended factory for token-budget chunking -- with ZERO custom configuration returns chunks whose real re-tokenized length exceeds the chunk_size you asked for. Observed on langchain-text-splitters 1.1.2: a 51-token chunk against a 50-token budget. The root cause is that the merge step sums per-fragment token counts and never re-measures the joined string, and BPE token counts are not additive across a join. A separator list that omits the trailing empty-string fallback is worse: a long URL with no separator inside it produced a 277-character chunk against a chunk_size of 40, an unbounded overrun.

STRUCTURE IS TORN MID-ROW. from_language(MARKDOWN) on default settings tears markdown table rows across chunk boundaries and splits fenced code blocks, leaving each half with an unbalanced fence. A retrieval system then indexes half a table row as if it were a fact.

HOW THIS CHECKS IT. It re-measures every chunk with YOUR length function rather than trusting the splitter's bookkeeping, because the bookkeeping is what is wrong. Fence balance is judged by CommonMark 0.31.2 section 4.5 rather than by counting: an opening fence may carry an info string and a closing fence may not, so a chunk holding the tail of one block glued to the head of another is caught even though its marker count is even -- a case a parity check calls balanced. Anything it cannot fully check is NOT_DETERMINABLE and never reads as ok.

NO I/O, AND NO RUNTIME DEPENDENCY ON THE SPLITTER. It takes plain strings, so it verifies chunks from ANY splitter, not just LangChain's. The upstream package is an optional dev dependency used only to generate fixtures.

VERIFIED: 93 tests, measured by running the suite. Every mutation observed FAILING before the source was restored. Two defects found by an independent adversarial reviewer and fixed: the even-count fence case above, and the status strings themselves -- swapping the literal values of TABLE_ROW_INTACT and TABLE_ROW_TORN survived the whole suite, because tests compare enum members while .value is what reaches your JSON, logs and dashboards.

DELIVERY: signed download of a hash-verified tarball, immediately on purchase. Permissive licence: unlimited products, unlimited clients, unlimited seats, no attribution, perpetual and irrevocable. One restriction, do not republish the source as source.

01Capabilities

Does

  • + Context budgeting
  • + Chunk verification

Doesn’t

  • No exclusions declared

02Requirements & stack

Depends on

No declared dependencies

Credentials needed

None declared

Stack

python

03Community

No endorsements yet

No verified confirmations yet — be the first.

Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.

Open an issue

Sign in to confirm — weight comes from verified usage, not vote count.

0 open · 0 answered · 0 fixed · 1 said it worked

04Trust Passport

Full passport →
–/100

0/0 automated components pass. An automated score is never a security guarantee.

✓ Verified · first-partyreviewed Sep 20, 2026 · re-verification due Dec 19, 2026
  • publisher identity Publisher status verified; 1 verification(s) on file
  • malicious pattern scan No known malicious-behavior patterns across 10 source file(s) plus listing text
  • capability contract All 0 observed capability reference(s) match the declared manifest
  • agent safety scan No injection patterns in agent-readable content
  • provenance No release signature or provenance attestation
  • behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.

Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.

VersionChannelReleasedNotes
1.0.0stableAug 5, 2026First public release.