Skip to content
Code Recycle

Component · for humans & their agents

Clause Span Grounding

verified · first-partyactively maintained$0 during beta (was $79)

A fluent paraphrase is not a low-confidence extraction — it is a wrong one. Verifies that every extracted clause exists in the source document, character-for-character.

by datawright · Code Recycle maintainer

Get it free — beta

Every claim on this page is refundable if it is untrue — refund policy.

21 tests · 4/4 deliberate defects caught. Zero runtime dependencies.

Not legal advice, and not a model

This answers one mechanical question: does the text this extraction claims to have found actually appear in the document, at the offsets it cites? It has no opinion about whether a clause is favourable, enforceable, standard, or worth signing. **It cannot tell you whether to sign anything.** Those are legal judgements, and nothing in this package is qualified to make them.

The defect

Ask a language model to "extract the termination clause" and you get something that reads exactly like a termination clause. Sometimes it is the one in the document. Sometimes it is a fluent paraphrase. Occasionally it is a clause from a different agreement entirely, reproduced with total confidence.

These two sentences differ by one word:

  The Tenant may not terminate this Lease prior to the Expiration Date.
  The Tenant may     terminate this Lease prior to the Expiration Date.

Both read as competent legal English. One is the opposite of the other. A reviewer reading a summary cannot tell which the document said — and the extraction pipeline reported success either way.

  verifyExtraction(doc, { label: "termination", quote: "The Tenant may terminate…" });
  // { status: "not_found", detail: "…one dropped negation reverses the meaning while
  //   still reading correctly" }

Normalisation is where a verifier stops verifying

Real documents arrive from PDFs with soft hyphens, ligatures, curly quotes, non-breaking spaces and line breaks mid-sentence. Compare raw strings and you reject almost everything genuine. So normalisation is necessary — and it is exactly the knob that, turned one notch too far, starts accepting may for may not.

Everything normalised here is typographic: quote and dash shapes, ligatures, whitespace, invisible characters. Nothing that carries meaning. No stemming, no stop-word stripping, no negation collapsing. A test asserts may not never normalises to may.

Citations that resolve

| status | meaning | |---|---| | grounded | found at the cited offsets, or found unambiguously | | offset_mismatch | the text is there, but not where the extraction said | | not_found | it is not in the document | | ambiguous | appears more than once, no offsets given |

offset_mismatch is rejected, not repaired. The text being present somewhere is not evidence the extractor found it there; silently correcting the offset hides that its citations cannot be trusted. Offsets are reported against the original string, not a normalised copy, so they can actually highlight the source.

Redlines where one word is the whole change

  materialDifference("The Tenant may not terminate.", "The Tenant may terminate.");
  // { differs: true, reversingTokens: ["not"] }

"3 words changed" is useless when one of them is not. This reports whether a change touches a meaning-reversing token — without deciding what the change means.

What it does not do

| not included | why | |---|---| | Extracting clauses | bring your own extractor; this checks its work | | Reading PDFs | pdf-manual-parse, pdf-reading-order | | Judging clause quality, risk or enforceability | legal judgement, out of scope, permanently | | Semantic similarity matching | that is the defect, not the feature |

The matcher is exact-after-typographic-normalisation. It will reject a genuine extraction whose quote was cleaned up by the extractor — that is the intended trade, because the alternative accepts the reversed clause.

01Capabilities

Does

  • + Contract analysis
  • + Clause extraction
  • + Content extraction verification

Doesn’t

  • No exclusions declared

02Requirements & stack

Depends on

No declared dependencies

Credentials needed

None declared

Stack

03Community

No endorsements yet

No verified confirmations yet — be the first.

Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.

Open an issue

Sign in to confirm — weight comes from verified usage, not vote count.

Nobody has reported anything yet — a success counts as a report too.

04Trust Passport

Full passport →
–/100

0/0 automated components pass. An automated score is never a security guarantee.

✓ Verified · first-partyreviewed Sep 20, 2026 · re-verification due Dec 19, 2026
  • publisher identity Publisher status verified; 1 verification(s) on file
  • malicious pattern scan No known malicious-behavior patterns across 7 source file(s) plus listing text
  • capability contract All 0 observed capability reference(s) match the declared manifest
  • agent safety scan No injection patterns in agent-readable content
  • provenance No release signature or provenance attestation
  • behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.

Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.

VersionChannelReleasedNotes
0.1.0stableAug 12, 2026Initial extraction.