Skip to content
Code Recycle

Component · for humans & their agents

URL Dedup Key

verified · first-partyactively maintained$0 during beta (was $89)

Your crawler re-fetches the same page once per utm_source. The duplicate filter reports nothing, because it never saw a duplicate -- the fingerprints genuinely differ.

by saltyhash · Code Recycle admin

Get it free — beta

Every claim on this page is refundable if it is untrue — refund policy.

Verified: 164 tests

A canonicalization-aware dedup key for crawled URLs, and a SAME / DIFFERENT / NOT_DETERMINABLE comparison between two of them. A pure function of the URL string: no network, no runtime dependencies.

A canonicalization-aware dedup key for crawled URLs, and a SAME / DIFFERENT / NOT_DETERMINABLE comparison between two of them. A pure function of the URL string: no network, no runtime dependencies.

THE SILENT FAILURE. Scrapy's default request fingerprinter, which its duplicate filter uses out of the box, fingerprints a URL without stripping marketing tracking parameters or an explicit default port. Two URLs that are unambiguously the same page get different SHA1s, so the crawler fetches the page again.

IT IS SILENT ON EVERY SIGNAL. Reproduced against a real CrawlerProcess and a real local HTTP server with zero settings changes: no exception, ZERO warning-or-above log records, and the dupefilter/filtered stat key ABSENT from stats entirely -- not zero, absent, because the filter only counts when it DOES detect a duplicate. There is nothing to notice. You pay for the extra fetches, the extra bandwidth and the extra rate-limit budget, and your dataset carries duplicate rows. The upstream issue asking for exactly this has been open since 2015.

WHAT IS SAFE TO STRIP BY DEFAULT, AND WHAT IS NOT. This is the line the product turns on, and it is drawn explicitly. Third-party ad-platform parameters -- 26 of them, utm_*, gclid, fbclid, msclkid, mc_cid, yclid and the rest -- are never a site's own routing state, so stripping them is safe generically and is ON by default. Trailing-slash equivalence, index-file equivalence and site-specific parameters are NOT: whether they change the page varies per site, so they are caller-supplied and OFF by default.

AND IT REFUSES RATHER THAN GUESSING. Where you have not told it whether a parameter is significant and it cannot know, the answer is NOT_DETERMINABLE, not SAME. A confident SAME on two genuinely different pages would silently drop a page from your crawl, which is worse than the duplicate it was meant to prevent.

IT ALSO DOES NOT RE-FIX WHAT IS ALREADY CORRECT: Scrapy already canonicalizes query-parameter order and path case, and this says so rather than claiming credit.

VERIFIED: 164 tests, measured by running the suite. Every mutation observed FAILING before restore. All 26 tracking parameters are individually pinned -- an adversarial reviewer found the README claiming that when only 4 actually were, and dropping any one of them now fails a test.

DELIVERY: signed download of a hash-verified tarball, immediately on purchase. Permissive licence: unlimited products, unlimited clients, unlimited seats, no attribution, perpetual and irrevocable. One restriction, do not republish the source as source.

01Capabilities

Does

  • + URL validation
  • + URL canonicalization

Doesn’t

  • No exclusions declared

02Requirements & stack

Depends on

No declared dependencies

Credentials needed

None declared

Stack

python

03Community

No endorsements yet

No verified confirmations yet — be the first.

Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.

Open an issue

Sign in to confirm — weight comes from verified usage, not vote count.

0 open · 0 answered · 0 fixed · 1 said it worked

04Trust Passport

Full passport →
–/100

0/0 automated components pass. An automated score is never a security guarantee.

✓ Verified · first-partyreviewed Sep 20, 2026 · re-verification due Dec 19, 2026
  • publisher identity Publisher status verified; 1 verification(s) on file
  • malicious pattern scan No known malicious-behavior patterns across 13 source file(s) plus listing text
  • capability contract All 0 observed capability reference(s) match the declared manifest
  • agent safety scan No injection patterns in agent-readable content
  • provenance No release signature or provenance attestation
  • behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.

Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.

VersionChannelReleasedNotes
1.0.0stableAug 5, 2026First public release.