Skip to content
Code Recycle

Silent failures

The bugs that don’t tell you they happened.

Every one of these ships green. No error, no failed deploy, no alert — the dashboard says everything worked. You find out weeks later, or a stranger does.

Nobody searches for a bug they don’t know exists. So your coding agent can ask us instead: it describes your project, we name what it’s exposed to. The free fix is always first — if a one-line change solves it, that’s what we tell you.

A user sees another tenant's data, occasionally, under load.

critical
What’s actually happening
Setting a tenant-scoping Postgres session variable at the connection level breaks under transaction-mode pooling (PgBouncer, Neon): the value either leaks to the next request that reuses the connection, or is gone before your query runs.
Why you never see it
It is load-dependent and non-deterministic. Local development uses a direct connection and never reproduces it.

Free fix

Set the scoping variable inside the SAME interactive transaction as the query, never on the connection. Also verify raw query escape hatches fail closed.

Documented and tested in a production multi-tenant app running on Neon pooled connections.

An order gets fulfilled twice, or a paid order never gets fulfilled at all.

critical
What’s actually happening
The payment provider's webhook and the browser success-redirect race each other; both try to fulfill. Separately, a failure in an optional side effect (a payout, an email) can abort the handler after payment but before entitlement.
Why you never see it
The happy path works in testing. The race only appears under real latency, and the stranded-order case looks like a provider problem.

Free fix

Make the claim atomic and exclusive (compare-and-set from unclaimed states only), dedupe by event id, return 200 on already-processed, and isolate every non-buyer-facing side effect so it can never abort fulfillment.

Hit in this marketplace's own Stripe integration, twice: a payout failure stranded a paid order, and a non-exclusive claim let two callers fulfill.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

Anyone with your public key can read your whole table.

critical
What’s actually happening
Row Level Security is off (or a policy is missing) on a table exposed through the client API, so the anon key can select everything.
Why you never see it
The app works perfectly. Nothing in the UI reveals that the data is world-readable.

Free fix

Enable RLS on every exposed table and add per-operation policies. Supabase's own dashboard Security Advisor flags missing policies for free.

CVE-2025-48757 — 170+ AI-built applications exposed via missing or disabled RLS; 303 vulnerable endpoints found by one researcher.

A background task disappears with no error when a worker is killed — or, with the setting many guides recommend as 'safer,' that same task instead runs twice from the beginning, including everything the first run had already done.

critical
What’s actually happening
Celery's acks_late controls WHEN the broker acknowledges a task, not how many times it can run. Default (acks_late=False) acknowledges on receipt, so a SIGKILL'd worker (OOM kill, spot reclaim, pod eviction, `kill -9` during a deploy) takes the task with it permanently. acks_late=True acknowledges on completion instead, so a killed worker's unacknowledged task is redelivered to another worker and re-run from scratch.
Why you never see it
The option's name reads like a timing/performance detail, not a binary choice between two failure modes — there is no third setting. A lost task leaves an empty queue, indistinguishable from a queue that finished normally; a duplicated task leaves two sets of side effects and one result row, which looks like one run that happened to be slow. Neither increments a failure count, a retry count, or fires an alert.

Free fix

Set task_acks_late=True together with task_reject_on_worker_lost=True, AND make the task idempotent (idempotency key, upsert not insert, downstream dedupe) — acks_late alone only trades silent loss for silent duplication; idempotency is what actually closes the gap. Record task starts somewhere the broker can't lose them so a vanished task is at least detectable.

Measured against real Celery 5.6.3 + Redis 8.x on 127.0.0.1:6399: worker SIGKILLed mid-task under both settings.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

An attacker who wrote their own JWT gets treated as an admin, and nothing in your logs, error handling, or monitoring ever fires.

critical
What’s actually happening
jwt.decode() is a base64 parser, not a signature check — it returns whatever claims object is embedded in the token's payload for ANY syntactically valid JWT, including one an attacker signed with a secret of their own choosing, one with alg:'none' and no signature at all, or a genuine token whose payload was edited after issuance. Only jwt.verify() checks the signature, and it throws on every one of those.
Why you never see it
decode() returns a well-formed object in exactly the shape the calling code expects (role, sub, etc.) with no exception, so a code path that reads decode() output for an authorization decision looks correct in every test that only ever hands it a real, unmodified token.

Free fix

Replace jwt.decode() with jwt.verify(token, key, { algorithms: [...] }) at every call site whose result influences an authorization decision — it's a same-shape function swap, free. Measured on jsonwebtoken 9.0.3, the library already infers the correct algorithm family from the key type, so classic algorithm-confusion forgeries are rejected even before you add an explicit algorithms allowlist (pin it anyway as defense-in-depth).

Own measurement against jsonwebtoken 9.0.3: four attacker-producible tokens (self-signed forgery, alg:none unsigned, tampered payload, expired-but-genuine) fed to both jwt.decode() and jwt.verify(), plus an RS256/HS256 algorithm-confusion probe.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A server-only secret (an internal API token, a signing key) ends up sitting in your page's HTML source, even though you never wrote NEXT_PUBLIC_ anywhere and grepping every JS file your browser downloads finds nothing.

critical
What’s actually happening
A server component builds an object that happens to include a server-only field (e.g. `{ id, name, apiToken: process.env.INTERNAL_API_TOKEN }`) and passes the whole object as a prop to a client component. Next.js serializes the ENTIRE prop into the RSC payload embedded in the page HTML — it does not matter that the client component only ever reads `user.name`. Whether a field is used has no bearing on whether it crosses the server/client boundary.
Why you never see it
There is no NEXT_PUBLIC_ prefix anywhere to grep for, the client component genuinely never reads the leaked field so nothing 'looks' wired to it, and the token lives inside .next/server/app/*.html and *.rsc — not in .next/static/ — so the standard mitigation of scanning the client JS bundle for secrets finds nothing at all.

Free fix

Pass only the fields the client actually needs — construct a narrowed object (`{ name: user.name }`) at the server/client boundary instead of forwarding the object you happen to be holding. Verify by grepping `.next/server/app/*.html` and `*.rsc` for the secret value after a production build.

Reproduced this session on Next 15.5.22, production `next build` + `next start`.

Your security review confirms Row Level Security is ON and the policy text is correct — and the application can still read every tenant's rows anyway.

critical
What’s actually happening
`ALTER TABLE ... ENABLE ROW LEVEL SECURITY` does not apply to the table's owner — and the owner is exactly the role that ran your migrations, which is the normal shape of a single-DATABASE_URL app. Only the separate `ALTER TABLE ... FORCE ROW LEVEL SECURITY` applies the policy to the owning connection too. Independently: a second policy added later (a debug view, an admin dashboard) is PERMISSIVE by default, which combines with OR — it silently widens access without changing the original, reviewed policy at all.
Why you never see it
`relrowsecurity=true` and the policy's `qual` expression both read correctly under direct inspection, so the standard 'is RLS enabled, and does the policy look right?' review passes cleanly. The widening second policy leaves the first, already-reviewed policy byte-for-byte unchanged, so diffing 'the policy' shows nothing — the access change is happening at the OR-combination level, not in any single artifact someone would think to re-review.

Free fix

Run `ALTER TABLE <table> FORCE ROW LEVEL SECURITY;` for every table your app's own connection role owns. When adding a second policy for an admin/debug/support path, declare it `AS RESTRICTIVE` (which ANDs) rather than accepting the default PERMISSIVE (which ORs), unless you specifically intend that policy to widen access.

Reproduced this session against a real PostgreSQL server: two roles (table-owning app_owner vs. non-owner app_user), one table, one standard-shaped RLS policy.

Two things happened around the same time — a deposit and a purchase, two people editing the same record — and only one of them actually stuck. No error, no conflict message, just a total that's wrong by exactly the amount of the change that got silently dropped.

critical
What’s actually happening
At PostgreSQL's default READ COMMITTED isolation level, two concurrent transactions that each SELECT a value, compute a new value in application code, then UPDATE, will silently lose one of the writes: the second UPDATE simply overwrites the first with no awareness that the row changed underneath it, and both transactions commit successfully.
Why you never see it
READ COMMITTED genuinely does guarantee you never read uncommitted data, so it 'sounds' safe on paper, and both writers get a normal success response. The failure is purely timing-dependent — it will not reproduce in serial manual testing or in a single-user QA pass, only under real concurrent load.

Free fix

Do the arithmetic inside the UPDATE statement itself (`SET balance = balance + 50`, no prior SELECT) — verified correct, because the row is locked for the single statement's duration. If the change genuinely can't be expressed as one SQL statement, wrap the read in `SELECT ... FOR UPDATE` inside the same transaction — also verified correct. Raising the isolation level to SERIALIZABLE alone is not the fix: it converts the silent wrong answer into a retryable error, but only an actual application-level retry recovers the lost write.

Reproduced this session against a real PostgreSQL server: one row starting at balance 100, two concurrent read-then-write transactions.

Two completely unrelated customers on a multi-tenant hosting platform — victim.github.io and attacker.github.io, or two different vercel.app project subdomains — get treated as 'the same site' by a cookie-scoping, same-site, or allowlist check, while a raw IP address like 192.168.1.1 gets treated as if it WERE a valid domain.

critical
What’s actually happening
The two most popular Public Suffix List libraries fail in opposite, complementary ways on real inputs. `tldts` collapses multi-tenant hosting subdomains — *.github.io, *.vercel.app — down to the shared suffix (github.io, vercel.app) instead of respecting the PSL's private-domain section, erasing exactly the tenant boundary that section exists to draw. `psl`, on IPv4 literals, doesn't refuse — it takes the last two dot-separated segments of the address and returns a plausible-looking two-label string (192.168.1.1 -> '1.1'), which passes a truthy `if (domain)` check as if it were a real registrable domain.
Why you never see it
Neither library throws or returns null for these specific inputs — both return an ordinary-looking string, so code doing `if (domain)` or `domainA === domainB` sees exactly the shape it expects and never suspects a boundary error. Manual security testing rarely includes a raw IP literal or a wildcard-hosting subdomain as a test case for domain-comparison logic.

Free fix

Use `psl` (not `tldts`) for anything security-relevant — psl correctly keeps victim.github.io and attacker.github.io as separate registrable domains — but add your own guard in front of it: reject any host shaped like a dotted-numeric address (including leading-zero octets, which some resolvers read as octal and others as decimal — an SSRF-relevant ambiguity) BEFORE calling psl.parse(), since psl itself will happily return a fake two-label 'domain' for an IP literal.

Reproduced this session: 12 real hosts run directly through both the actual tldts and psl npm packages, side by side.

A customer gets charged twice for the same purchase after a slow connection retries or they double-click the pay button — two separate successful charges, two separate PaymentIntents, and nothing in Stripe or your own logs flags it as a duplicate.

critical
What’s actually happening
Creating a PaymentIntent (or Charge) without passing an idempotency key means Stripe treats a retried request as an entirely new one. Two calls with identical parameters — same amount, same customer, same everything — produce two separate, independently-successful PaymentIntents rather than being collapsed into one.
Why you never see it
Both calls return a normal 200 with a well-formed, valid PaymentIntent object — there is no error, no warning field, nothing that looks different from a correct single charge. It only shows up later as a support ticket or a reconciliation mismatch, and only on the minority of requests that actually retry (a client timeout, a flaky proxy, a double-click), which is exactly the path manual and automated testing rarely exercises.

Free fix

Pass a stable idempotency key derived from something that does NOT change across retries of the same logical attempt — e.g. the order id, or a key generated once and persisted before the first attempt — `stripe.paymentIntents.create(params, { idempotencyKey })`. Stripe already refuses a reused key sent with different parameters (a typed StripeIdempotencyError); it is only the missing-key case that is silent.

Measured against the real Stripe API in test mode (evidence/probe.mjs) — real PaymentIntent creations, not inferred from documentation.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

Your scheduled jobs never run — and your dashboard says they succeeded.

high
What’s actually happening
Vercel invokes cron paths with GET. A route that exports only POST returns 405, and Vercel records the invocation as a successful function run.
Why you never see it
There is no error anywhere: no exception, no failed deploy, no alert. The dashboard shows green ticks for a job that has never once executed.

Free fix

Add `export const GET = POST` to every route listed in vercel.json's crons[]. That single line is the entire fix.

Reproduced in production: five cron routes in a live Next.js app silently 405'd from their creation commit until a CI guard caught them.

A number you typed saved as a much smaller number, and nothing warned you.

high
What’s actually happening
Stripping non-digits from '$25M' yields '25', which parses cleanly to 25. The save succeeds and the UI confirms it.
Why you never see it
There is no parse error — the wrong value is a perfectly valid number. It is caught weeks later by someone reading a report.

Free fix

Parse K/M/B/percent suffixes explicitly, refuse ambiguous input instead of guessing, and echo the parsed value back ('Saves as $25,000,000') rather than masking the field.

Reproduced in production: '$25M' stored as 25 with a success toast; found only when a downstream total looked wrong.

The assistant keeps talking over you when you interrupt it.

high
What’s actually happening
Cancelling a turn aborts the outer promise but not the audio playback and model stream beneath it, so speech continues to the end of the buffered sentence while a new turn starts underneath. The two turns then interleave.
Why you never see it
It is invisible in short demo replies, which finish before anyone tries to interrupt. It only shows up with real users and real-length answers — by which point the turn loop is load-bearing and awkward to restructure.

Free fix

Thread one AbortController through transcription, the model stream, AND playback, and abort it when a new utterance starts. Verify by interrupting a long reply — if the voice does not stop within a word, the cascade is incomplete.

Reproduced while building this marketplace's own assistant; the cancellation cascade is covered by a mutation-verified test in that listing.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

Your video export has a frozen frame or a gap that was not in the editor.

high
What’s actually happening
A ripple trim shifts downstream clips by the delta the user REQUESTED rather than the delta actually APPLIED. When the trim is clamped — by the end of the source media or by minimum clip length — the difference becomes a silent gap in the sequence.
Why you never see it
The editor renders the gap as empty track, which reads as normal spacing at most zoom levels. Nothing errors. It surfaces only in the rendered output, usually as a frozen frame or a black flash.

Free fix

Have every trim return what it actually applied after clamping, and shift downstream clips by that value. Then assert zero gaps after each edit — a cheap invariant check catches the whole class.

Reproduced and mutation-tested while building that listing: removing the applied-delta guard opens exactly this gap.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

The same event shows up three times in your feed — or worse, two different events got merged into one.

high
What’s actually happening
Deduplicating on `title | venueId | startHour`. Titles differ across sources by design (presenters prepend, tours append, support acts get listed), venue ids are per-source with no shared identifier, and start times disagree by 30–90 minutes between doors and showtime. Loosening the match to fix the duplicates then merges a tribute act, a parking pass, or a matinee into the headline event.
Why you never see it
Duplicates look like a cosmetic annoyance rather than a matching failure, so the fix is usually a looser comparison — which trades visible duplicates for invisible false merges. A merged event shows fewer results, and nothing anywhere reports that a performance disappeared.

Free fix

Stop truncating start times — compare instants with an explicit tolerance. Then block merges on tokens that change the KIND of event (tribute, karaoke, parking, VIP, afterparty) before scoring similarity at all.

Reproduced while building that listing; the matinee/evening false merge was a real scoring bug caught by its own test suite.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A job runs its side effects twice — a customer is charged twice, an email goes out twice — even though it was added with attempts: 1 and BullMQ's own event log shows only one 'completed'.

high
What’s actually happening
When a processor blocks the Node event loop longer than lockDuration, the lock-renewal timer — which also runs on the event loop — can't fire, so BullMQ's stalled-job detection reassigns the job to a second worker while the first is still running it. Both workers execute the job to completion; only the second worker's 'completed' event fires.
Why you never see it
attempts:1 reads as 'runs at most once,' but it only bounds retries after a FAILURE — a stall is not a failure, so redelivery happens outside that budget entirely. The only trace is the first worker's 'error' event ('Missing lock for job... moveToFinished'), a channel most codebases never subscribe to because completed/failed look like the complete set.

Free fix

Make the job idempotent (idempotency key on the side effect, upsert not insert, dedupe downstream) and move CPU-bound work off the event loop (worker_threads, chunk it with setImmediate/setTimeout breaks). Subscribe to the worker's 'error' event so a lost lock is at least visible. Raising lockDuration alone does not fix it — the renewal timer runs on the same blocked loop.

Measured against real BullMQ 6.0.8 + Redis on 127.0.0.1:6399: two workers, one job, attempts:1, worker-A blocks its event loop 4s with a 1s lock.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A hand-rolled HTTP chunked-transfer-encoding parser returns a normal-looking body for input a real web server rejects with 400 — including input that has a second, hidden HTTP request appended after the first one.

high
What’s actually happening
A short decoder built from indexOf('\r\n') + parseInt(line,16) + slice tolerates bytes RFC 9112 section 7.1 requires rejecting: whitespace or a tab after the chunk-size, a bare LF with no preceding CR, a negative or hex-prefixed ('0x5') size, a doubled CR. Worst case: bytes trailing the terminating chunk are simply dropped from the returned body instead of triggering a rejection, so a second request smuggled after a well-formed chunked body vanishes with zero indication it was ever there.
Why you never see it
Every one of these inputs returns a plausible string, never an exception. Tests written against well-formed input all pass; the gap only shows up by diffing the naive decoder's output against a real server (Node's own http.createServer/llhttp) on adversarial bytes, which nobody does for a function this short.

Free fix

Don't hand-roll chunked-body parsing — decode via the platform's own HTTP stack (Node's http module / undici), which already implements RFC 9112 correctly, instead of a bespoke indexOf/parseInt/slice function. If raw bytes must be parsed outside a real HTTP server, reject (don't return a body) when trailing bytes remain after the terminating chunk, reject whitespace after the chunk-size line, and reject any LF not immediately preceded by CR.

Measured: a naive decoder (indexOf/parseInt/slice, written cold) run against 11 adversarial byte sequences drawn from RFC 9112's ABNF and cross-checked against a real, unmodified Node http.createServer over a raw socket.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A cache keeps serving a value that's older than the database, permanently, with no error anywhere — even though the write path correctly calls cache-invalidate on every write.

high
What’s actually happening
A cache-aside pattern has three events, not two: a DB write, a cache invalidate, and — on a separate request, sometimes much later — a cache repopulate by whoever next reads the key. If a repopulating read started before a write landed, its (now stale) result can still be written into the cache AFTER the invalidate meant to guard against exactly this, leaving the cache holding pre-write data with nothing left to correct it. A detector or a person checking only 'did invalidate fire after this write' never looks at when the repopulating read actually started, and calls this fixed.
Why you never see it
Every individual step is correct code in isolation: the write invalidated the cache exactly as intended, and the stale read was a valid read of an earlier, real state. The bug lives only in the relationship between the read's start time and the write's time, which ordinary logging doesn't capture — and the wrong outcome is a normal-looking, successfully-returned cached value, never an error.

Free fix

Track when a repopulating read actually STARTED (not when it lands) and compare that against write timestamps for the same key — a populate is stale if any write happened after its read started, regardless of when the populate itself arrives. Cheaper backstops that also work without that plumbing: a short TTL, or delete-then-write with a delayed second delete ('double delete') so a racing populate gets cleared shortly after it lands.

Measured: a naive write-then-invalidate detector (the literal one-line reading of 'check whether cache_invalidate happened after db_write') run against 5 adversarial event-log cases modeling this exact race.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A follow-up retry batch quietly stops retrying certain failed items forever, or re-runs items that already succeeded (a customer gets billed twice, an email goes out twice) — and the batch call never throws or logs anything either way.

high
What’s actually happening
Code reads a per-item batch result (bulk email, bulk row writes, an agent tool call over N items) to decide what to retry. The obvious one-line implementations are each wrong in a way that returns a normal-looking result: treating any non-throwing call as fully successful drops every per-item failure; retrying the whole batch on any failure double-runs the already-succeeded items; and the most careful-looking version, `if (item.retryable)`, reads an upstream API's omitted/undefined retryable field as falsy and silently drops that failure from retry forever.
Why you never see it
None of the three shapes throw; each produces a validly-typed JSON object, so nothing downstream rejects it. The `if (item.retryable)` version looks the most correct of the three — it does check per item — which is exactly why its undefined-vs-false bug survives review: it reads as the fix, not as a second instance of the bug.

Free fix

Classify every item into exactly two buckets — retry or done — with an explicit truth table: success -> done, confirmed-permanent-failure -> done, everything else (unconfirmed, unknown, retryable-unspecified) -> retry. This mirrors how AWS Lambda's SQS/Kinesis partial-batch-response and Stripe's idempotent-retry guidance both handle it. Never let an omitted field silently default to 'skip' — an absent retryable must mean retry-eligible, not done.

Measured: three naive retry-split implementations (transport-only success check, any-fail-retry-all, per-item `if (item.retryable)`) written cold from the same one-line brief, each run against the same 5-case adversarial batch-result set.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

Your data pipeline dashboard is all green, but a number that summarizes a column has been wrong for weeks and there is nothing anywhere in the warehouse that reveals it.

high
What’s actually happening
dlt's default schema_contract, 'evolve', absorbs any upstream shape change automatically: a type change creates a shadow column (e.g. amount__v_text) and leaves the original column NULL for those rows, and an unrequested new field is created and populated with no review. The other two non-freeze modes lose data differently — discard_row drops the whole row, discard_value nulls just the changed field — but all three report success.
Why you never see it
All three non-freeze contracts load successfully and leave a value of the right type in the right place — a NULL is indistinguishable from a value that was legitimately absent, a shadow column looks like a deliberate schema addition, and a dropped row leaves no trace it ever existed. The downstream aggregate lands on the identical wrong number regardless of which of the three losses happened, so there's no discrepancy for anyone to notice.

Free fix

Pass schema_contract explicitly — use 'freeze' if any shape change should stop the run rather than silently absorb it — instead of relying on dlt's implicit 'evolve' default, and add a row-count or null-rate check between source and destination. That turns any of these three losses into an alert instead of a permanently wrong aggregate. No purchase needed for either half of this fix.

Own measurement against dlt 1.29.1 (Apache-2.0): a two-row pipeline sending {id, amount} received a type change (amount becomes a string) plus an unrequested new field on row 2, run under all four schema_contract settings.

You edit your .env file, restart the app, and it keeps using the old value — with no error, warning, or log line anywhere telling you why.

high
What’s actually happening
dotenv's config() defaults to override:false: if a variable of the same name already exists in process.env (from a shell rc file, docker run -e, launchd/systemd, a CI secret, a parent process, or an editor launched from Dock/Spotlight with a different inherited environment), the .env file is parsed but its value is discarded. config() still reports success, and result.parsed still contains the file's new value even though it was never applied to process.env.
Why you never see it
The obvious debugging step — print what dotenv parsed — shows exactly the value you just typed into the file, which reads as confirmation the fix worked. The only place the truth is visible is process.env itself, the one thing nobody thinks to check because they believe they just set it.

Free fix

Call dotenv.config({ override: true }) if the .env file should always win, or check process.env[NAME] itself — not result.parsed — after the call before concluding a config change took effect. dotenv's own 'injected env (N)' tip line is a free built-in tell when N is 0; it's just easy to miss next to the library's ad line beside it.

Own measurement against dotenv 17.x on Node 26: a shell-set MYKEY survives a .env file declaring a different value, with zero errors reported at any point.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A numeric column read from a CSV comes out full of zeroes, and every check you have — schema, type, range, NOT NULL — passes, because zero is a valid number.

high
What’s actually happening
csv-parse's cast:true decides whether a cell is numeric with `value - parseFloat(value) + 1 >= 0`, whose subtraction coerces the cell using Number() — which honours 0x, 0b and 0o prefixes. It then converts the cell with parseFloat(), which does NOT honour them and stops at the letter following the leading zero. The gate admits the cell on Number()'s reading and a different function converts it, so 0xFF passes the test as 255 and is stored as 0.
Why you never see it
Nothing raises and nothing is logged — cast:true reports success. The result is not malformed, it is a perfectly ordinary zero, so it survives every downstream defence: a numeric schema accepts it, a range check accepts it, NOT NULL accepts it. The loud neighbours mislead too: cells the gate rejects (0xZZ, 1_000, Infinity) come back as untouched strings and blow up on the next arithmetic, which trains people to believe bad values announce themselves.

Free fix

Leave cast off for columns that may carry radix-prefixed values and convert them yourself with Number(), which honours 0x/0b/0o — or pass a cast function that handles those prefixes explicitly. To detect it after the fact, compare Number(rawCell) against the cast result for a sample of raw cells; where both are finite and they disagree, the cast silently changed the value.

Own measurement against csv-parse 7.0.2 on Node v26.0.0, plus the cast implementation read directly from dist/esm/index.js. Re-implementing that gate and conversion and replaying them over all 20 measured literals reproduces csv-parse's actual output with 0 disagreements.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

An id or amount you look up doesn't match the one you were sent, even though nothing errored and both look like ordinary numbers.

high
What’s actually happening
JavaScript numbers are IEEE-754 doubles, exact only up to 2^53-1. JSON.parse() converts every JSON integer into a JS number regardless of size, so a 64-bit id (a database bigint, a snowflake-style id) or a large integer currency amount above that threshold is silently rounded to the nearest representable double during parsing.
Why you never see it
The parsed value is a completely normal-looking number — right type, plausible digit count — that simply isn't the number that was sent. JSON.parse's contract gives no signal of this; the corruption is invisible until something downstream compares the round-tripped value against the original and finds they disagree.

Free fix

Have the source send large integers as strings (many APIs already offer this, e.g. an *_str-suffixed id field) and keep them as strings/BigInt throughout, rather than letting JSON.parse coerce them into doubles. Free — no dependency change required if the API already offers a string form.

Own measurement in plain Node, no dependencies: JSON.parse on a payload with an id above 2^53-1 and a large integer amount.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A page shows the exact same numbers forever after you deploy it, even though the data behind it changes constantly — and the build log doesn't flag anything wrong, it just marks the route "○ Static," which reads like a performance win, not a bug.

high
What’s actually happening
`await fetch(url)` in a server component, with no cache option and no route-segment config, gets cached by Next's Data Cache and the route is prerendered at build time — it serves the build-time snapshot forever. On Next.js 14, adding `export const dynamic = "force-dynamic"` does not fix this: it forces the ROUTE to re-render on every request (provable — a per-render marker moves), but the underlying fetch is still served from the cache, so the value printed on the page stays frozen while the route visibly re-executes around it.
Why you never see it
Nothing throws and nothing warns. The build output labels the route Static, which most developers read as a good sign, not a red flag. On Next 14 specifically, the force-dynamic 'fix' looks like it worked — the build output now says the route is dynamic — while the actual bug (a cached fetch) is completely untouched.

Free fix

Use `fetch(url, { cache: "no-store" })` or `export const revalidate = 0` — both verified live (i.e. the number actually updates) on Next.js 14.2.35, 15.5.22, and 16.3.0. force-dynamic alone is not sufficient on Next 14.

Reproduced this session: identical app code, same server component pattern, run against Next 14.2.35, 15.5.22, and 16.3.0 with a live origin that increments on every read.

A user uploads an image and your server's memory spikes or the process gets OOM-killed — and there's a warning buried in your logs from before it happened that nobody treated as an error, because nothing about it stopped the code from continuing.

high
What’s actually happening
Pillow's `MAX_IMAGE_PIXELS` guard only RAISES an exception above 2x its configured limit. Between 1x and 2x the limit, `Image.open(...).load()` succeeds and allocates the full image in memory — it only emits a `DecompressionBombWarning`, which under Python's default warning filters does not raise or abort anything. Code that checked 'do we have a decompression-bomb guard configured' (yes, the default is non-zero) is still fully exposed in that entire 1x-2x band.
Why you never see it
The commonly-published fix for the warning message — `Image.MAX_IMAGE_PIXELS = None` — is a single global, process-wide assignment that any imported module can execute at import time, before your own code runs. It doesn't silence the warning; it removes the check entirely, with zero trace afterward that it was ever disabled.

Free fix

Check `width * height` against your own application-appropriate limit BEFORE calling `.load()` — don't rely on the process-global default catching it, since the default's dangerous band (1x-2x) loads fine and only warns. And never let `MAX_IMAGE_PIXELS = None` exist anywhere in your dependency tree; there is no per-call override, so the process-wide value is the entire mechanism.

Reproduced this session against Pillow 12.3.0: PNGs generated locally at chosen pixel dimensions, opened and load()ed with warnings captured.

You open a spreadsheet with openpyxl, save it right back out — maybe just to normalize formatting or add a sheet — and every formula cell is now permanently blank, even though the file still opens fine in Excel.

high
What’s actually happening
`load_workbook(data_only=True)` does not evaluate formulas — it returns whatever value Excel last cached inside the file for that cell. openpyxl (like xlsxwriter and pandas) never writes that cache when it saves a file, so any workbook Excel hasn't touched has no cached values at all: every formula cell reads back as `None`. If that workbook is then saved by openpyxl, the `None` values are written back in place of the formulas, deleting them permanently.
Why you never see it
Nothing raises anywhere in the load-compute-save sequence. The workbook opens normally, the sheet names and constant values are all correct, and only the formula cells are silently `None` — which reads exactly like empty cells a user forgot to fill in, not like a parsing failure with a root cause to trace.

Free fix

Load the file twice — once with `data_only=True` to read cached values for display/computation, once without to preserve the actual formula text — and never call `.save()` on the data_only=True workbook. When the file's origin is unknown, treat 'were these values actually cached by Excel' as an open question rather than an assumption baked into the read path.

Reproduced this session against openpyxl 3.1.5: a workbook built with a formula cell, loaded both with and without data_only, and round-tripped through save.

A recurring billing date, reminder, or meeting slowly drifts off its original day — a Jan 31 monthly charge starts landing on the 28th every month after that and never returns to the 31st — or a scheduled time is off by exactly one hour after a daylight-saving change. No library ever raised anything.

high
What’s actually happening
There are two natural ways to write 'generate the next N recurring occurrences,' and both are silently wrong. (1) Chaining: deriving occurrence N from occurrence N-1 instead of from the original anchor, so a month-end clamp (Jan 31 -> Feb 28) permanently 'sticks' and every later cycle inherits the already-clamped date instead of recomputing from the 31st. (2) Fixed-epoch-add: converting the anchor to a UTC epoch once, then adding `n * intervalMs` — correct until a DST boundary is crossed, after which every subsequent occurrence is off by exactly the DST offset.
Why you never see it
Every mainstream JS date library (date-fns, dayjs, luxon) computes a SINGLE call from the anchor correctly — the bug isn't in any library, it's in the loop structure wrapped around it, so 'we use a trusted date library' gives false confidence. The output at every step is a valid, plausible-looking date or time; nothing throws, and the drift only becomes visible if someone compares a later occurrence directly back to the original anchor.

Free fix

Recompute every occurrence directly from the original anchor (`anchor + n * interval`), never by chaining off the previous occurrence, and do date arithmetic in the IANA timezone via wall-clock components rather than fixed UTC-millisecond deltas — verified correct across date-fns, dayjs, and luxon when done this way.

Reproduced this session in Node 26 (ESM) against the real date-fns, dayjs, luxon, and rrule packages, both as single calls and inside both flawed loop patterns.

A number extracted from a scanned or rotated PDF table is a completely different, still-perfectly-plausible number — a quantity of 1250 comes out as 521 — and the extraction call that produced it never raised or warned about anything.

high
What’s actually happening
On PDF pages whose `/Rotate` entry is 180 or 270, pdfplumber returns every text cluster's characters in EXACTLY REVERSED order via `extract_text()`, `extract_words()`, and `extract_table()`, while the row/column bounding-box geometry stays completely correct — so the table structure looks entirely right and only the text content within each cell is backwards. A reversed digit string is still a syntactically valid number (leading zeros in the reversed form simply vanish on parse), so a downstream `int()`/`float()` call succeeds silently on the wrong value.
Why you never see it
Only the character order within a cluster is affected, and only at 2 of the 4 possible rotation values (180 and 270 — not 0 or 90), so testing happens to use an unrotated fixture, or testing uses non-numeric text where 'backwards' is visually obvious, both miss it entirely. pdfplumber's own suggested workaround parameter, `text_use_text_flow=True`, was measured to make the corruption WORSE (fragmenting cells character-by-character), and neither `extract_table` nor `extract_text` has any docstring describing the behavior at all.

Free fix

Check `page.rotation` before trusting any extracted text; on rotation 180 or 270, reverse each character cluster back to its correct order before parsing numbers out of it (the row/column geometry itself needs no correction — only the per-cluster text does). Do not use `text_use_text_flow=True` as a workaround — it was measured to make the corruption worse, not better. No documented pdfplumber fix or parameter avoids this.

Reproduced this session against real pdfplumber output on a page with a controlled /Rotate value, including a real extract_table() call whose corrupted output string was captured and pinned into a test fixture.

A background side effect fired from inside an API route — an audit-log write, a notification, an analytics event — sometimes never happens, and other times shows up minutes late attributed to a completely unrelated request. The route itself always returns 200 with no error.

high
What’s actually happening
On Vercel's Node serverless runtime the execution environment is frozen the instant the response is returned. An unawaited promise is not cancelled — it is left mid-flight and only resumes if and when a LATER, unrelated request thaws the same execution context, at which point it runs charged against that request's duration and timeout. If no later request lands before the environment is recycled, the frozen work never resumes at all.
Why you never see it
The handler returns a clean 200 whether the deferred work eventually runs or not, so nothing in the response signals a problem. Worse, testing the identical code with `next dev`/`next start` locally hides it completely: a local Node process never freezes between requests, so a dangling unawaited promise simply resolves on schedule and every local test passes.

Free fix

Either `await` the work before returning, or wrap it in `waitUntil()` from `@vercel/functions` (`import { waitUntil } from '@vercel/functions'; waitUntil(sideEffectPromise);`) so the platform keeps the environment alive until it finishes. A longer timeout does not help — the environment is frozen, not slow, so extra time just waits for a thaw that may never arrive.

Measured by deploying three probe routes to real Vercel serverless (region iad1, Next.js 15.5.22) and reading the platform's own runtime logs back over HTTPS — not simulated, not inferred from docs.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A date column imported from an uploaded Excel file turns into a plain number like 46085 somewhere downstream, or a date that should read the 4th gets filed as the 3rd for some users — and the import itself never throws an error.

high
What’s actually happening
Excel does not store a date type — it stores a plain number with a display format attached. SheetJS's default read (`cellDates:false`) hands back that raw number, a perfectly valid float that gets stored, compared and summed as a number. Turning on `cellDates:true` fixes that but converts to a JS Date at LOCAL midnight expressed as a UTC instant, so reading the UTC calendar day back shifts the date by one for any reader in a negative UTC offset.
Why you never see it
Both outputs are fully valid values of their own type — a legitimate number, a legitimate Date — so nothing throws or warns either way, and the two SheetJS read options actively disagree with each other on edge cases (Excel's serial 60, the fictitious Feb 29 1900 kept for Lotus 1-2-3 compatibility), so there is no single flag that is simply 'the correct one' to flip.

Free fix

Read with `cellDates: true`, and before reading the date back out, force the parsing process's own clock to UTC (e.g. `process.env.TZ = 'UTC'` set before any parsing runs, or an equivalent per-process override) — then use the UTC getters (getUTCFullYear/getUTCMonth/getUTCDate). Neither UTC nor local getters are safe on their own: UTC getters only avoid the day-shift when the parsing machine's own timezone is at or west of UTC, and read a day early on machines east of UTC (confirmed for Tokyo, Auckland, Kiritimati). For any data that could reach pre-March-1900 serials (historic records, birth dates), add an explicit check for serial 60, since SheetJS's own two read options disagree about which calendar day it is.

Measured against xlsx 0.18.5 (SheetJS Community Edition) by running evidence/measure-cell-types.mjs and printing the actual returned cell values — not inferred from documentation.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

An OCR-extracted number is occasionally wrong by a factor of ten, or has an extra digit inserted, even though your confidence-based quality gate passed it through at 80%+ confidence.

high
What’s actually happening
Tesseract's per-word confidence measures how sure the classifier is about the glyphs it segmented, not whether it segmented the number correctly in the first place. A blurred number can be SPLIT into two separate word tokens, each individually a passable confidence, even though the resulting parsed number is off by 10x; a mildly-degraded number can gain a spurious inserted digit and still score in the high 70s. Aggregating by mean compounds this: one genuinely bad token gets diluted by good neighbors, so a mean-based gate admits a line a minimum-based gate would have caught.
Why you never see it
Every wrong value reported is a syntactically valid word at a valid confidence score. There is no error path and no word-level signal distinguishing a segmentation failure from a correct read, so the number simply flows downstream looking exactly like a trustworthy one.

Free fix

Gate on the MINIMUM word confidence within a field, not the mean — one bad token should fail the whole field. For numeric fields specifically, add a sanity check independent of OCR confidence (checksum, expected digit count, or cross-reference against another field), since confidence alone cannot detect a segmentation error.

Measured against tesseract 5.5.3 / pytesseract 0.3.13 by running evidence/measure-tesseract-confidence.py against real degraded test images — not inferred from documentation.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

Your AI app looks like every other AI app, and more prompting does not fix it.

medium
What’s actually happening
Models converge on the same layout and visual defaults regardless of prompt detail, so iterating burns time without moving the design.
Why you never see it
Nothing is broken. Each attempt looks fine in isolation; only side-by-side with a distinctive product does the sameness show.

Free fix

Start from a real design system rather than prompting for one; free options include shadcn/ui primitives and the Vercel AI Elements set.

Independently documented (2026 'AI-generated UI curse' analyses) and reproduced by this marketplace's own founder across many attempts.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A reordered, deduplicated event log looks perfectly clean — no gaps flagged, no duplicates flagged — but silently splices a crashed-and-respawned worker's new output into the middle of its old, stale output as if it were one continuous stream.

medium
What’s actually happening
Out-of-order events from a resumable producer (a queue consumer, a reconnecting WebSocket client, a respawned worker) get reordered and deduplicated keyed only on (writerId, seq). A bare sequence number can't distinguish 'this writer already sent seq 1' from 'a NEW incarnation of this writer legitimately restarted its own counter at 1' after a crash — so a global-watermark implementation drops the new incarnation's entire stream as duplicates, and even the natural per-writerId fix keeps the events but interleaves old-epoch and new-epoch payloads into one gapless, dedup-clean-looking sequence.
Why you never see it
The global-watermark version's failure (silently discarding a whole second producer's stream) is easy to spot once you look for concurrent writers, but the natural fix — per-writerId state — resolves 4 of 5 adversarial concurrency cases correctly and only fails on writer restart, precisely the case least likely to be in anyone's test suite. The failing output isn't an error or a visible gap; it's a tidy, contiguous-looking log with two unrelated process incarnations' payloads merged into it.

Free fix

Key every event on (writerId, instanceId, seq) — mint a fresh instanceId (a UUID, a process-start timestamp, anything unique per process lifetime) every time a writer starts or restarts. Never reuse writerId alone as the dedup key; this is the same problem TCP solves with Initial Sequence Numbers, for the same reason.

Measured: a naive single-global-watermark implementation written cold and run against 5 concurrency scenarios, then the 'obvious fix' (per-writerId state) built and measured against the identical 5.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A list you just sorted comes back in almost the same order it went in, as if .sort() silently did nothing.

medium
What’s actually happening
Array.prototype.sort() requires its comparator to return a negative, zero, or positive NUMBER. A comparator written as (a, b) => a > b returns a boolean, which JavaScript coerces to 1 or 0 — never negative — so the engine concludes almost every adjacent pair is already 'equal' and leaves the array close to its input order.
Why you never see it
Nothing throws or warns. The output is an array of the right length, containing the right elements, in an order that reads as plausible — indistinguishable from a list that genuinely was already sorted unless you diff it against the correctly-sorted result.

Free fix

Return a signed number from the comparator — (a, b) => a - b for ascending numbers — never a boolean. This is the entire fix, and it's exactly the mistake reached for when asked to "sort numbers" without being told to subtract.

Own measurement in plain Node, no dependencies: a mixed-order numeric array run through a boolean comparator vs. a subtraction comparator.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A date-only value files under the wrong calendar day for some users — one day earlier — depending purely on which part of the codebase constructed it.

medium
What’s actually happening
The ECMA-262 Date spec parses a dash-separated date-only string ('2026-03-04') as UTC midnight but a slash-separated one ('2026/03/04') as LOCAL midnight. For anyone west of UTC those are different instants, so reading the calendar day back with .getDate() disagrees by one day depending purely on which separator the source string used.
Why you never see it
Both new Date('2026-03-04') and new Date('2026/03/04') are valid Date objects — no exception, no NaN, no warning. It's invisible in UTC-based CI environments and only surfaces for users or servers running west of UTC.

Free fix

Never hand a date-only string straight to new Date(); parse explicitly with a library (date-fns parseISO, Temporal.PlainDate) or build the date yourself with new Date(Date.UTC(y, m-1, d)) so the string's separator can't silently pick the timezone.

Own measurement in plain Node, host timezone America/Los_Angeles: new Date('2026-03-04') vs new Date('2026/03/04').

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A photo or scanned document is missing part of its content — the bottom section is replaced with flat grey — but the file 'opened fine,' nothing was logged, and nobody would know to look unless they happened to view that exact image.

medium
What’s actually happening
`ImageFile.LOAD_TRUNCATED_IMAGES = True` (commonly set to silence a real `OSError: image file is truncated` that was crashing on genuine partial uploads) makes Pillow fill in the missing image data with a flat placeholder color instead of raising. It is a process-global flag: any module anywhere in the process — including a dependency you've never read — can set it, silently changing the load behavior of code elsewhere that never opted in.
Why you never see it
`Image.open()` itself is lazy and never raises on a truncated file at all — size and mode both report correctly at that point; only the later `.load()` call reveals anything, and with the flag set, even `.load()` succeeds outright with the wrong pixel data. Nothing about the object handed back looks wrong at any inspection point.

Free fix

Don't set LOAD_TRUNCATED_IMAGES=True as a blanket workaround for the crash-on-truncation error; instead catch the OSError from `.load()` at the default (False) setting and explicitly reject the truncated upload. If truncated images must be tolerated, use `.verify()` on a SEPARATE file handle first, not the one you intend to `.load()` — verify() consumes the handle it's called on, so a subsequent load() on that same object fails with an unrelated AssertionError instead of the real error.

Reproduced this session against Pillow 12.3.0: a real JPEG truncated to 50% of its byte length, opened and loaded with both settings of the flag.

A perfectly clean, correctly-read document gets silently thrown out by your OCR confidence gate — a page a human confirms was read 100% correctly reports roughly 30% confidence.

medium
What’s actually happening
pytesseract's image_to_data returns one flat table that interleaves four non-word hierarchy levels (page, block, paragraph, line) with the actual word rows. Every non-word structural row carries a confidence sentinel of -1. Averaging the raw 'conf' column — the natural thing to do with a numeric column of that name — drags a perfect read's average far down by counting four -1 sentinels as low-confidence readings alongside two genuinely high-confidence words.
Why you never see it
pytesseract's own docstring for image_to_data reads, in full, 'Returns string containing box boundaries, confidences, and other information' — it never documents the multi-level structure or the -1 sentinel. Averaging the column returns a syntactically valid float in-range, so nothing errors; it just quietly reports the wrong number for a page that was read perfectly.

Free fix

Filter to level==5 (word) rows — or simply drop any row with conf == -1 — before computing an average, and gate on that true word-level mean, never the raw column.

Measured against tesseract 5.5.3 / pytesseract 0.3.13 by running evidence/measure-tesseract-confidence.py and printing the full returned table — not inferred from documentation.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

A scraped-content pipeline's 'article' text turns out to be a navigation menu or a paywall notice repeated across many pages of the index or dataset — and the extraction step itself never errored or returned empty.

medium
What’s actually happening
trafilatura's extract() returns a plain string or None, and returning a non-empty string is the only signal most callers check. There is no confidence, score, or quality field anywhere in its output. A page that is pure navigation/footer, or an article hidden behind a paywall stub, produces an ordinary non-empty string that passes any `if text:` guard exactly like a real article does.
Why you never see it
Length cannot separate the two cases: a measured paywall stub (39 characters) and a genuinely short real article (41 characters) are nearly the same size, so the obvious guard — `len(text) > 50` — either lets paywall junk through or rejects legitimate short content depending on where the threshold lands, consistently and silently, which is exactly why nobody re-checks it once the pipeline ships.

Free fix

Don't trust a truthy string. Compute link-text ratio (characters inside <a> tags divided by total extracted characters) and sentence count as a post-extraction filter — the measured cases separate cleanly on those two signals (nav-only page: ratio 0.90, 0 sentences; real content: ratio 0.00, 6+ sentences) even though they do NOT separate on length alone. Treat None from extract() as its own explicit outcome, distinct from 'extracted but low quality.'

Measured against trafilatura 2.2.0 by running evidence/measure-extraction.py against constructed HTML fixtures and printing trafilatura's real output — not inferred from documentation.

If you’d rather not hand-roll it, we sell a tested pack: see the listing → (the free fix above works either way)

For agents

Public, no key required. Post what you can observe about the project; get back the failure modes it’s exposed to, most severe first.

curl -X POST https://coderecycle.ai/api/v1/detect-risks \
  -H "Content-Type: application/json" \
  -d '{
    "stack": ["next.js"],
    "hosting": ["vercel"],
    "deps": ["stripe", "@supabase/supabase-js"],
    "signals": ["vercel.json contains crons[]", "webhook handler"]
  }'

Detection is deterministic — no model call, no network beyond this request, and nothing about your project is stored.