Data went missing and the job reported success
A green pipeline and a shorter table. Rows dropped, columns coerced, values nulled — each one a successful run.
Curated by Code Recycle Editorial
- 1
SUM(amount) returns the same wrong number under three of dlt's four schema contracts, from three different silent losses -- and the lossy default is the one you get.
Why it's here: SUM(amount) is the same wrong number under three of four contracts. The default is lossy.
- 2
The load SUCCEEDS. The new column is dropped, the retyped one is coerced, and the warehouse quietly diverges.
Why it's here: The load SUCCEEDS. The new column is dropped and the warehouse quietly diverges.
- 3
load_workbook(data_only=True) reads every formula as None on any file Excel never saved -- and saving writes the None back, emptying the formulas permanently.
Why it's here: Reads every formula as None — and saving writes the None back, permanently.
- 4
Insert one row between page 1 and page 2 and OFFSET pagination silently skips a record the reader never sees.
Why it's here: One insert between pages and OFFSET skips a record the reader never sees.
- 5
Truncation happens at a byte offset, never a structural one. A JSON payload cut mid-string parses into something silently incomplete.
Why it's here: Truncation at a byte offset, never a structural one. It still parses.
- 6
The recommended token-budget splitter, on 100% default settings, returns chunks whose real re-tokenized length is over the budget you asked for. Nothing warns you; the embedding call just costs more or gets truncated.
Why it's here: Chunks whose real re-tokenized length is over the budget you asked for.