Writing
How the data behind this site is collected and served. The code that does it is for sale.
Scraping TikTok's Mobile API
Four unrelated things have to be right before TikTok answers you, and getting any one of them wrong returns a clean HTTP 200 with an empty body and no hint which. Device registration and activation, the X-Argus cipher stack, regional hosts, TLS fingerprints, and the 24 endpoints on the other side.
Read the write-up →Tracking a Billion TikTok Sounds
A billion reads a day, and the obvious way to do it needs four million throwaway devices. Device economics, a ClickHouse schema with four silent-corruption traps, and fourteen measurements that turned out to be measuring my own instruments.
Read →Querying Billions of Rows in Milliseconds
A search box over a few billion scraped rows, answered in 20 ms. ClickHouse keeps every row, Elasticsearch serves the page, and a nightly loader moves between them. Includes the four watermark designs I wrote that disagree with each other.
Read →Searching Video by What It Looks Like
Most captions on short video are emoji and hashtags, so the frame is the only honest description. A hundred million thumbnails through SigLIP2, why the vectors ended up in the search index rather than a vector database, and the scoring bug I shipped twice.
Read →