· via Hacker News – Front Page (native)
ParadeDB answers PlanetScale's TIN with Postgres full-text search speedups
ParadeDB says two weeks of profiling and optimization closed the BM25 performance gap with PlanetScale's new TIN Postgres extension, without rearchitecting how its index identifies documents.

ParadeDB has published a detailed technical response to PlanetScale's TIN, a full-text search extension for Postgres that launched two weeks earlier with benchmarks showing large performance wins over ParadeDB's own search functionality. According to the ParadeDB blog post by Ming Ying, TIN measured at least eight times faster than ParadeDB 0.25 in every benchmark PlanetScale ran. Two weeks of profiling and optimization later, ParadeDB says it has closed that gap for BM25-ranked Top K queries — and that the way it closed it undercuts PlanetScale's explanation of why TIN was faster in the first place.
The dispute over document identifiers
ParadeDB's index is powered by Tantivy, a search library that identifies documents with compact sequential numbers assigned in insertion order. Postgres, meanwhile, locates rows through ctid values, which point to a row's physical position in the database's block-based storage. Because ParadeDB bridges the two systems, it must maintain a mapping between Tantivy document IDs and Postgres ctids.
PlanetScale's launch post attributed much of TIN's advantage to eliminating that mapping: by keying its index directly on ctids, TIN skips the translation step and can run efficient bitmap operations and visibility checks. ParadeDB concedes the logic for COUNT queries, where millions of matches each need a visibility check and the translation cost compounds. For BM25 Top K queries, though, ParadeDB was skeptical, because its engine already delays ctid lookups until the final results are assembled — a top-10 query needs only ten lookups, far too small in the profile to explain an orders-of-magnitude gap.
Fixing fieldnorm locality
The first optimization targeted how ParadeDB reads fieldnorms, single-byte values that encode document length and let BM25 normalize scores. A new ParadeDB feature that attributes page accesses to the data structures behind them showed fieldnorms drove 83 percent of page reads on a single-term query against a 28.7-million-document Hacker News dataset — roughly 1,500 pages, against about 300 for everything else.
The problem was locality. Tantivy stores fieldnorms in one shared array indexed by document ID, separate from the postings lists, so scoring a term means jumping around that array. That layout is cheap in Tantivy's usual memory-mapped setting, but inside Postgres it meant touching hundreds of distinct pages per query. ParadeDB's fix was to store a fieldnorm array alongside each postings list, in the same order, so both are read sequentially. Fieldnorm page accesses fell from about 1,500 to 30. The tradeoff is storage: a document's fieldnorm is now repeated for every distinct term it contains, which grew the Hacker News index by about 9 percent, though ParadeDB notes most terms have short postings lists in real corpora.
An algorithmic bottleneck in Blockmax WAND
The second fix addressed multi-term disjunction queries. After the fieldnorm change, buffer reads for such queries fell by roughly 80 percent, yet query times dropped only about 5 percent — a sign the remaining bottleneck was algorithmic rather than I/O. Profiling placed most of the time in the Blockmax WAND loop, the standard technique search engines use to skip over chunks of postings while processing ranked queries. The post, which focuses on Top K optimizations, promises COUNT improvements in a second part.
How the comparison was run
ParadeDB reran the before-and-after benchmark on the same 150-million-document StackExchange dataset, harness and machine types as PlanetScale's tests, though TIN itself is not open source and ran on PlanetScale's platform. ParadeDB also says several benchmark configuration settings in the original comparison were not fully fair, with details promised in a follow-up. It credits the PlanetScale team for its engineering, and notes that PlanetScale adopted the benchmarker tool ParadeDB built for exactly this kind of testing.
Why it matters
The exchange is a live case study in Postgres-native search engineering. Rather than accepting that a rival's architecture was fundamentally superior, ParadeDB used page-level profiling to turn an orders-of-magnitude gap into two concrete, fixable problems — and closed it within two weeks. For teams running BM25 search inside Postgres, this competition is rapidly raising the baseline of what such extensions can do. It also illustrates how benchmark claims should be read: architectural explanations hold up only until someone profiles the actual workload, and configuration choices can skew results as much as design decisions. Finally, with TIN closed-source and ParadeDB open, independent verification of future comparisons will depend on shared tooling like the benchmarker both teams now use.
- #postgres
- #full-text-search
- #performance
- #databases
- #paradedb