deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Node.js typed arrays beat JSON.parse on a million-row report and end OOM crashes

A dev.to engineering write-up shows how two typed arrays replaced a million JSON objects on a 2 GB Node.js server, going from out-of-memory crashes to flat memory and faster-than-JSON.parse loads.

Node.js typed arrays beat JSON.parse on a million-row report and end OOM crashes

A query that kept killing the process

An air-quality monitoring backend at Oizom ran on a 2 GB DigitalOcean droplet, and report queries spanning roughly a million sensor readings kept crashing the process with out-of-memory errors. According to a dev.to write-up by Panth Patel, who leads the company's software team, each reading was a small JSON-style object holding a device identifier, an epoch timestamp, a map of numeric measurements such as CO2 or PM2.5, and a map of free-form labels like location or firmware.

The post attributes the blow-up to three costs that never show up in application code. Property access on such objects follows chains of references instead of a single read at a known offset. Each object is backed by a hash table, so one pass over a million ten-key objects means something on the order of ten million hashed lookups. And every object is a tracked allocation, so the garbage collector ends up working harder than the query itself, with boxed numbers scattered across the heap.

Laying readings out like a grid

The core idea in the post is to put timestamps on one axis and measured quantities on the other, turning the whole dataset into a rectangle of numbers — essentially a sparse image — indexed with the familiar row-times-width-plus-column calculation.

The storage breaks down like this:

  • A single Float64Array holding every decimal value, sized rows times columns.
  • An Int32Array for label values, where each entry is an index into a dictionary of unique strings. Labels such as a city name repeat on almost every row, so there is no reason to store them a million times.
  • Three small axis arrays: timestamps, gas names and label names.

Sentinel values mark the gaps: timestamp zero denotes an unused row and an empty string an unused column. Reading a cell is plain offset arithmetic, and the entire table is one allocation instead of a million objects.

One class hides the raw buffers

Because the buffers are awkward to use directly, everything sits behind a DataPointTable class. It offers factory methods for creating a table, loading from JSON and hydrating from an ArrayBuffer; axis lookups and mutation helpers; cell get and set; row iteration that yields name-value pairs; plus toJSON, toBuffer for shipping to the browser, a compress step that reclaims unused rows and columns, and a stats method. The rest of the codebase never has to know a Float64Array exists underneath.

Version one was correct — and 8x slower

The first implementation kept rows sorted by descending timestamp, since queries want time order. That made insertion expensive: adding a row meant locating its slot and shifting everything in between, and adding a column meant relocating cells in every row. Patel over-allocated five spare rows and columns at a time to soften reallocation. All tests passed, including round-trips through the JSON conversion, and at a million points with thirty columns the process never exhausted memory — which the object version always did at that size. Loading, however, came out eight times slower than JSON.parse.

The diagnosis: the typed arrays were not the problem, the shifting was. The workload loads rows from the database, transforms them and sends them on; it almost never inserts into the middle of a table. Two secondary costs were also addressed or discarded — resolving column names went from a linear scan on every cell write to a Map over the roughly thirty axis entries, and a binary search over sorted timestamps turned out never to matter.

Dropping the order requirement

The rewrite stops keeping rows sorted entirely. The table instead maintains free lists of unused row and column indexes, plus two Maps linking timestamps to rows and gas names to columns. Adding a timestamp is either a Map hit or a pop from the free list; removal zeroes the row and returns its index to the list. Because queries know their result size before reading, the table is sized correctly up front and the grow-and-copy branch almost never executes. Ordering is restored only where it is read: the times method collects live timestamps and sorts them once per query rather than shifting on every insert.

As reported in the post, JSON objects ran out of memory on the 2 GB droplet, version one never did but was 8x slower than JSON.parse, and version two kept memory flat and beat JSON.parse outright.

Why it matters

This is a concrete demonstration that at large row counts the overhead of a million small objects — hash lookups, reference chasing, garbage collection pressure — can dwarf the size of the data itself, and that contiguous typed-array storage combined with string interning changes the arithmetic entirely. It is equally a lesson in matching a data structure to the access pattern: the first version paid for ordered insertion on every load to serve a read-time concern, and removing that single assumption is what made the fast version fast. For any Node.js service that builds large tabular result sets, the pattern — columnar buffers, sentinels for missing cells, one wrapper class at the boundary — is directly reusable, and the post is accompanied by a four-part video series covering the build.

  • #node-js
  • #typed-arrays
  • #memory-management
  • #performance
  • #javascript

Related posts