deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

DuckDB's DuckLake extension reads and writes an open lakehouse built on SQL and Parquet

A new DuckDB extension attaches DuckLake, an open lakehouse format that keeps metadata in a SQL catalog database and data in Parquet files, adding updates, time travel and change feeds.

DuckDB's DuckLake extension reads and writes an open lakehouse built on SQL and Parquet

DuckDB gains a native lakehouse format

The DuckDB project has published DuckLake, an extension that lets the database read and write an open lakehouse format directly. The repository, duckdb/ducklake, surfaced on the Hacker News front page. According to the project's documentation, DuckLake is built from two familiar layers: metadata lives in a catalog database, while table data is stored as Parquet files.

After installing the extension with a single INSTALL command — or FORCE INSTALL ducklake FROM core_nightly for the latest development build — a lake can be attached and used:

ATTACH 'ducklake:metadata.ducklake' AS my_ducklake (DATA_PATH 'file_path/');

From that point on, tables are created, modified and queried with ordinary SQL. The README walks through CREATE TABLE, INSERT, UPDATE and plain SELECT against the attached database. Notably, UPDATE works at the row level, rewriting a single value in the example — something raw Parquet files do not provide on their own.

Where the metadata goes

Splitting the catalog from the data is the format's core design decision. The usage example keeps the catalog in a DuckDB database file, metadata.ducklake, and points DATA_PATH at a directory of Parquet files, so the two layers can sit in different places. The test configurations confirm the catalog is not tied to DuckDB itself: the suite can run with PostgreSQL or SQLite acting as the catalog database, and there is a separate configuration for running with deletion vectors enabled.

The test tree also signals the feature coverage the team considers important — transaction conflict handling, partitioning, and a mode that runs DuckDB's own core test suite with DuckLake attached as the storage backend.

Snapshots, time travel and change feeds

Because changes are recorded in the catalog, DuckLake exposes version-oriented querying:

  • Time travel: querying FROM my_ducklake.my_table AT (VERSION => 2) returns the table as it looked at an earlier snapshot. In the README's walkthrough, this brings back the pre-update value even after an UPDATE has replaced it.
  • Schema evolution: ALTER TABLE my_ducklake.my_table ADD COLUMN new_column VARCHAR; goes through, and existing rows come back with NULL in the new column.
  • Change Data Feed: the table_changes function, called as table_changes('my_table', 2, 2), produces one row per change with snapshot_id, rowid and change_type alongside the column values, so consumers can replay inserts and updates between snapshots.

Building and contributing

Anyone can compile the extension from source by initialising and updating the git submodules and running make. The submodules are pinned to the DuckDB version recorded in .github/duckdb-version, which is the version CI builds against; a make pull target moves them to the tips of their branches instead, where the build can fail. External contributions are welcome and should target the main branch. Tests run through ./build/release/test/unittest, with options to filter to a single file or pattern, to swap in PostgreSQL, SQLite or deletion-vector configurations, and to exercise DuckDB's core tests on top of DuckLake storage.

Why it matters

DuckDB is an embedded analytical engine, and lakehouse formats have mostly been the territory of larger distributed systems. With DuckLake, a single-process database can attach a lake, run standard SQL against it, mutate rows, evolve schemas and read historical snapshots — without a separate compute cluster or a bespoke metadata service in the critical path.

The choice of building blocks matters too. By keeping the catalog in an ordinary SQL database and the data in Parquet, the format leans on technologies that other engines can already speak, and the fact that the test suite runs against PostgreSQL and SQLite catalogs suggests the design is deliberately engine-agnostic rather than a DuckDB-only convenience. Running DuckDB's own core test suite with DuckLake as the storage backend is perhaps the strongest signal of intent: this is positioned as first-class storage for DuckDB, not just a reader bolted onto Parquet files.

  • #duckdb
  • #parquet
  • #lakehouse
  • #open-source
  • #data-engineering

Related posts