· via Hacker News – Front Page (native)
Yantra brings LALR(1) parsing to C++ with built-in lexer and top-down AST walking
A new open-source LALR(1) parser generator written in C++, Yantra combines lexer, parser and AST generation in one grammar and runs semantic actions top-down after the full tree is built.
What Yantra is
Yantra, a new MIT-licensed parser generator written in C++, has been shared on Hacker News. It is a compiler-compiler in the LALR(1) family that bundles the three pieces a language implementer usually assembles separately: an integrated lexer, the parser itself, and AST generation with built-in tree walkers. The project is authored by Renji Panicker.
According to the project's repository, Yantra depends on nothing beyond the C++ standard library. A plain CMake build produces the generator executable, ycc, and generated parsers compile with any C++23 compiler, including clang, gcc and MSVC. The name is Sanskrit for "machine", a nod to the state machines at the heart of the technique.
One grammar, lexer included
The lexer supports UNICODE and UTF-8 input and is multi-mode, which the project notes is useful for constructs such as nested multi-line comments. Parsing can also run in a push-based, lexer-driven mode: input is read one character at a time and tokens are handed to the parser as they complete, which suits processing a stream as it arrives, for example from a socket.
Output comes in two shapes. In an optional amalgamated mode, the entire parser is generated as a single .cpp file complete with a working main() function; in non-amalgamated mode, it emits separate .hpp and .cpp files meant to drop into an existing project.
Parse bottom-up, walk top-down
The design decision that separates Yantra from the classic tools is when semantic actions run. Bison, Yacc and Lemon, with Lemon from SQLite being Yantra's stated inspiration, execute actions during the parse, bottom-up, as each rule is reduced. Yantra instead always builds the complete AST first, then walks it top-down in a separate pass calling your actions, so a parent rule's action runs before its children are visited.
The repository illustrates this with a calculator grammar: for "1 + 2 + 3", parsed left-associatively as (1 + 2) + 3, the root node's action fires first, then its left child, then the right subtree. A hand-written recursive-descent or bottom-up parser would need extra AST classes plus a separate traversal pass to get that ordering; here, the project says, it falls out of the grammar directly. A single grammar can also define multiple walkers, for example one emitting C++ and another emitting Java from the same parse. Getting either behaviour out of the Bison family means hand-building the AST and walker on top.
Against ANTLR and tree-sitter
The comparison section of the project's documentation is unusually direct. ANTLR also walks a fully built parse tree, but that comes naturally from its LL(*) algorithm. Yantra delivers the same top-down walk on top of LALR(1), a bottom-up algorithm with no moment during parsing at which the whole tree exists, while, per the project's claim, keeping LALR(1)'s time and space efficiency. It targets C++ only, and its generator is a native executable, whereas ANTLR's generator is a Java program, so a C++ project otherwise needs a JVM in its build toolchain just to run the generator. The documentation is frank that ANTLR is far more mature and widely used, and that Yantra is a much smaller, newer, single-maintainer effort. tree-sitter is set aside as a different problem entirely: incremental, error-tolerant parsing for editors and IDEs, which Yantra does not attempt.
Ecosystem and maturity
Around the core tool there is a standalone sample project, lingo, and a third-party language server by Raj Chaudhuri that adds Yantra syntax highlighting to VS Code, Qt Creator and any other IDE supporting the Language Server Protocol. The repository also links a Known Limitations page covering what the tool does not do yet.
Why it matters
For compiler and DSL work in C++, the incumbent paths each carry friction: the Bison family pushes AST construction and traversal onto you, while ANTLR imports a JVM dependency and a larger runtime. Yantra's pitch is that one grammar yields lexer, parser, AST and top-down walkers, generated by a dependency-free native tool. That is a genuine convenience for small language projects. It is early software from a single maintainer, so anyone betting a production compiler on it should read the limitations list and weigh the bus factor, but the core idea, top-down walking layered on an LALR(1) parse, addresses a real gap for C++ developers.
- #parser-generator
- #c-plus-plus
- #open-source
- #compilers
- #developer-tools