deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Trail of Bits open-sources SequenceHash for safe multihashing with any hash function

Trail of Bits has released SequenceHash and SequenceMAC, open-source constructions that make hashing multiple values together unambiguous with SHA2, BLAKE, RIPEMD and other hashes, where NIST's TupleHash only works with Keccak.

Trail of Bits open-sources SequenceHash for safe multihashing with any hash function

What Trail of Bits released

Trail of Bits has published SequenceHash and its keyed counterpart SequenceMAC, a pair of open-source hash constructions that make multihashing — hashing several values together — safe to perform with almost any modern hash function. According to the company's blog post, the specification has been added to the Community Cryptography Specification Project (C2SP), and initial implementations in Rust, Go and Python are available alongside a large set of test vectors covering multiple hash functions, including intermediate values to help developers debug their own ports.

The problem with hashing several things at once

Multihashing appears all over cryptography, whenever a paper or protocol instructs you to compute something like N = Hash(X, Y, Z). The intuitive approach — feeding each value into the hash function in separate calls — is insecure, because the hash function sees nothing but the concatenation of the inputs. Trail of Bits demonstrates this with SHA256: hashing "Test 0", "Test 1" and "Test 2" in three separate updates yields exactly the same digest as hashing "Test 0Test 1" followed by "Test 2", or "Test 0", an empty string and "Test 1Test 2".

That ambiguity has real consequences. Multihashing is a critical component of the Fiat-Shamir transform, a core technique in zero-knowledge proofs, and errors there can introduce proof forgeries — mistakes that, in the cryptocurrency world, have sometimes been measured in millions of dollars, according to Trail of Bits. Other common uses include authenticating the files in an archive, grouping several transactions under a single hash, and commitment schemes where a secret is hashed together with a random blinding value; if the boundary between the two is unclear, a commitment can potentially be opened in more than one way.

Despite this, no consistent standard has emerged. In its audits and across open source, Trail of Bits reports finding ad hoc fixes such as separator characters that can also occur inside the inputs, encodings like Base64 strings joined with dollar signs that add complexity and subtle timing risks, and protocols that length-prefix some inputs but not others.

Why TupleHash wasn't enough

The best-known existing answer is TupleHash, defined in NIST SP 800-185. Trail of Bits calls it a genuinely good tool: it uses straightforward length-prefix encoding, handles effectively unlimited input sizes and naturally works as an extendable output function. The catch is that TupleHash is defined only for Keccak, the permutation behind SHA3. Substituting a different hash function can strip away important properties such as resistance to length-extension attacks. With SHA3 adoption described as lackluster over the past decade — and CNSA 2.0 mandating SHA384 and SHA512 for nearly everything in the US government contracting sector — many developers have been left without a standard option. Existing tools such as Merlin and Trail of Bits' own decree target Fiat-Shamir transforms specifically and do not generalize to everyday multihashing.

What SequenceHash offers

SequenceHash is hash-agnostic in roughly the way HMAC is: it works out of the box with the SHA2 family, BLAKE, RIPEMD and other secure hash functions, and it avoids requiring developers to implement fiddly computations that are not byte-aligned. Trail of Bits says the construction provides four features that keep hashed values from being mixed, extended or replayed across contexts, and details two in the announcement: unambiguous input encoding, meaning no other sequence of inputs of any length can produce the same input to the underlying hash, in line with the Horton principle of hashing what you mean; and length-extension prevention, closing the classic attack where someone who knows H(A) can compute H(A || B) without learning A.

The construction's security rests entirely on the underlying hash — it cannot rehabilitate broken functions such as MD4 or SHA0, and the specification assumes a reasonable choice like SHA256. SequenceMAC, the keyed variant, supports keys of 32 bytes or longer, up to a theoretical ceiling of 2^128 − 1 bytes.

Why it matters

Ambiguous input encoding is a recurring source of cryptographic failures, and until now developers outside the Keccak ecosystem had no standardized, well-tested way to avoid it. SequenceHash gives the majority of systems still running SHA2, BLAKE or RIPEMD a drop-in construction with a public specification, cross-language implementations and verification vectors. That lowers the barrier for the protocol designers who need it most — including zero-knowledge proof systems and government-aligned work constrained to SHA384 and SHA512 — and replaces the current patchwork of ad hoc encodings with something that can be reviewed, tested and reused.

  • #cryptography
  • #hashing
  • #open-source
  • #security
  • #zero-knowledge-proofs

Related posts