deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Jane Street intern project tests autoregressive diffusion on order-book data

A Jane Street intern project found autoregressive diffusion can model order-book events, but only after flow matching replaced DDPM and a hand-built categorical head absorbed the data's sharp discontinuities.

Jane Street intern project tests autoregressive diffusion on order-book data

A generative model for order books

Jane Street has published a research write-up asking whether autoregressive diffusion, a technique with roots in image generation, can be used to synthesize order-book market data rather than merely predict prices. The post, part of the firm's 2026 summer intern project series and surfaced on the Hacker News front page, describes work by an intern named Kavish who built an event-level generative model trained on four years of US equities data.

The motivation is a gap in current practice. Quantitative finance already has many models that consume a stream of book events for a symbol — resting orders added, cancellations, executions — and output a future price. A generative model would go further, producing the events themselves, including when they arrive at the exchange. Rollouts from such a model would carry far more texture than a single price trajectory.

That framing raises a question the post keeps returning to: is market data more like video or text? According to the post, it is genuinely both. The order book evolves through a discrete action space, yet key parameters of an order, such as price, have cardinality so high that they behave continuously, and arrival timing is continuous outright. The distributions are also spiky: improving a price by one tick, known as pennying, is far more common than improving by two, and events cluster at zero elapsed time, at the earliest possible reaction time, and around whole-number seconds.

Encoder, categorical head, diffusion head

Following the paper Autoregressive Image Generation without Vector Quantization, the project uses an encoder–diffuser design. A causally masked transformer encoder produces a latent embedding at each position. A small categorical head uses that latent to predict whether the next event is a trade or a best-bid-and-offer update, while a diffusion head generates the continuous targets, such as price and elapsed time, conditioned on the latent and the event kind via AdaLN conditioning. At inference time, the last event's latent drives sampling of the next event, which is then appended to the stream so generation continues autoregressively.

DDPM diverged, flow matching did not

The most concrete technical finding concerns the choice of diffusion objective. The initial head predicted noise on a 1,000-step cosine schedule, as recommended in Improved DDPM, with sampling done via DDIM. According to the post, this configuration was unstable: denoising trajectories exploded, with 88 to 95 percent of values landing more than eight standard deviations from the distribution mean across all seven continuous targets.

More sampling steps shrink the per-step increment and adding stochasticity to sampling regularizes the denoising process, so the instability can be tuned around. But no workaround was needed in the end: a rectified-flow approach, where the model interpolates linearly between noise and data and learns to predict velocity, performed well out of the box, and the project switched to flow matching. The post notes the two methods are different parameterizations of the same objective, yet the difference mattered a great deal in practice.

Hybrid data needs hybrid heads

The deeper problem is that market data has continuous snapshots but evolves in a discrete way governed by microstructure. Prices cluster at the midpoint of the bid and ask, or at the current bid or ask, and timing carries point masses, so a purely continuous diffusion head fits the data poorly. There was also a class imbalance — only 8 percent of events are trades, the rest being BBO changes — that made the two-way categorical head overpredict trades.

The remedy was to fracture the categorical head into 20 classes, such as size-only changes with no price move, ask up, and bid down. This lets sharp atoms, like an elapsed time of exactly zero, be predicted as their own categories, leaving the diffusion head to model smoother residuals such as the magnitude of a nonzero time gap or the magnitude of a bid-price change. The post reports materially better marginal distributions for both event kind and continuous targets, but acknowledges the approach involves substantial hand-engineering that becomes unscalable with larger feature sets.

Why it matters

A working generative model of order-book events would be a practical instrument, not just a research curiosity: simulated markets with realistic microstructure could be used to train and stress-test trading strategies, and the post suggests such rollouts would be far richer than price forecasts alone. The project's partial results also map where off-the-shelf diffusion machinery breaks on hybrid discrete-continuous data — unstable noise-prediction sampling, point masses in the target distributions, and categorical imbalance — and point toward what a working design likely requires: flow-matching-style objectives plus explicit handling of discontinuities. Those lessons generalize beyond finance to any sequential domain that mixes discrete events with continuous values, and the post signals that quant firms are now applying frontier generative modeling directly to market microstructure.

  • #machine-learning
  • #diffusion-models
  • #quant-finance
  • #market-data
  • #generative-models

Related posts