deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

OpenDLSS-NR rebuilds NVIDIA's DLSS 5 neural rendering in Vulkan, bit-exact to the original

An open-source project has reimplemented the DLSS 5 neural rendering network in Vulkan and again in WebGPU, matching NVIDIA's original byte for byte at every block boundary.

OpenDLSS-NR rebuilds NVIDIA's DLSS 5 neural rendering in Vulkan, bit-exact to the original

A byte-for-byte rebuild of DLSS 5's neural renderer

An open-source project called OpenDLSS-NR has reimplemented the neural network behind NVIDIA's DLSS 5 neural rendering in Vulkan, and it claims something rare: bit-exact agreement with the original. According to the project's README, which reached Hacker News's front page on 1 October 2026, the code reproduces the same 71-block Swin/ViT network as DLSS-NR build 310.8.0, running FP8 on tensor cores. The match extends to the internals: all 75 block boundaries line up byte for byte, not just the final image.

The repository holds two independent implementations. The native path pairs a C++20 Vulkan host with GLSL kernels, plus a faster route built from generated PTX. A second implementation under ports/browser-webgpu/ runs the same network in a browser with no tensor cores, no FP8 and no fusion between blocks — and still produces identical bytes against the same captures, at 72 ms per frame at 512x512 versus 2.7 ms natively. The lesson, per the README, is that the agreement comes from precisely specifying the computation, not from the underlying hardware.

What the network actually does

One clarification the README insists on: DLSS-NR is not an upscaler. Input and output resolutions are identical. NVIDIA's term is a generative neural rendering network — it re-renders the frame the engine already drew, inventing detail from injected noise and shifting tone, structure and skin according to a style setting. The architecture is a U-net of shifted-window transformer blocks with a global ViT at its base: 71 blocks across six pooling levels, FP8 (E4M3) activations with FP16 accumulation, and 141 MiB of weights. Per frame it takes a low-dynamic-range proxy of the rendered image, three channels of Gaussian noise, the reprojected output of the previous frame and five conditioning scalars, and emits four f32 channels per pixel: an RGB residual plus one temporal-blend logit. NVIDIA documents the model in its report DLSS 5: Generative Neural Rendering.

Weights are not included

The project ships no weights. Users supply a model directory — a manifest of stages and tensors with E4M3 weights as packed bytes — and the loader refuses anything that is not exactly the 71-block graph of build 310.8.0. Parity is checked against fixtures, recorded captures of NVIDIA's original that are also excluded from the repository. A parity command gates on those fixtures; a verify command bisects block 0 kernel by kernel. The verdicts are strict: output equal only up to the sign of zero counts as a failure, 8-bit captures may pass within one code, and everything else is a mismatch. Fixtures are validated before the GPU runs, and any fixture that declares a check without its reference, or leaves a comparable boundary with neither reference nor explanation, is refused.

Performance and requirements

On an RTX 4070 SUPER, with the whole network per frame and 241 dispatches at every resolution, the project reports minimum timings over 40 frames of 2.8 ms at 768x768, 7.8 ms at 1920x1080, 12.6 ms at 2560x1440 and 29.3 ms at 3840x2160. The README advises comparing minima because the GPU alternates between two clock states under sustained load, which pushes medians a few percent higher. The fast path rests on generated PTX — mma.sync E4M3 with f16 accumulation, cp.async rings, barrier-free chaining through device counters and split-K GEMMs — while the GLSL cooperative-matrix kernels form a complete reference route on their own. Requirements are narrow: Windows, an NVIDIA Ada or newer GPU, a driver exposing VK_KHR_cooperative_matrix, VK_NV_cooperative_matrix2, VK_EXT_shader_float8 and VK_NV_cuda_kernel_launch, and Visual Studio 2022. A demo embeds the network in a patched Filament renderer, adding per-object motion vectors and a Vulkan interop hook, with glTF scene loading.

What is left out

DLSS-SR, the upscaling network, is not implemented; it is a different model. The temporal path exists only in the demo, where the history input lanes and the per-pixel blend logit drive a reprojected feedback loop. The command-line tool runs single frames with no history, which is how the reference captures were made.

Why it matters

Faithful reimplementations of shipping vendor neural networks are uncommon; ones verified bit-exactly at every block boundary, rather than by eyeballing screenshots, are rarer still. OpenDLSS-NR demonstrates that a hardware-coupled inference stack can be rebuilt on open APIs with provable numerical parity, and the WebGPU port pushes the argument further: if the same bytes fall out of a browser with none of the tensor-core machinery, the network's behaviour has effectively become a portable specification that other implementations and vendors could target. The caveats are real — no weights are distributed, the fast path needs recent NVIDIA silicon on Windows, and the upscaler most people associate with the DLSS name is absent. As a proof that DLSS 5-class neural rendering can be specified, ported and audited outside its original stack, though, it is a genuine milestone.

  • #open-source
  • #vulkan
  • #nvidia
  • #graphics
  • #machine-learning
  • #webgpu

Related posts