deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

NVIDIA's CUDA Rust compiles GPU kernels natively to PTX with compile-time memory safety

NVIDIA's CUDA Rust lets developers write GPU kernels in Rust compiled natively to PTX, via two tracks: cuda-oxide for SIMT-level control and cutile-rs for tile-based programming, per a dev.to write-up.

NVIDIA's CUDA Rust compiles GPU kernels natively to PTX with compile-time memory safety

According to a dev.to article by Ethan Vance, NVIDIA has begun letting developers write GPU kernels directly in Rust and compile them natively to PTX, the virtual instruction set CUDA drivers execute on the hardware. The capability, described as CUDA Rust and announced in September 2026, closes a gap that previously forced Rust-based GPU programs to keep their kernel code in C++ or another language: you could launch kernels from Rust, but the kernels themselves had to be written elsewhere.

Two programming tracks

The article lays out two routes that mirror the execution models already present in CUDA, and the choice comes down to how much manual hardware control a workload needs.

  • cuda-oxide (SIMT track): a custom rustc codegen backend that intercepts compilation and routes kernel functions through Rust's MIR, the Pliron IR framework and LLVM IR before translating them into PTX. The developer manages thread indexing and memory allocation directly. To reconcile Rust's rule that only one mutable reference may exist at a time with thousands of parallel threads, it introduces a DisjointSlice type that splits a single mutable borrow into per-thread pieces. It requires Linux, a GPU with Compute Capability 8.0 or higher, the CUDA 12.x toolkit and a pinned nightly Rust toolchain.
  • cutile-rs (Tile track): a higher-level abstraction where computation operates on tiles — sub-tensors of data — rather than individual scalars. A #[cutile::module] macro embeds the kernel's AST into the host binary, which is JIT-compiled through CUDA Tile IR at launch time. The Tile IR compiler, not the developer, decides how tiles map onto the physical GPU, which the article says keeps source code hardware-agnostic. Instead of DisjointSlice, the host side calls a .partition() method that assigns each tile block exclusive ownership of its data chunk. It runs on stable Rust 1.89 or newer and needs CUDA 13.3 plus Compute Capability 8.0+, with no nightly toolchain or custom LLVM setup.

The borrow checker arrives on the GPU

The central appeal, per the article, is that classic GPU concurrency bugs are caught before the code ever runs. In C++ CUDA, thousands of threads can touch the same buffers in no guaranteed order, and aliasing or data-race mistakes often pass unit tests only to fail in production. In CUDA Rust, both tracks enforce Rust's borrowing rules at compile time: inputs are shared references, while an output buffer must be exclusively owned by a single writer. If a developer passes an output buffer as one of its own inputs, the build stops with a borrowing error such as E0502 instead of producing a runtime fault on the device.

Because kernels compile straight to PTX rather than passing through Python wrappers or abstraction layers, the article also claims zero runtime overhead and full use of the GPU's compute cycles for the workload itself.

Which track to pick

NVIDIA's guidance, as relayed by the article, is to start new projects on the Tile track and drop down to SIMT only when a workload demands manual thread indexing, precise control over the execution grid or custom shared-memory management. The company also plans inter-language interoperability between the tracks, so the initial choice is not a lock-in.

Early tooling and infrastructure demands

The article notes the technology is still in early alpha and that its compilation pipeline is heavy, arguing that compiling and benchmarking CUDA Rust workloads calls for substantial CPU resources, stable OS-level configurations and isolated environments. It is worth reading that section in context: the piece doubles as a pitch for bare-metal dedicated GPU servers, so its infrastructure claims carry a commercial framing.

Why it matters

The systems layer of AI — inference engines, serving infrastructure and drivers — has been steadily shifting to Rust, and per the article NVIDIA is part of that shift: the Nova Linux driver is written in Rust and NVTX already ships Rust bindings. The kernel itself was the last component that could not be written in the language. If CUDA Rust matures as described, an entire class of memory and concurrency bugs that currently surface as GPU crashes, silent data corruption or out-of-memory failures in production would instead become compile-time errors, without giving up the bare-metal performance that made C++ the default. The caveat is provenance and maturity: the details above come from a single dev.to write-up of an alpha-stage toolchain, and requirements such as CUDA 13.3 and a pinned nightly compiler suggest the interface will keep moving.

  • #rust
  • #cuda
  • #nvidia
  • #gpu
  • #memory-safety

Related posts