· via Hacker News – Front Page (native)
Rgpu adds a PyTorch device that keeps tensors on a remote NVIDIA GPU
An open-source project called Rgpu lets PyTorch programs treat a remote NVIDIA GPU as a local device, sending tensor operations over TCP while the Python application stays on the client.
What Rgpu does
A project called Rgpu, published on GitHub under the handle ymcrcat and surfaced on Hacker News' front page on October 8, 2026, addresses a familiar constraint: you need an NVIDIA GPU for a workload, but the machine you work on does not have one. According to the repository's README, Rgpu runs GPU work on a remote NVIDIA machine while the application stays on the client — including, as the project's own summary puts it, from a Mac with no CUDA installation.
The centerpiece is a custom PyTorch device. A program opts in by requesting "rgpu" as its device, after which tensor allocations and operations travel over TCP to a server process on the GPU machine. The README's smoke test creates a small tensor of ones on the rgpu device, doubles it, sums it and prints 8.0 — the remote hop is otherwise invisible to the script.
Two integration paths
Per the README, Rgpu currently offers two routes. The PyTorch device is aimed at PyTorch programs that can be adapted to select the new device, and carries torch operations over TCP. The second path is a CUDA shim for existing Linux CUDA programs, including a stock CUDA build of PyTorch; it intercepts libcuda, the CUDA Runtime, cuBLAS, cuBLASLt and cuDNN, and forwards the calls to the remote GPU.
The README is upfront about the trade-off between them: the PyTorch device is the simpler integration, while the CUDA shim can run unmodified binaries but has to emulate far more of the CUDA stack.
Setup for the PyTorch path is a pip install, a server deployment, and a bundled launcher (rgpu-run) that opens an SSH tunnel to the GPU host and wires up the connection before running your script. The repository also ships a documentation site with a quickstart, a training guide, a nanoGPT example, the CUDA shim guide, an operations reference, configuration options, performance notes and troubleshooting.
Read the security notes before deploying
The README contains a warning worth highlighting: neither protocol authenticates or encrypts connections. The recommended posture is to keep the operation server bound to localhost and reach it through SSH. The CUDA server is more exposed out of the box, listening on all IPv4 interfaces, so the project advises restricting port 9713 with firewall rules before starting it — even when it sits behind an SSH tunnel. In short, this is tooling designed for trusted, tunnelled setups rather than open network exposure.
Implementation details
The codebase spans C++ and Python: the client shims and transport, the server with its dispatch logic, and a shared protocol definition live on the C++ side, with generated C++ committed to the tree and regenerated from a CUDA header parser. The Python package holds the PyTorch device and the launcher. Around the code sit engineering records and experiment reports, plus an experimental JAX directory that the README explicitly labels as unsupported. The project is licensed under Apache 2.0.
Why it matters
Developers without local CUDA hardware typically choose between a full remote development environment, awkward sync-and-run loops, or managed notebooks, and each of those pulls the workflow away from the machine where the code is actually written. A device-level abstraction collapses that to roughly a one-line change, so training scripts keep their shape and tooling while the compute lands elsewhere — useful for laptop-based developers, and equally handy for keeping a powerful desktop GPU busy from anywhere. The CUDA shim extends the same idea to binaries that were never designed for remote execution.
The candid security caveats set expectations: this fits personal or trusted-team infrastructure, not multi-tenant service. And while the smoke test is trivial, the presence of a nanoGPT example and a dedicated performance section signals that the author expects real training workloads — how well the TCP transport holds up under them is the question to watch as the project matures.
- #pytorch
- #gpu
- #open-source
- #remote-compute
- #cuda