· via Hacker News – Front Page (native)
Janus runs local GGUF models from a single Go binary with an OpenAI-compatible API
Janus is a single Go binary that runs GGUF models locally through Vulkan or on CPU and exposes an OpenAI-compatible API with a web UI, without Python, Docker or Ollama.
Janus, a local LLM server written in Go, has reached Hacker News's front page through a Show HN post. Its pitch is packaging: one compiled binary loads GGUF model files onto your GPU or CPU, serves them through an OpenAI-compatible API, and ships with a web UI. According to the project's GitHub README, the usual local-AI prerequisites — a Python runtime, a container setup, or an Ollama install — are unnecessary, though Ollama can serve as a backend if you already run it. The stated design split is that the model decides what to do while Go handles inference, request routing and tool execution, keeping everything on the machine.
How it runs models
Inference goes through llama.cpp built against Vulkan, which the README says covers AMD, Intel and NVIDIA GPUs, with a CPU fallback when no suitable GPU is present. Configuration lives in a .env file: you point it at a .gguf model, pick a backend (vulkan, cpu or ollama), and can tune GPU layer offloading — the default of -1 places every layer on the GPU — plus a VRAM budget hint defaulted to 9216 MiB and a 4096-token reply cap. Models typically take 2–8 GB of disk, with roughly 50 MB more for the binary and its llama library. A bundled downloader utility fetches GGUF files from Hugging Face; the README's example grabs a Q8_0 quant of Llama-3.2-3B-Instruct.
The API and the tooling around it
The OpenAI-compatible surface includes chat completions with streaming and a model list, so any client that speaks the OpenAI protocol — the README names Cursor and Cline as examples — can point at the local server, which listens on port 8990 by default, and treat local models like any other endpoint. Beyond chat, the server exposes endpoints for listing and invoking tools, running what the README calls a ReAct kernel on a task, and uploading files up to 50 MB.
The built-in tool set is broad for a model server: reading and writing files, executing shell commands, math, Word document import and export, PDF output, and OCR through Tesseract. The web UI adds Assistant and Chat modes, a Kernel tab that streams tool calls into a live log, a Memory tab for storing facts the model can reuse, and a Skills tab for installing community tools from JSON. Models can be hot-swapped from the UI without a restart, and two safety levers exist: a safe mode that blocks shell commands inside tools, and optional Basic Auth on admin endpoints.
Platforms and rough edges
Windows 10/11 is the primary platform; Linux and macOS are also documented, with macOS limited to the CPU backend because Vulkan support varies by hardware. Building requires Go 1.22+, and the Windows build script downloads prebuilt llama.cpp Vulkan libraries before compiling the executable. If Ollama is set as the backend, Janus proxies chat requests to it while its tools and web UI keep working.
The README also carries a frank pitfalls table built from repeated experience, according to its own framing: stale processes holding port 8990, .env files picked up from the wrong directory, model paths that break when launching from another folder, and first replies that take 10–60 seconds while weights load into memory.
Why it matters
Most local LLM stacks still ask users to assemble parts: a Python environment or container image, a model server, and a client to tie them together. Collapsing that into one binary plus a model file removes the biggest source of setup friction, and the repo's documented air-gapped install path signals intent for locked-down environments. The OpenAI-compatible API is the practical glue — because it is the de facto standard, Janus slots into existing tooling without client-side changes. Choosing Vulkan over CUDA also makes AMD and Intel GPUs first-class citizens rather than an afterthought. The space is crowded, with Ollama, llama.cpp's own server and LM Studio covering similar ground, but Janus's bet is architectural: it moves routing, tools, memory and an agent loop into the server itself, so any client, from a curl script to an editor plugin, inherits those capabilities.
- #local-llm
- #gguf
- #vulkan
- #go
- #openai-api