· via dev.to (home feed)
llamadart 0.11 packs offline model loading, chat history and runtime checks into one Dart API
Dart and Flutter package llamadart 0.11 introduces a single-call model loader, history-aware chat sessions, non-disruptive model switching and runtime capability reporting for offline AI apps.

Loading becomes a single call
According to a post on dev.to, llamadart began as an attempt to give a writing assistant an offline mode — something usable on a flight, on a flaky connection, or in a network where outbound API access is blocked. Version 0.11, the post explains, is less about new capabilities and more about compressing everyday integration work into a smaller, clearer API.
The biggest change is model loading. In 0.10 an app created an engine around a backend and then invoked a separate load method. In 0.11, LlamaEngine.load accepts a LlamaModel whose ModelSource may be a local path, a URL or a Hugging Face reference, and hands back an engine that is ready to use. Ownership is now explicit: if loading throws, the engine and backend it created have already been disposed; on success, the caller owns the finished engine. Download progress and cancellation surface through the same operation via onProgress and download options, so the download UI a user sees before the first generated token shares a lifecycle with loading itself.
The post also notes that the package's native build hook resolves runtime assets, meaning a local C++ toolchain is unnecessary for the common setup. Model files remain an app-level concern.
From prompt to conversation
A follow-up request such as 'make it shorter' only works if the earlier exchange is still in context. ChatSession keeps that history; the new convenience method session.send returns the completed reply, and streaming via session.create now exposes text as chunk.text rather than reaching into a chunk's choices array.
The session also manages context limits: it drops the oldest turns to stay within the context budget while preserving the system prompt. Apps that maintain their own transcripts can bypass the session and hand a full message list straight to the engine.
The post's worked example loads a deliberately tiny model — SmolLM2-135M-Instruct in GGUF form at Q2_K quantization, with a 1024-token context and no GPU layers — and the author is careful to note that choosing a model that actually follows instructions well is a separate evaluation task.
Switching models without going dark
Model pickers raise a lifecycle problem: unloading the current model before downloading a replacement leaves the user with nothing available on a slow connection. On native targets, dev.to explains, setModel keeps the existing model loaded while the replacement's files are fetched; if that fetch fails or is cancelled, the old model stays in place. The boundary is documented honestly — a failure during the later load stage can still end with nothing loaded, and web runtimes unload before fetching.
The post also offers placement advice: a Flutter engine belongs in a service, provider or state object rather than inside build(), disposed when its owner finishes.
More runtimes, explicit differences
llamadart has outgrown its original llama.cpp path. LiteRT-LM support arrived in 0.7; 0.8 split Apple runtime companions out of the core package so pure Dart users are not forced onto a Flutter SDK; 0.10 added preview image generation through an opt-in stable-diffusion.cpp runtime.
To let apps reason about those differences, 0.11 introduces engine.runtime to identify the active runtime and await engine.capabilities to report which operations it supports — for instance, whether an image-attachment button makes sense. ModelParams.device provides shared CPU, GPU or NPU selection: an explicit request either runs there or fails as unsupported, while auto defers to the runtime's default.
The limits are stated rather than hidden: Android's llama.cpp Vulkan path is experimental and device-dependent, web support is experimental, and automatic tool loops refuse the pinned LiteRT-LM runtimes because they cannot reliably report token-limit truncation. Running it requires Dart 3.10.7 or newer, added via llamadart: ^0.11.0.
Why it matters
Much on-device AI coverage centres on which model fits in memory, but the dev.to post is a useful reminder that shipping offline AI is mostly application engineering: downloads that can be cancelled, engines that own resources deterministically, chat state that fits a context window, and capability checks across runtimes that genuinely behave differently. By making failure modes explicit — a load that disposes itself, a model switch that keeps the old model warm, a device request that fails loudly instead of silently falling back — llamadart 0.11 offers a pattern any developer wiring local inference into an app can borrow, whatever their stack.
- #dart
- #flutter
- #local-ai
- #on-device-inference
- #llama-cpp