Candle: Pure-Rust LLM Inference in Production
A tour of Hugging Face's Candle — a minimalist Rust ML framework for serverless and GIL-free inference: crates, backends, supported formats and models, and how it compares to tch-rs and Burn.
Published on • October 3, 2026
AI Assistant

Python owns ML tooling, but it’s increasingly not where inference runs. If you’re shipping a serverless function, an edge binary, or a latency-critical service, dragging in CPython and the GIL to run a transformer is a strange kind of tax. Candle is Hugging Face’s answer: a minimalist ML framework for Rust with first-class GPU support and no Python in the runtime.
Source: github.com/huggingface/candle (README; ~21k stars, MIT / Apache-2.0).
Why it exists
The README’s motivation is blunt: serverless inference wants small, fast binaries — and production services shouldn’t carry the GIL. Candle also aligns the Rust ecosystem with Hugging Face’s own Rust-first infrastructure: safetensors for weights and tokenizers for BPE are already Rust crates, so a pure-Rust stack has no foreign boundary to cross.
PyTorch-like ergonomics are part of the pitch:
use candle_core::{Device, Tensor};
let device = Device::new_cuda(0)?; // or Device::Cpu / new_metal(0)
let a = Tensor::randn(0f32, 1., (2, 3), &device)?;
let c = a.matmul(&b)?;
Same mental model as Torch, no Torch in your dependency tree.
The crate layout
Candle is deliberately modular — you depend on what you use:
| Crate | Role |
|---|---|
candle-core | Tensors and ops (the foundation) |
candle-nn | Layers, loss, optimizers |
candle-transformers | Model implementations |
candle-datasets | Data loading |
candle-kernels / candle-metal-kernels | CUDA / Metal kernels |
candle-flash-attn | Flash-attention v2 |
candle-onnx | ONNX model support |
candle-pyo3 | Python bindings (for tooling, not runtime) |
candle-examples | Runnable examples |
Backends and formats
- CPU — optimized, with optional MKL/Accelerate acceleration; capable enough for quantized small models on a laptop.
- CUDA — multi-GPU via NCCL (
--features cuda), cuDNN, andcutileJIT-compiled kernels. - Metal — via the
candle-metal-kernelscrate for Apple silicon. - WASM — runs in the browser, which pairs neatly with Flutter’s WebAssembly-first web story.
Weights: safetensors (the native choice), plus npz, ggml/GGUF, and PyTorch checkpoints. Quantization reuses llama.cpp’s quant types, so GGUF models you already have are fair game — a big deal given how much of the local-model world is GGUF.
What runs out of the box
The model zoo covers the classics: LLaMA v1–v3, Mistral 7B, Mixtral 8x7B, Falcon, Phi 1–3, Gemma, StarCoder, Qwen, Whisper, Stable Diffusion, BERT, T5, YOLO and SAM. In practice this means “clone, cargo run, get tokens” for most well-known architectures — and candle-transformers as a reference when you’re implementing a newer one.
Training is supported too, but the honest positioning is inference-first: this is where Candle’s performance story is strongest.
How it compares
vs tch-rs — tch-rs are bindings to libtorch. The README’s own framing: they’re “extremely versatile, but they bring in the entire torch library into the runtime.” Choose tch-rs when you need PyTorch’s exact op coverage and can afford the binary size; choose Candle when you want a lean dependency graph. (Fun fact: the main tch-rs contributor also works on Candle.)
vs Burn — Burn is a multi-backend DL framework with a stronger training story and a different abstraction style (including compile-time shapes in other experiments). Candle leans minimalist and HF-ecosystem-native; Burn leans batteries-included and backend-portable.
vs llama.cpp — llama.cpp is a C++ inference engine; Candle is a framework you build on. If you only need to serve GGUF models, llama.cpp/vLLM-class tools may be less work; if you want inference embedded in a Rust service with your own post-processing pipeline, Candle composes better.
A realistic production checklist
- Right-size the model. Candle makes 7B quantized models feasible on one GPU — measure tokens/s on your hardware before promising latency.
- Load with
safetensors, pin versions. Reproducibility starts with weight format and model revision hashes. - Feature-flag your backend (
cuda,metal) so CPU-only CI still builds. - Watch memory, not just speed. KV-cache growth dominates long conversations; batch and truncate deliberately.
- Benchmark against a Python baseline once — it’s the only way to know whether “no GIL” actually showed up in your p95.
- Keep a fallback. Hybrid deployments (Candle for hot paths, hosted API for the long tail) are the boring, correct answer for most teams.
When it’s the right call
Candle is at its best when your constraint is deployment shape: a single static-ish binary, sub-second cold starts, no Python runtime, GPUs optional. That’s serverless inference, edge appliances, CLIs, in-process RAG, and latency-sensitive Rust services. Reach for PyTorch when you’re training; reach for Candle when you’re serving — and let the safetensors + tokenizers crates carry the weights the rest of the way.