Skip to content
Blog

Candle: Pure-Rust LLM Inference in Production

A tour of Hugging Face's Candle — a minimalist Rust ML framework for serverless and GIL-free inference: crates, backends, supported formats and models, and how it compares to tch-rs and Burn.

Published on • October 3, 2026

AI Assistant

Python owns ML tooling, but it’s increasingly not where inference runs. If you’re shipping a serverless function, an edge binary, or a latency-critical service, dragging in CPython and the GIL to run a transformer is a strange kind of tax. Candle is Hugging Face’s answer: a minimalist ML framework for Rust with first-class GPU support and no Python in the runtime.

Source: github.com/huggingface/candle (README; ~21k stars, MIT / Apache-2.0).

Why it exists

The README’s motivation is blunt: serverless inference wants small, fast binaries — and production services shouldn’t carry the GIL. Candle also aligns the Rust ecosystem with Hugging Face’s own Rust-first infrastructure: safetensors for weights and tokenizers for BPE are already Rust crates, so a pure-Rust stack has no foreign boundary to cross.

PyTorch-like ergonomics are part of the pitch:

use candle_core::{Device, Tensor};

let device = Device::new_cuda(0)?;                  // or Device::Cpu / new_metal(0)
let a = Tensor::randn(0f32, 1., (2, 3), &device)?;
let c = a.matmul(&b)?;

Same mental model as Torch, no Torch in your dependency tree.

The crate layout

Candle is deliberately modular — you depend on what you use:

CrateRole
candle-coreTensors and ops (the foundation)
candle-nnLayers, loss, optimizers
candle-transformersModel implementations
candle-datasetsData loading
candle-kernels / candle-metal-kernelsCUDA / Metal kernels
candle-flash-attnFlash-attention v2
candle-onnxONNX model support
candle-pyo3Python bindings (for tooling, not runtime)
candle-examplesRunnable examples

Backends and formats

  • CPU — optimized, with optional MKL/Accelerate acceleration; capable enough for quantized small models on a laptop.
  • CUDA — multi-GPU via NCCL (--features cuda), cuDNN, and cutile JIT-compiled kernels.
  • Metal — via the candle-metal-kernels crate for Apple silicon.
  • WASM — runs in the browser, which pairs neatly with Flutter’s WebAssembly-first web story.

Weights: safetensors (the native choice), plus npz, ggml/GGUF, and PyTorch checkpoints. Quantization reuses llama.cpp’s quant types, so GGUF models you already have are fair game — a big deal given how much of the local-model world is GGUF.

What runs out of the box

The model zoo covers the classics: LLaMA v1–v3, Mistral 7B, Mixtral 8x7B, Falcon, Phi 1–3, Gemma, StarCoder, Qwen, Whisper, Stable Diffusion, BERT, T5, YOLO and SAM. In practice this means “clone, cargo run, get tokens” for most well-known architectures — and candle-transformers as a reference when you’re implementing a newer one.

Training is supported too, but the honest positioning is inference-first: this is where Candle’s performance story is strongest.

How it compares

vs tch-rs — tch-rs are bindings to libtorch. The README’s own framing: they’re “extremely versatile, but they bring in the entire torch library into the runtime.” Choose tch-rs when you need PyTorch’s exact op coverage and can afford the binary size; choose Candle when you want a lean dependency graph. (Fun fact: the main tch-rs contributor also works on Candle.)

vs Burn — Burn is a multi-backend DL framework with a stronger training story and a different abstraction style (including compile-time shapes in other experiments). Candle leans minimalist and HF-ecosystem-native; Burn leans batteries-included and backend-portable.

vs llama.cpp — llama.cpp is a C++ inference engine; Candle is a framework you build on. If you only need to serve GGUF models, llama.cpp/vLLM-class tools may be less work; if you want inference embedded in a Rust service with your own post-processing pipeline, Candle composes better.

A realistic production checklist

  1. Right-size the model. Candle makes 7B quantized models feasible on one GPU — measure tokens/s on your hardware before promising latency.
  2. Load with safetensors, pin versions. Reproducibility starts with weight format and model revision hashes.
  3. Feature-flag your backend (cuda, metal) so CPU-only CI still builds.
  4. Watch memory, not just speed. KV-cache growth dominates long conversations; batch and truncate deliberately.
  5. Benchmark against a Python baseline once — it’s the only way to know whether “no GIL” actually showed up in your p95.
  6. Keep a fallback. Hybrid deployments (Candle for hot paths, hosted API for the long tail) are the boring, correct answer for most teams.

When it’s the right call

Candle is at its best when your constraint is deployment shape: a single static-ish binary, sub-second cold starts, no Python runtime, GPUs optional. That’s serverless inference, edge appliances, CLIs, in-process RAG, and latency-sensitive Rust services. Reach for PyTorch when you’re training; reach for Candle when you’re serving — and let the safetensors + tokenizers crates carry the weights the rest of the way.