Skip to content
Blog

Burn: A Rust-Native Deep Learning Framework Comes of Age

Burn 0.22 drops backend type parameters for runtime Device selection, unifies CUDA, ROCm, Metal, Vulkan and WebGPU behind CubeCL, and ships a batteries-included training story.

Published on • October 9, 2026

AI Assistant

Python owns deep learning tooling, but the production inference and training niche keeps moving to Rust. Burn is the most ambitious of the Rust-native frameworks: a pure-Rust deep learning library whose single codebase targets CUDA, ROCm, Metal, Vulkan, WebGPU, CPU, WebAssembly, and even no_std.

Version 0.22.0, released October 6, 2026, is the release where Burn’s ergonomics caught up with its ambitions.

What Burn is

Burn is organized around three pillars:

  1. A tensor abstraction with autodiff built in
  2. A backend system — you swap execution engines without rewriting models
  3. Training tooling — burn-train provides a full Learner loop with metrics, checkpointing, and early stopping

The headline of 0.22: backend type parameters are gone from the user API. Execution is now selected at runtime through a Device value, and Tracel claims rebuilds are up to 15× faster as a result.

use burn::tensor::Device;

let device = Device::cuda(0);                    // or Device::wgpu(Default::default())
let device = Device::flex().autodiff();          // pure-Rust backend + autodiff

let model = Model::new(&device);
let input = Tensor::<2>::ones([8, 4], &device);
let gradients = model.forward(input).sum().backward();

Autodiff is now runtime state (device.autodiff()), and training/dataset APIs became fallible — device failures return Err instead of panicking.

The backend matrix after 0.22

BackendEngineNotes
CubeCL (CUDA, ROCm, Metal, Vulkan, WebGPU, CPU)JIT via LLVMFusion, autodiff, remote execution decorators
burn-flexPure Rust eagerstd, no_std, Wasm
burn-tch (LibTorch)Deprecated”will be removed”
NdArrayDeprecatedlegacy
Candle backendRemovedfeature flags dropped

Important consequence: burn ships no default execution backend. You must enable a Cargo feature per backend constructor, and Device::default() picks from compiled-in backends rather than detected hardware — enabling another feature can silently change what default() resolves to.

Only burn-flex works under no_std.

Importing models

Burn’s interop story is ONNX and safetensors, not GGUF:

  • burn-onnx runs at build time via build.rs, generating Rust code plus burnpack weights from an ONNX graph. Opsets 1–24, graph simplification on by default:
use burn_onnx::ModelGen;

fn main() {
    ModelGen::new()
        .input("src/model/my_model.onnx")
        .out_dir("model/")
        .run_from_script();
}
  • SafeTensors support landed in 2025.
  • GGUF has no first-party support — it remains an open design discussion (issue #1187). GGUF carries metadata plus weights, so it maps to the weights-loading path rather than the graph-import path, and model structure must still be defined manually.

Training with burn-train

The Learner bundles model, optimizer, and LR scheduler, driven by SupervisedTraining over train/validation dataloaders:

let optim = AdamWConfig::new().with_weight_decay(5e-5).init();
let result = training.launch(Learner::new(model, optim, lr_scheduler));

Included: loss, accuracy, precision/recall, F-scores, AUROC, BLEU, CER, WER, perplexity, and system-usage metrics — plus PSNR/SSIM/LPIPS/FID behind the vision feature. Checkpointing supports periodic and metric-based saves (burnpack format, keep-last-two plus best-val), with early stopping, multi-device/DDP training, and a tui terminal dashboard by default.

Ecosystem

burn-train, burn-onnx, burn-store, burn-dataset, burn-bench, burn-rl, burn-vision, burn-fusion, burn-autodiff, and the CubeCL crates (burn-cubecl, burn-linalg) all version together at 0.22.

Burn vs. Candle

The two Rust ML stacks have clearly diverged (and the Candle post on this blog covers the other side):

  • Candle — minimalist, Hugging Face ecosystem native, GIL-free serverless inference, leans inference-only.
  • Burn — batteries-included, backend-portable, with a real training story and compile-time kernel fusion.

If you want training loops, multi-GPU, and one model definition targeting six hardware stacks, Burn is the answer. If you want a minimal HF-compatible inference crate, Candle still wins.

Gotchas

  • 0.22 is a breaking migration: Tensor<B, D> → Tensor<D>, server → remote-server features, and more. Budget migration time.
  • No default backend — audit your Cargo features.
  • Legacy PyTorch-interop users must move to CubeCL backends or burn-onnx/safetensors; tch and NdArray are on the way out.
  • No GGUF path — convert through ONNX or safetensors instead.
  • MSRV is now Rust 1.95.

Wrapping up

Burn 0.22 removes the last ergonomic tax on multi-backend Rust deep learning: no more backend generics haunting every function signature. Combine runtime Device selection with CubeCL’s six hardware targets and burn-train’s complete loop, and Rust stops looking like a prototype language for AI — it looks like a deployment strategy.

References