Burn: A Rust-Native Deep Learning Framework Comes of Age
Burn 0.22 drops backend type parameters for runtime Device selection, unifies CUDA, ROCm, Metal, Vulkan and WebGPU behind CubeCL, and ships a batteries-included training story.
Published on • October 9, 2026
AI Assistant

Python owns deep learning tooling, but the production inference and training niche keeps moving to Rust. Burn is the most ambitious of the Rust-native frameworks: a pure-Rust deep learning library whose single codebase targets CUDA, ROCm, Metal, Vulkan, WebGPU, CPU, WebAssembly, and even no_std.
Version 0.22.0, released October 6, 2026, is the release where Burn’s ergonomics caught up with its ambitions.
What Burn is
Burn is organized around three pillars:
- A tensor abstraction with autodiff built in
- A backend system — you swap execution engines without rewriting models
- Training tooling —
burn-trainprovides a fullLearnerloop with metrics, checkpointing, and early stopping
The headline of 0.22: backend type parameters are gone from the user API. Execution is now selected at runtime through a Device value, and Tracel claims rebuilds are up to 15× faster as a result.
use burn::tensor::Device;
let device = Device::cuda(0); // or Device::wgpu(Default::default())
let device = Device::flex().autodiff(); // pure-Rust backend + autodiff
let model = Model::new(&device);
let input = Tensor::<2>::ones([8, 4], &device);
let gradients = model.forward(input).sum().backward();
Autodiff is now runtime state (device.autodiff()), and training/dataset APIs became fallible — device failures return Err instead of panicking.
The backend matrix after 0.22
| Backend | Engine | Notes |
|---|---|---|
| CubeCL (CUDA, ROCm, Metal, Vulkan, WebGPU, CPU) | JIT via LLVM | Fusion, autodiff, remote execution decorators |
burn-flex | Pure Rust eager | std, no_std, Wasm |
burn-tch (LibTorch) | Deprecated | ”will be removed” |
| NdArray | Deprecated | legacy |
| Candle backend | Removed | feature flags dropped |
Important consequence: burn ships no default execution backend. You must enable a Cargo feature per backend constructor, and Device::default() picks from compiled-in backends rather than detected hardware — enabling another feature can silently change what default() resolves to.
Only burn-flex works under no_std.
Importing models
Burn’s interop story is ONNX and safetensors, not GGUF:
burn-onnxruns at build time viabuild.rs, generating Rust code plusburnpackweights from an ONNX graph. Opsets 1–24, graph simplification on by default:
use burn_onnx::ModelGen;
fn main() {
ModelGen::new()
.input("src/model/my_model.onnx")
.out_dir("model/")
.run_from_script();
}
- SafeTensors support landed in 2025.
- GGUF has no first-party support — it remains an open design discussion (issue #1187). GGUF carries metadata plus weights, so it maps to the weights-loading path rather than the graph-import path, and model structure must still be defined manually.
Training with burn-train
The Learner bundles model, optimizer, and LR scheduler, driven by SupervisedTraining over train/validation dataloaders:
let optim = AdamWConfig::new().with_weight_decay(5e-5).init();
let result = training.launch(Learner::new(model, optim, lr_scheduler));
Included: loss, accuracy, precision/recall, F-scores, AUROC, BLEU, CER, WER, perplexity, and system-usage metrics — plus PSNR/SSIM/LPIPS/FID behind the vision feature. Checkpointing supports periodic and metric-based saves (burnpack format, keep-last-two plus best-val), with early stopping, multi-device/DDP training, and a tui terminal dashboard by default.
Ecosystem
burn-train, burn-onnx, burn-store, burn-dataset, burn-bench, burn-rl, burn-vision, burn-fusion, burn-autodiff, and the CubeCL crates (burn-cubecl, burn-linalg) all version together at 0.22.
Burn vs. Candle
The two Rust ML stacks have clearly diverged (and the Candle post on this blog covers the other side):
- Candle — minimalist, Hugging Face ecosystem native, GIL-free serverless inference, leans inference-only.
- Burn — batteries-included, backend-portable, with a real training story and compile-time kernel fusion.
If you want training loops, multi-GPU, and one model definition targeting six hardware stacks, Burn is the answer. If you want a minimal HF-compatible inference crate, Candle still wins.
Gotchas
- 0.22 is a breaking migration:
Tensor<B, D>→Tensor<D>,server→remote-serverfeatures, and more. Budget migration time. - No default backend — audit your Cargo features.
- Legacy PyTorch-interop users must move to CubeCL backends or burn-onnx/safetensors;
tchand NdArray are on the way out. - No GGUF path — convert through ONNX or safetensors instead.
- MSRV is now Rust 1.95.
Wrapping up
Burn 0.22 removes the last ergonomic tax on multi-backend Rust deep learning: no more backend generics haunting every function signature. Combine runtime Device selection with CubeCL’s six hardware targets and burn-train’s complete loop, and Rust stops looking like a prototype language for AI — it looks like a deployment strategy.
References
- tracel-ai/burn on GitHub - framework source, issues, and examples
- Burn 0.22.0 release announcement - what changed in the release
- Migrating to Burn 0.22 - breaking changes and migration table
- The Learner API - training loop documentation