Skip to content
Blog

GGUF Demystified: Why Local AI Runs on This Format

GGUF is the single-file format behind llama.cpp, Ollama, and LM Studio. Here is what is inside it, how quantization levels trade size for quality, and how to convert and pick the right quant for your hardware.

Published on • October 4, 2026

AI Assistant

What GGUF Actually Is

GGUF is the model file format of llama.cpp, the C/C++ LLM inference project that powers most of the local-AI tooling you have probably used: Ollama, LM Studio, llamafile, KoboldCpp, and the llama-server OpenAI-compatible endpoint. If you have ever downloaded a .gguf file and dragged it into a desktop app, you have handled the de facto standard of on-device inference.

It is the successor to GGML, llama.cpp’s original format. GGML used a flat, sequential tensor layout that made adding new metadata awkward — every new model architecture meant a format question. GGUF fixed that by pairing the tensor data with a self-describing metadata section: key-value pairs that say what architecture this is, what the rope scaling parameters are, which chat template to use, and how the tokenizer works.

The result is a format designed around one idea: everything the runtime needs travels in one artifact.

Why Single-File Matters

Before GGUF, running a Hugging Face checkpoint locally meant juggling config.json, tokenizer.json, tokenizer_config.json, model shards, and generation configs — and hoping they all described the same model. Miss one file or mix two directories and you get a tokenizer mismatch or a silently wrong chat template.

GGUF collapses that into one file containing:

  • Model metadata — architecture (llama, qwen2, gemma, …), context length, rope parameters, block count, attention heads
  • Tokenizer data — the full vocabulary, merges, special tokens, chat template
  • Tensor info — names, shapes, dtypes, byte offsets
  • Tensor data — the actual weights, quantized

Copy the file, you copy the model. A tool can read the metadata header in kilobytes without loading gigabytes of weights, reject an unsupported architecture before allocating memory, and render the correct chat template without a sidecar JSON. That is why llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF can pull and run a model with one command — the format is self-contained.

Quantization Levels Explained

Quantization stores each weight with fewer bits. llama.cpp supports a wide range — the project lists 1.5-bit through 8-bit integer quantization — but in practice you will meet four families:

QuantBits/weight (approx)QualityTypical use
Q8_0~8.6 (8 + scale)Near-losslessWhen you have plenty of VRAM; archival
Q6_K~6.6Excellent, hard to distinguish from fp1624GB cards, “just make it good”
Q5_K_M~5.5Very goodThe quality/size sweet spot for 16GB
Q4_K_M~4.8Good; slight loss on hard reasoningThe default recommendation for 8–12GB
Q4_0 / Q3_K_M~4 / ~3.5Noticeable degradationTight 8GB budgets, experimentation

The K in Q4_K_M means k-quant: a later generation that groups importance-aware blocks, giving measurably better perplexity than older fixed-block formats at the same bit width. The suffix is granularity: Q4_K_S (small) is slightly smaller and slightly worse than Q4_K_M (medium); _L is large.

The trade-off curve is steep at the bottom and flat at the top. Dropping from fp16 to Q4_K_M typically costs a percent or two of relative perplexity while cutting size roughly in half — that is why Q4_K_M became the community default. Q4 to Q2 costs real capability: repetition, weaker instruction following, broken long-context behavior. As a working rule: use the highest quant that fits your memory, and stop worrying at Q5_K_M or above.

Memory Math: Weights + KV Cache

Two numbers decide whether a model runs:

weights   = params × bytes_per_param
kv_cache  = 2 × layers × n_kv_heads × head_dim × seq_len × bytes
total     = weights + kv_cache + activations (~10–20% headroom)

For a 7B-parameter model, bytes_per_param is roughly 0.6 for Q4_K_M (≈4.2GB), 0.82 for Q5_K_M (≈5.7GB), and 1.08 for Q6_K (≈7.5GB). Add the KV cache, which grows linearly with context — hundreds of MB at 8k, rivaling the weights at 128k — plus runtime overhead.

Concrete picks:

  • 8GB VRAM/RAM: 7–8B at Q4_K_M, or 3B at Q5_K_M
  • 16GB: 14B at Q4_K_M or Q5_K_M, 8B at Q6_K
  • 24GB: 32B at Q4_K_M/Q5_K_M, 14B at Q6_K
  • 32GB+: 32B at Q6_K or Q8_0; 70B at Q4_K_M with tight context

llama.cpp’s CPU+GPU hybrid inference lets you partially offload layers that do not fit in VRAM, so a 32B model can run on a 12GB card — slower, but functional.

File Structure Basics

A GGUF file is laid out linearly:

  1. Magic — the 4 bytes GGUF (little-endian 0x46554747)
  2. Version — uint32 (v3 is current; v2 still appears in old files)
  3. Tensor count and metadata KV count — uint64 each
  4. Metadata KV pairs — string key, typed value (uint8/16/32/64, float, bool, string, array), recursively nested
  5. Tensor info — name, dimensions, dtype, offset, per tensor
  6. Padding — up to the file alignment (default 32 bytes)
  7. Tensor data — raw quantized blocks at the offsets declared above

Because metadata comes first, any tool can cheaply inspect a file:

from gguf import GGUFReader

r = GGUFReader("model-Q4_K_M.gguf")
arch = r.get_field("general.architecture")
print(arch.contents())            # e.g. "qwen2"
print(r.get_field("tokenizer.chat_template").contents()[:120])

print(f"{len(r.tensors)} tensors")
for t in list(r.tensors)[:3]:
    print(t.name, t.shape, t.n_elements)

gguf-py, the reference Python library shipped in the llama.cpp repository, is what Hugging Face uploaders and quantization tools use to write and read these files.

Converting a Hugging Face Model to GGUF

The workflow lives in llama.cpp and is two steps: convert to a full-precision GGUF, then quantize.

# 1. Convert HF weights to f16 GGUF
python convert_hf_to_gguf.py `
  --outfile my-model-f16.gguf `
  --outtype f16 `
  ./hf-checkpoint-dir

# 2. Quantize to your target level
llama-quantize my-model-f16.gguf my-model-Q4_K_M.gguf Q4_K_M

Notes that save hours:

  • Keep the f16 GGUF around; it is the master you quantize every target from.
  • --outtype q8_0 produces an already-quantized file if you want to skip the f16 intermediate.
  • The converter validates the tokenizer and writes metadata — if it warns about an unknown chat template, fix it before distributing the file.
  • Modern llama.cpp can pull GGUF straight from Hugging Face with -hf, so for popular models you can skip conversion entirely.

Ecosystem Compatibility

GGUF’s reach extends well beyond llama.cpp itself:

  • llama.cpp — native; llama cli, llama-server, GBNF grammars
  • Ollama — imports GGUF (a Modelfile can reference one directly)
  • LM Studio — drag-and-drop local inference over GGUF
  • llama-cpp-python — Python bindings exposing an OpenAI-style API and local embeddings
  • HF transformers — loads GGUF via the gguf package, though llama.cpp remains the reference implementation

That shared format is why an ecosystem formed: publish one file, and every local runtime can consume it.

Choosing a Quant for Your Hardware

A practical decision procedure:

  1. Compute your budget — available VRAM/RAM minus ~15% for the runtime and KV cache.
  2. Filter by fit — find the largest model whose params × bytes/param fits that budget.
  3. Pick the highest quant that still fits — if a 14B fits at Q4_K_M and Q5_K_M, take Q5_K_M; if only Q4_K_M fits, take it.
  4. Check the KV cache — long context (32k+) may force you down a quant level to leave room.
  5. Benchmark on your own tasks — perplexity charts are averages; your domain (code, math, multilingual) may weight differently.

Also match the quant to the hardware path: CPU-heavy setups favor Q4_K_M (fewer bytes moved per token), while fast VRAM cares about total fit over bandwidth.

Common Pitfalls

  • Wrong chat template in metadata. Symptom: the model answers a system prompt as if it were user text, or never stops. The template lives inside the GGUF — verify tokenizer.chat_template rather than assuming your UI will apply one.
  • Vocabulary mismatches. Merging an edited tokenizer.json with an unrelated GGUF, or using a base-model GGUF with a chat finetune’s prompt style, produces gibberish special tokens. Always convert tokenizer and weights together.
  • Mixing architectures. Metadata says qwen2; you force llama prompts. Let the runtime read general.architecture.
  • Chasing Q2. It fits more, but the quality cliff is real. If Q3 is the only thing that fits, consider a smaller model at Q4_K_M instead.
  • Context blow-ups. A model that runs fine at 4k OOMs at 32k because the KV cache was never in your arithmetic.

GGUF is boring in the best way: one file, explicit metadata, predictable math. Learn it once and every local-model workflow — download, convert, quantize, deploy — becomes three commands.

Sources