Quantization calculator

By Michael Lip. The preset sizes and the ladder on this page are measured values read from published file listings and hard-coded here; only the multiplication that combines them with your inputs runs in your browser. Your inputs are never sent to a server.

Estimate a quantized file size

Pick one of the seven Qwen2.5 Instruct presets, or enter a custom parameter count in billions. For a preset the tool shows the real size of Qwen's published GGUF file at that level, read from the file listing and cross checked on a second host, together with the measured bits per original parameter. For a custom count the tool estimates a size by reusing the measurement of the closest Qwen2.5 size at the same level. The load estimate adds the fp16 KV cache for the context you choose; treat it as an estimate, since runtime compute buffers are not included.

The measured ladder

Each cell shows the real file size in decimal GB (10^9 bytes, the unit the Hugging Face file browser shows) with the measured bits per original parameter under it. Bits per parameter is the byte count times 8 divided by the release parameter count. Sizes were read from the Hugging Face tree API on 2026-09-18 and cross checked against the ModelScope file listing.

Measured GGUF sizes for Qwen2.5 Instruct, read 2026-09-18 from the Hugging Face tree API and cross checked on ModelScope.
Level0.5B1.5B3B7B14B32B72B
fp161.27 GB
20.51 bpw
3.56 GB
18.45 bpw
6.80 GB
17.63 bpw
15.24 GB
16.01 bpw
29.55 GB
16.00 bpw
65.54 GB
16.00 bpw
145.93 GB
16.06 bpw
q8_00.68 GB
10.94 bpw
1.89 GB
9.82 bpw
3.62 GB
9.37 bpw
8.10 GB
8.51 bpw
15.70 GB
8.50 bpw
34.82 GB
8.50 bpw
77.53 GB
8.53 bpw
q6_k0.65 GB
10.53 bpw
1.46 GB
7.59 bpw
2.79 GB
7.24 bpw
6.25 GB
6.57 bpw
12.12 GB
6.57 bpw
26.89 GB
6.56 bpw
59.86 GB
6.59 bpw
q5_k_m0.52 GB
8.46 bpw
1.29 GB
6.66 bpw
2.44 GB
6.32 bpw
5.44 GB
5.72 bpw
10.51 GB
5.69 bpw
23.26 GB
5.68 bpw
51.67 GB
5.69 bpw
q5_00.49 GB
7.94 bpw
1.26 GB
6.53 bpw
2.38 GB
6.18 bpw
5.32 GB
5.58 bpw
10.27 GB
5.56 bpw
22.64 GB
5.53 bpw
50.34 GB
5.54 bpw
q4_k_m0.49 GB
7.96 bpw
1.12 GB
5.79 bpw
2.10 GB
5.46 bpw
4.68 GB
4.92 bpw
8.99 GB
4.87 bpw
19.85 GB
4.85 bpw
44.01 GB
4.84 bpw
q4_00.43 GB
6.94 bpw
1.07 GB
5.53 bpw
2.00 GB
5.18 bpw
4.43 GB
4.66 bpw
8.52 GB
4.61 bpw
18.64 GB
4.55 bpw
41.37 GB
4.55 bpw
q3_k_m0.43 GB
7.00 bpw
0.92 GB
4.79 bpw
1.72 GB
4.47 bpw
3.81 GB
4.00 bpw
7.34 GB
3.98 bpw
15.94 GB
3.89 bpw
35.47 GB
3.90 bpw
q2_k0.42 GB
6.72 bpw
0.75 GB
3.90 bpw
1.38 GB
3.57 bpw
3.02 GB
3.17 bpw
5.77 GB
3.13 bpw
12.31 GB
3.01 bpw
27.33 GB
3.01 bpw

Why a small model pays more bits per parameter

At q4_k_m the 0.5B model measures 7.96 bits per original parameter and the 72B model measures 4.84. The gap is mostly an accounting fact. The 0.5B config sets "tie_word_embeddings": true (see the Qwen2.5-0.5B-Instruct config.json), with vocab_size 151936 and hidden_size 896, so the release stores one matrix for the input embedding and the output. The GGUF records a separate output matrix, so the file stores more weights than the release counts: the GGUF tensor count for the 0.5B model is 630167424 while the release parameter count is 494032768. Dividing the file bits by the release's smaller count gives the higher figure.

Nominal widths versus what the files measure

The Hugging Face GGUF documentation describes Q4_K as "4-bit quantization (q). Super-blocks with 8 blocks, each block has 32 weights.", "resulting in 4.5 bits-per-weight". The same documentation lists nominal widths of "6.5625 bits-per-weight" for Q6_K, 5.5 for Q5_K, 3.4375 for Q3_K and 2.625 for Q2_K. Measured files land near but not on those nominal numbers, and the distance varies with model size. The llama.cpp quantize README shows the same effect from the other direction: it publishes measured bits/weight for Llama-3.1-8B, with Q4_K_M at 4.8944 bits/weight and 4.58 GiB, and F16 at 16.0005 bits/weight and 14.96 GiB. The calculator above never uses a nominal width; preset outputs come from the measured files only.

What the level does not change

Moving between levels of the same release does not change the architecture keys in config.json, the fp16 KV cache per token, or the attention structure. What changes is how many bytes each stored weight takes. The KV cache per token is read from each release's config.json the same way as on the KV cache size calculator. On the load question, the llama.cpp quantize documentation states: "At the moment, memory and disk requirements are the same." For full inference and fine-tuning memory, use the LLM memory calculator.

Worked example, Qwen2.5-7B-Instruct at q4_k_m

The file is 4683073632 bytes, which is 4.68 GB. The release's parameter count, safetensors.total from the Hugging Face model-info API, is 7615616512. Bits per original parameter is the byte count times 8 divided by the parameter count, which gives 4.92. The fp16 KV cache at a 128k context is 7 GiB, so the load estimate at that context is the file plus that cache. The load figure is an estimate; runtime compute buffers are not included.

Where these numbers come from

Two hosts publish these files. The first is the Hugging Face tree API, which lists every part of every split GGUF with its byte size. The second is the ModelScope file listing, an independent mirror. Each size on this page was read from the first host and cross checked against the second, and the two matched. Every figure is computed in code from those listings and never typed. A parameter count comes from each release's safetensors.total in the Hugging Face model-info record. Architecture facts, such as the tied embeddings of the small models, come from each model's published config.json.