Chinchilla scaling calculator

By Michael Lip. Enter a training compute budget C in FLOPs. The page gives the compute-optimal parameter count N and token count D under two published fits of the same loss formula: the Hoffmann et al. (2022) fit and the 2024 replication fit by Besiroglu et al. It also shows the gap between the two fits, the tokens-per-parameter ratio of each, and the predicted loss for a plan that you enter. All arithmetic runs in this browser tab.

Compute a plan

Compute budget
Your plan (optional)

Enter N or D and the page takes the other one from C = 6ND. Enter both and the page ignores the C field: it uses C = 6ND of your plan for the plan and for the optimum.

Constants of the two fits

Both fits use the loss formula L(N, D) = E + A / Nα + B / Dβ. The exponent α goes with the parameter count N and β goes with the token count D. The calculator uses the first two columns. The third column is the rounded set printed in the body of the 2022 paper, shown for comparison only. The rows a, b and G are computed on this page from the constants with Eq. 4 of the 2022 paper.

ConstantHoffmann et al. 2022, preciseReplication 2024Hoffmann et al. 2022, rounded
SourcearXiv 2404.10102, Eq. 4arXiv 2404.10102, Eq. 3arXiv 2203.15556, Eq. 10
E1.69341.81721.69
A406.4482.01406.4
B410.72085.43410.7
α (N)0.33920.34780.34
β (D)0.28490.36580.28
a = β / (α + β)0.45650.51260.4516
b = α / (α + β)0.54350.48740.5484
G1.30040.11961.3447

The replication paper says that its Eq. 4 values come from comments in the TeX source of the 2022 paper and are more precise than the values in the body of that paper. The 2022 paper states a = 0.46 and b = 0.54 for this approach. The precise column rounds to these values.

Formulas

The 2022 paper finds the optimum by minimizing L(N, D) under the constraint FLOPs(N, D) ≈ 6ND (its Eq. 4 and the text above it). The factor 6 comes from Kaplan et al. (2020), who estimate the non-embedding training compute as C ≈ 6N floating point operations per training token, with N the number of non-embedding parameters and the backward pass at about twice the cost of the forward pass. The 2022 paper counts the embedding matrices in its parameter count and in its FLOPs (its Appendix F). The closed form is:

N_opt(C) = G × (C / 6)^a
D_opt(C) = G^(-1) × (C / 6)^b
G = (αA / (βB))^(1 / (α + β)), a = β / (α + β), b = α / (α + β)
tokens per parameter = D_opt / N_opt
loss of a plan = E + A / N^α + B / D^β
gap = value of the replication fit / value of the Hoffmann fit

When you enter N only, the page sets D = C / (6N). When you enter D only, it sets N = C / (6D). The loss penalty of a plan is the loss of the plan minus the loss at the optimum of the same fit for the same C.

Worked examples

Gopher budget. The 2022 paper gives the Gopher training budget as 5.76 × 1023 FLOPs and writes: "We project the optimal model size given the Gopher FLOP budget to be 40B parameters." With the precise constants, this page gives N_opt = 4.03 × 1010 and D_opt = 2.38 × 1012, which is 59.1 tokens per parameter. With the rounded constants of Eq. 10, the same budget gives N_opt = 3.22 × 1010, so the rounding alone moves the answer away from 40B. The replication fit gives N_opt = 7.22 × 1010 and D_opt = 1.33 × 1012, which is 18.4 tokens per parameter. The gap is 1.79 for N and 0.558 for D.

Check of C = 6ND with Table 3 of the 2022 paper. 400 million parameters and 8.0 billion tokens give 1.92 × 1019 FLOPs. 1 billion parameters and 20.2 billion tokens give 1.21 × 1020. 10 billion parameters and 205.1 billion tokens give 1.23 × 1022. These three products agree with the FLOPs column of Table 3. Table 3 comes from Approach 1 of the paper, not from the fitted loss formula, so it tests only the 6ND rule.

Plan loss. For N = 1 × 1010 and D = 2.051 × 1011, the plan uses C = 1.23 × 1022. The predicted loss is 2.1041 under the Hoffmann fit and 2.1293 under the replication fit. At the optimum for the same C, the Hoffmann fit gives 2.1015, a penalty of 0.0026. The replication fit gives 2.1293 at its optimum, a penalty below 0.0001.

Why the two fits disagree

The replication paper reconstructs the data of Figure 4 of the 2022 paper and fits the same formula again. It states that the Hoffmann parameters suggest "approximately 70 tokens per parameter" and that its own fit implies "around 20 tokens per parameter". It gives two reasons for the difference: the constants in the body of the 2022 paper are rounded, and the optimizer of the 2022 fit stopped before convergence. Under the precise Hoffmann constants, a is less than b, so the ratio D_opt / N_opt grows with C. This page computes 27.8 tokens per parameter at 1020 FLOPs, 59.1 at the Gopher budget and 92.5 at 1026 FLOPs. Under the replication fit, a is more than b, and the ratio falls slowly: 22.9 at 1020, 18.4 at the Gopher budget and 16.1 at 1026. The replication paper also states that its fit is consistent with a ratio between 4 and 40 for models trained on 1026 FLOPs or more.

Limits

Each fit describes the data and the tokenizer of the 2022 paper only. That paper trained "over 400 language models ranging from 70 million to over 16 billion parameters" (its abstract). Its Table A9, the list of all its models, runs from 44 million to 16,183 million parameters. Outside that range the formulas are an extrapolation, and the calculator shows a warning when a parameter count is below 44 million or above 16,183 million. The page does not predict the loss or the optimal size of any named model. It does not model cost, hardware time or prices. The loss values are in the units of the 2022 paper and do not transfer to a different dataset.

Sources

All three sources were fetched on 2026-10-11. Hoffmann et al. (2022), "Training Compute-Optimal Large Language Models", arXiv 2203.15556: Eq. 4, Eq. 10, Table 2, Table 3, Appendix F and the Gopher budget of 5.76 × 1023 FLOPs, with Table A9 for the range of model sizes. Besiroglu et al. (2024), "Chinchilla Scaling: A replication attempt", arXiv 2404.10102 (version 2): Eq. 3, Eq. 4, Table 1 and Section 3.3. Kaplan et al. (2020), "Scaling Laws for Neural Language Models", arXiv 2001.08361: Section 2.1, C ≈ 6N per training token. Links: arXiv 2203.15556, arXiv 2404.10102, arXiv 2001.08361.