ml0x.com/triage
Your run failed at hour six and the internet will tell you what a learning rate is
Seven guides and four scripts for the part nobody writes down. Not what the knobs mean, but what to set them to, and the source each figure came from with the conditions it was measured under. Ninety nine dollars, paid once.
What you get
$99 once. Seven PDF guides totalling 183 pages, four Python scripts, two offline calculators, and a source map beside every guide.
When the loss goes to NaN
A decision tree for divergence. What to check first, what each observation rules out, and where gradient clipping actually belongs relative to the scaler.
Picking a learning rate and a schedule
The range test as it was actually described, why the chosen point is not the minimum of the curve, and how batch size moves it. Every starting value carries the conditions it was measured under.
Reading metrics that disagree
Micro against macro against weighted, why F1 can sit below both precision and recall, and why ROC-AUC stays high while everything else collapses on imbalanced data. Worked by hand so you can check the arithmetic.
Choosing the optimizer
Where L2 and decoupled weight decay actually diverge, what each optimizer costs in bytes per parameter, and what the published comparisons do not establish.
GPU memory for training and inference
The four terms, counted. Which ones move with batch size and which do not, and the three that are routinely misremembered.
Reading PyTorch shape errors
Real error strings, the actual shapes behind them, and the fix. The longest guide in the set.
The preflight checklist
What to verify before you commit hours of compute, each item with the cost of skipping it. Includes the single batch overfit test, which is the fastest correctness test there is.
Four scripts and two calculators
A memory budget calculator, a learning rate range test, a threshold sweep, and the overfit runner. Plus the KV cache and model memory calculators, offline, no network.
The part that took the time
Every guide ships with a source map. Each formula, constant and default is traced to a primary reference, with a URL, and with the conditions under which it was measured. Arithmetic that is mine is marked as derived and reproduced so you can check it without a calculator.
Writing it that way is slower, and it is the whole product. A default with no conditions attached is the thing that made your run fail.
What this is not
- Not a course. No video, no cohort, no community.
- Not a tutorial for a first model. It assumes you have trained things before and are debugging, not learning.
- Not a benchmark suite. It carries no throughput or speedup claims, because I did not measure any.
- Not a replacement for the free calculators on this site. Those stay free and they stay free after you buy.
No buyers yet. This went on sale today, so there are no testimonials on this page and there will not be invented ones.
Ninety nine dollars, once
Payment runs through Stripe. You land on a download page immediately, the archive is yours to keep on every machine you own, and there is no licence key and no account to create.
Buy Training Run Triage, $99Stripe checkout opens. Card, Apple Pay or Google Pay. Then a download page.
Sixty day refund, no questions asked, and you keep the files either way. If a formula in here turns out to be wrong I would rather hear about it than not, and corrections go to everyone who bought.