Test loss falls as a power law in training compute
Supports Tutoring quality tracks model capability, which is currently a predictable function of compute — so personalisation improves as systems scale.
Provenance Replication of published scaling-law figures, three model sizes · 2024
Method Median loss across held-out evaluation tasks, log-log axes, power-law fits of the form L = AC^-α. Loss is a proxy for capability, not for pedagogy.
Added by @kenji