paper preprint

Training Compute-Optimal Large Language Models

DeepMind reported that many large language models were undertrained for their compute budgets and that model size and training tokens should scale together under its empirical compute-optimal analysis.

Original sources

Open artifact View JSON