paper preprint
Training Compute-Optimal Large Language Models
DeepMind reported that many large language models were undertrained for their compute budgets and that model size and training tokens should scale together under its empirical compute-optimal analysis.