{"event":{"id":"evt-emergent-abilities-20220615","dedupe_key":"scaling_behavior_research:2022-06-15:emergent-abilities-paper-formalizes-discontinuous-capability-observations-at-model-scale","event_type":"scaling_behavior_research","title":"Emergent Abilities paper formalizes discontinuous capability observations at model scale","summary":"Researchers documented benchmark tasks on which larger language models appeared to acquire abilities only beyond scale thresholds, including prompted behaviors not smoothly predictable from smaller-model performance.","occurred_at":"2022-06-15T17:32:01.000Z","published_at":"2022-06-15T17:32:01.000Z","observed_at":"2026-10-01T18:14:00.000Z","reconstructed_at":"2026-10-01T18:14:00.000Z","ingest_type":"live_world_scan","locations":null,"status":"verified","confidence":0.95,"metadata":{"actors":["Google Research","Stanford","DeepMind and collaborators"],"what_happened_at_time":"A contemporaneous research framing explicitly treated some capability gains as discontinuous with scale and difficult to predict by smooth extrapolation.","evidence_at_time":"The June 15 preprint compared multiple large model families and catalogued prompted tasks with threshold-like behavior.","affected_layers":["layer-research","layer-models","layer-compute","layer-ai"],"change_kind":"epistemic_uncertainty_increase","q1_continuity":"Q1 Chinchilla emphasized compute-optimal scaling. Q2 added evidence that capability behavior might not scale smoothly even when resource scaling itself is systematic.","retrospective_inference":"Capability forecasting became a distinct uncertainty inside the models/compute stack: more scale could reveal qualitatively different benchmark behavior.","uncertainty":"Later research debated whether apparent emergence can result from metrics and evaluation choices. Q2 reconstruction records what this paper established then, not a later verdict."},"created_at":"2026-10-01T19:22:17.750Z","updated_at":"2026-10-01T19:22:17.750Z"},"artifacts":[{"id":"art-emergent-abilities-20220615","canonical_url":"https://arxiv.org/abs/2206.07682","artifact_type":"paper_preprint","title":"Emergent Abilities of Large Language Models","summary":"The paper documented benchmark abilities that appeared only above model-scale thresholds and argued that some prompted capabilities were not predictable by smooth extrapolation from smaller models.","creator_entities":["Google Research","Stanford","DeepMind and collaborators"],"released_at":"2022-06-15T17:32:01.000Z","content_hash":null,"metadata":{"historical_scope":"2022-Q2"},"created_at":"2026-10-01T19:22:17.743Z","updated_at":"2026-10-01T19:22:17.743Z","relation":"documented_by"}],"signals":[{"id":"sig-q2-capability-predictability-uncertainty","statement":"Q2 research increased uncertainty about capability forecasting by documenting benchmark behaviors that appeared only above model-scale thresholds.","signal_type":"epistemic_uncertainty","direction":"increasing","confidence":0.89,"epistemic_status":"supported_inference","mapping_mode":"retrospective","reconstructed_at":"2026-10-01T18:14:00.000Z","created_at":"2026-10-01T19:22:17.755Z","updated_at":"2026-10-01T19:22:17.755Z"}]}