{"event":{"id":"evt-minerva-20220629","dedupe_key":"quantitative_reasoning_model_release:2022-06-29:minerva-extends-large-language-model-reasoning-into-technical-mathematics-and-science","event_type":"quantitative_reasoning_model_release","title":"Minerva extends large-language-model reasoning into technical mathematics and science","summary":"Google introduced Minerva, a PaLM-based model further trained on technical content and using step-by-step prompting plus majority voting to improve mathematics and science benchmark performance without external tools.","occurred_at":"2022-06-29T18:54:49.000Z","published_at":"2022-06-29T18:54:49.000Z","observed_at":"2026-10-01T18:14:00.000Z","reconstructed_at":"2026-10-01T18:14:00.000Z","ingest_type":"live_world_scan","locations":null,"status":"verified","confidence":0.98,"metadata":{"actors":["Google Research"],"what_happened_at_time":"A large language model achieved much stronger quantitative benchmark results through domain data, chain-of-thought-style prompting and sampling/majority-vote inference.","evidence_at_time":"The June 29 preprint and June 30 Google Research post documented the method, benchmark results and limitations.","affected_layers":["layer-ai","layer-models","layer-research","layer-software"],"change_kind":"reasoning_capability_maturation","q1_continuity":"Q1 chain-of-thought was a prompting precursor. Minerva showed that it could be combined with large-scale pretraining, domain data and inference-time sampling to materially improve technical reasoning benchmarks.","retrospective_inference":"Reasoning capability was becoming a stack composition of model scale, specialized data, prompt scaffolding and inference strategy rather than only a property of raw weights.","uncertainty":"Minerva explicitly lacked formal verification and could reach correct final answers through incorrect reasoning; it did not use tools or establish reliable mathematical agency."},"created_at":"2026-10-01T19:22:17.751Z","updated_at":"2026-10-01T19:22:17.751Z"},"artifacts":[{"id":"art-minerva-20220629","canonical_url":"https://arxiv.org/abs/2206.14858","artifact_type":"paper_preprint","title":"Minerva quantitative reasoning model","summary":"Google introduced Minerva, a PaLM-based model further trained on technical content and using chain-of-thought-style prompting and majority voting to achieve strong results on mathematics and science benchmarks without external calculation tools.","creator_entities":["Google Research"],"released_at":"2022-06-29T18:54:49.000Z","content_hash":null,"metadata":{"historical_scope":"2022-Q2"},"created_at":"2026-10-01T19:22:17.744Z","updated_at":"2026-10-01T19:22:17.744Z","relation":"documented_by"}],"signals":[{"id":"sig-q2-reasoning-composition","statement":"Quantitative reasoning gains increasingly came from composing model scale with specialized training data, chain-of-thought-style prompting and inference-time sampling rather than from raw parameter count alone.","signal_type":"capability_composition","direction":"increasing","confidence":0.96,"epistemic_status":"supported_inference","mapping_mode":"retrospective","reconstructed_at":"2026-10-01T18:14:00.000Z","created_at":"2026-10-01T19:22:17.756Z","updated_at":"2026-10-01T19:22:17.756Z"}]}