Multi-Label LLM Pretraining With A Smaller Teacher Model

CSEF · 2026 Computational Science (Senior Division)

Overview

Standard causal language model pretraining uses a single-label cross-entropy objective that ignores the existence of multiple valid next-token continuations resulting in sample inefficiency. In this work, we introduce a multi-label pretraining objective that modifies the loss to append a small set of context-conforming auxiliary tokens selected by a lightweight surrogate language model. Distinct from existing knowledge distillation methods, the surrogate is used only for token selection rather than full distribution matching. Upon training OLMo2-1b for ~52 billion tokens, we found that this method achieves comparable benchmark performance (e.g. PIQA and WinoGrande) with significantly fewer training tokens and optimization flops in our tested setting. The results demonstrate that fixing token-level label inefficiencies through multi-target objectives can reduce pretraining expenses when subjected to constrained compute.

Competition history

  • CSEF 2026 Computational Science (Senior Division) · Entry S-07-22

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: California Science & Engineering Fair public projects

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google