· AI research
The Student Becomes the Teacher
Why this mattersModels that teach themselves could reduce the need for labelled data, but they can also reinforce their own mistakes.
In battery research, a single ground-truth requires assembling a cell, cycling it for weeks, and extracting one electrochemical data point. This consumes both time and money and post-training large language models has the exact same problem. Verified reasoning traces, gold answers, and external reward models are the labels, and they are just as scarce relative to the volume of data a model could learn from. This week, both Guibin Zhang from NUS and Yijiang Li from UC San Diego pushed the question "what if the model supplies its own supervision". Short answer: the external teacher (OPSD), the ground-truth labels (U-OPSD), and now handcrafted privileged context have all been replaced by the model itself.
Background: one model, two contexts
The line of work starts with On-Policy Self-Distillation (OPSD, January 2026, arXiv:2601.18734). A single LLM acts as both teacher and student under different contexts. The teacher conditions on privileged information (verified reasoning traces) while the student sees only the question. Training minimizes the per-token divergence between the two distributions over the student's own rollouts, on trajectories it would actually produce.
Self-Distillation Fine-Tuning (SDFT) applies the same trick to demonstrations: the demonstration-conditioned model becomes its own teacher, generating on-policy training signals that, in the original experiments, preserved prior capabilities better than ordinary SFT while acquiring new skills. It was one of the papers put through the ICML 2026 open reproduction effort (Hugging Face report), I come back to that below.
Remove or learn the supervision
U-OPSD, On-Policy Self-Distillation without any Supervision arXiv:2608.06296 removes the last external crutch. Sample k rollouts, majority-vote a pseudo-solution, gate it by a self-consistency threshold, then condition the model on that pseudo-solution and distill only on the disagreeing completions. The model corrects itself where it is confidently wrong. An answer the model produces consistently is treated as correct, the agreement among samples is taken as evidence (in place of any external check). On Qwen3 math in non-thinking mode, U-OPSD beats supervised OPSD by up to 3.2 points and GRPO by up to 8.9 points (summary), using no labels at all.
L-OPSD, Latent On-Policy Self-Distillation arXiv:2608.13040 steps in the opposite direction. Instead of removing the privileged context, it makes privilege learnable: retrieved experiences are composed into continuous latent tokens that condition the self-teacher and dense supervision is provided at every visited prefix. Because the composer is trainable, the distillation loss has a degenerate minimum (the teacher collapses into the student for uninformative context). The privileged-margin objective forbids this by requiring the teacher to maintain a verifiable log-probability advantage over the student on successful trajectories. The latent context must stay informative, not merely easy to match. Without the constraint, joint training performs worse than freezing the composer outright. With this constraint, L-OPSD outperforms GRPO, OPSD, SDPO, and Skill-SD on agentic tool use and code generation, using under 30% of GRPO's rollout budget.
One paper removes privileged information entirely, the other learns to construct it. Which algorithm wins depends on the failure mode you're more afraid of; the model confidently agreeing with its own mistakes vs. the teacher collapsing into a mirror of the student.
Some limitations
The privilege illusion. The DAPD analysis (coverage) argues that OPSD quietly teaches the student to behave as if the reference were still present at inference time. The measured consequence: gains of +5.19 points at 1.7B parameters collapse to ≤ +0.28 at 8B–32B. What works at small scale may be an artifact of scale.
Decoding collapse. In agentic settings, feedback-augmented self-distillation can produce trajectories that look diverse but are largely input-agnostic templates. The KL-based supervision signal becomes uninformative, and standard evaluation metrics miss it (arXiv:2607.17558). An exponential-moving-average teacher is a partial fix.
The awkward reproduction. In the ICML 2026 open reproduction challenge, at least one independent attempt found no SDFT advantage over matched SFT on a 507-example science evaluation, leaving toy-scale support only for the reduced-forgetting claim (reproduction logbooks). One reproduction is not a refutation, but it is exactly the kind of result that should be cited next to the original.
The meta-point: every failure above is an examination failure, not a teaching failure. Self-generated supervision inherits the model's blind spots, and internal consistency measures confidence, not correctness. Majority vote repairs the cases where the model is right often and wrong sometimes; it amplifies the cases where it is confidently, systematically wrong.
Why I care
Process–property models in electrode manufacturing have the same structure as this problem: abundant process data from coating and calendering lines, scarce and delayed electrochemical labels from cycling tests. Any method that extracts training signal from internal consistency rather than external labels is directly relevant to surrogate models that have to stay useful between cycling campaigns.
The SDFT thread also connects to my continual-learning work. On-policy self-signals reduce forgetting relative to off-policy SFT, which makes them a candidate complement to structural protection mechanisms like the interference-gated approach I wrote about earlier. A speculative extension: could majority-vote pseudo-labels over process-parameter rollouts bootstrap a surrogate model during the weeks while a cycling campaign is still running? The failure modes above, especially confident systematic error, are exactly what a battery surrogate would exhibit under process drift, so the caveats transfer along with the method.
A geometric caution
The temptation, after reading this literature, is a geometric intuition: self-distillation keeps updates KL-close to the current policy, KL-close means small parameter drift, small drift means the model stays in a flat region of its loss landscape, and flat regions retain. It is a good story.
A recent theory paper is a useful antidote. Flat Minima and Generalization: Insights from Stochastic Convex Optimization (arXiv:2511.03548) constructs settings (in the friendliest possible regime, convex and smooth) where provably flat minima generalize trivially, and where a SAM-style algorithm converges quickly to exactly such points. It was the most-reproduced paper in the ICML 2026 challenge, with twelve independent teams verifying every claim, so this is not a fringe counterexample. Geometry is suggestive; it is never sufficient.
Models can teach themself. The student has become the teacher. It has not yet become the examiner.
Use your identity on another device
Save an encrypted identity file and restore it in another browser. Keep the file and its passphrase private.