Home
All posts

· Continual learning

Finding a Fair Evaluation Under Changing Representations

Why this mattersA learning system should be tested on whether it can adapt and remember when its internal representation changes.

LLMs handle well above 3 billion prompts per day (traditional search engines handle 8.5–16 billion queries, for comparison). Those models, however, do not improve during those interactions: their weights do not change meaningfully. Some information may be stored in databases, but the LLM itself forgets every interaction outside context windows and related mechanisms.

Training systems that remember first requires a way to evaluate retention. This is trickier than it seems.

Raw accuracy is one possible comparison. But when a frozen model already obtains perfect scores, there is no incentive for adaptation. The math for frozen models is easy, and quite a bit harder for an adapting representation. The unsolved problem is thus certified retention under a changing representation.

Why changing representations make retention hard

Retention is easy to guarantee if you freeze the model’s perception and only update a readout on top. Then the features are constants, and the sufficient statistics are additive in the fixed coordinate system. This is why frozen-feature continual learning works so well, and why most approaches inherit or rely on a fixed backbone.

Let the representation move, and those assumptions fail. Sufficient statistics stop being sufficient because they are expressed in stale coordinates, and updates become misaligned once the update distribution shifts. Certified retention under a changing representation is therefore a fundamentally different problem.

This creates a simple but important benchmark-design challenge: how can a task make adaptation actually necessary?

First, make adaptation necessary

Since strong frozen encoders win nearly everywhere against adapting representations, the first task is to create a problem where adaptation is advantageous or, one step further, where frozen encoders provably cannot succeed but the adapting representation can.

On existing benchmarks, a good encoder plus a moderately sized head exceeds the performance of most adaptive systems. In that regime, any claimed benefit of adaptation can often be removed simply by strengthening the frozen baseline. A benchmark that only reflects this regime will not reveal whether certified adaptive retention is possible.

That led me through a lot of failed ideas.

  • High-frequency gratings above the encoder’s sampling limit failed because the frozen, strided encoder solved the problem through aliasing: high-frequency shapes became visible moiré patterns.
  • Dot positions within a patch blinded the encoder, but also blinded a from-scratch model, making the task unusable.
  • Relative phases of multiple gratings produced signals that survived only under specific nonlinearities and did not generate a clean separation between a failing frozen encoder and a winning adapter.
  • Texture pairings, number of marks, and spatial layouts all looked plausible and turned out to be linearly separable in the frozen feature space.

The frozen encoder is surprisingly hard to trick in any way that still leaves a realistic path for a trainable model.

Designing the task bottom-up

When all task-relevant information is eliminated, adaptation becomes necessary. Let’s give each group and class a label: coarse stripes for the group and fine stripes for the class. The encoder gets only a blurred copy, with only the group label visible. The highest achievable accuracy is therefore fixed by the grouping, which makes the representation itself testable.

The grouping can be designed around how much of the label the model is allowed to see, so that the ceiling is related to the group size. Each group consists of fine classes that differ only in the fine stripes. The anti-aliased channel feeding the frozen encoder removes those fine frequencies entirely, so all classes within a group collapse to the same representation.

Any readout on top of that representation is therefore limited to random guessing within the group. Every method, whether adaptive or not, receives the same backbone and the same bandwidth restriction. Any differences downstream can therefore be attributed to the mechanisms being evaluated rather than to architectural shortcuts.

Two classes in one group carry identical coarse (group) stripes but different fine (class) stripes. The anti-aliased channel removes the fine frequencies, so both classes collapse to the same blurred representation, capping any readout at within-group chance. Class A (input) Class B (input) anti-aliased channel (blur) What the encoder sees A = B within-group: chance
Task construction and the representation bottleneck. Two classes in one group share the coarse (group) stripes but differ only in the fine (class) stripes. The anti-aliased channel feeding the frozen encoder strips the fine frequencies, so both classes collapse to one representation and any readout is capped at within-group chance. Only a representation that is allowed to change can recover the fine labels.

Measuring the accuracy–retention tradeoff

I evaluate the task with CARG, the certified adaptive-retention gap: the best accuracy of any method minus the best accuracy of any method whose retention certificate actually holds.

CARG  =  max⁡m acc(m)  −  max⁡m : cert(m) acc(m)\mathrm{CARG} \;=\; \max_{m}\ \mathrm{acc}(m)\;-\;\max_{m\,:\,\mathrm{cert}(m)}\ \mathrm{acc}(m)

This measures the tradeoff between raw accuracy and adaptation. Best accuracy of anyone is the unconstrained optimum: whatever method reaches the highest score on the stream, regardless of whether it can justify its retention claims. Best accuracy of anyone where retention holds is the maximum score among methods whose retention bound actually covers their behaviour on held-out checks.

A high CARG means that enforcing a correct retention certificate eliminates most of the achievable accuracy. That is precisely the regime I want to expose. The point is not simply to produce a harder benchmark; it is to create one where adapting the representation while retaining previous capabilities is measurable and cannot be bypassed by a stronger frozen feature extractor.

What happens to existing continual-learning methods?

Parameter-space methods — Elastic Weight Consolidation (EWC), Synaptic Intelligence (SI), Learning without Forgetting (LwF), Averaged Gradient Episodic Memory (A-GEM), and Gradient Projection Memory (GPM) — collapse under an adapting representation, while data-replay methods perform well.

Methods that regularise or project parameter updates are built around preserving functions in a representation that is assumed to remain sufficiently stable. Once the representation has to move far enough to resolve the fine labels, those parameter-space constraints can no longer keep the old functions intact. Replay and class-incremental schemes that retain explicit examples or proxies close to them continue to track the unconstrained optimum more closely, because they can re-fit the old tasks in the new coordinates.

Can replay solve the scaling problem?

Exemplar-free settings are nevertheless necessary for scaling this approach, where storage, bandwidth, and privacy costs become prohibitive at deployment scale. Common approaches store the old model parameters or old raw data, and a NeurIPS’25 paper demonstrated decent results with only 0.3% of the data size [1]. A 2026 ICLR paper argues that “exemplar memory is no longer the limiting factor” and that simple replay baselines can outperform SOTA methods “at a fraction of the GPU cost” [2].

That rules out the most straightforward way of reconciling adaptation and retention. The burden then shifts to methods that summarise history into a compressed representation, embed data, or find adaptive representations of the data and tasks themselves. That is exactly where certified retention under changing representations becomes interesting: the challenge is remembering what mattered about those examples while the coordinate system used to represent them is itself changing.

Closing remarks

This benchmark is, unfortunately, not applicable to real images or audio spectrograms, even blurred. Its construction deliberately isolates a regime in which a frozen encoder is provably insufficient and an adaptive one is necessary. It does this by altering what the encoder is allowed to see rather than relying on natural variation in datasets.

That is also the point. It is not intended to be a realistic vision dataset; it is intended to establish a clean experimental regime in which the usual frozen-feature advantage disappears and adaptive retention can actually be tested. A certified adaptive-retention regime would greatly advance continual-learning research by making adaptation mandatory and retention measurable.

The broader lesson is that fair evaluation of continual learning cannot stop at asking “how accurate is the model?” Evaluation also needs to ask “how much of that accuracy survives under a valid retention certificate?” and “what happens when the representation itself has to change?”

Once there are sufficiently strong adaptive models, the models that gain the most from changing their representations are also the models for which simple retention certificates become hardest to verify.

Sources

  1. NeurIPS 2025. Exemplar-free continual learning at roughly 0.3% of the data footprint. proceedings.neurips.cc
  2. ICLR 2026. “Exemplar memory is no longer the limiting factor”: simple replay baselines that outperform SOTA at a fraction of the GPU cost. iclr.cc/virtual/2026/poster/10008192

What should I call you?

Choose a display name for your comments. No email or account signup.

Your commenting identity

Use at least 12 characters. You’ll need this passphrase to restore the file.