Local Support Learning

1Tel Aviv University   2MIT CSAIL
†Work done while visiting the Spoken Language Systems group at MIT CSAIL

Modern neural networks learn in phases: large-scale pretraining followed by smaller finetuning phases. Catastrophic forgetting hurts this process, as each new phase can overwrite capabilities acquired in previous ones. We propose Local Support Learning (LSL), which keeps weight updates local to their training data, minimizing interference with prior capabilities and overcoming forgetting in LLMs of up to 7B parameters.

FIG. 1Avoiding catastrophic forgetting with Local Support Learning. Left: A gating function enables the weight adapter only on input activations from its training distribution. Interestingly, conventional MLP classifiers are not well suited for this task. Instead, we propose a GMM-based gate that tends to stay closed on prior data it has never encountered. Right: a 1D illustration of the GMM-based gate. Φpos (orange) captures the finetuning data, while the wider Φneg (blue) dominates elsewhere, so data outside the current distribution is routed only to the pretrained weights.

Abstract

We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.

How does it work?

In each new learning phase, LSL pairs a standard weight adapter (e.g. LoRA) and a GMM-based gate. The weight adapter is trained with gradient descent as usual. The gate fits two Gaussian mixtures: Φpos on the current phase’s input activations, and a wider Φneg, fit once on about one million generic pretraining tokens. A token passes through the adapter only where it is more likely under the current phase:

Φpos(x) − Φneg(x) > 0

For any other token, the weight matrix behaves exactly as before the current training phase.

Results

We finetune Qwen2.5-7B-Instruct on diverse tasks: chemistry, translation to a low-resource language (English→Igbo) and cybersecurity, and measure retention of pretraining capabilities such as math, coding, and instruction following (Fig. 2).
LSL achieves almost optimal retention while also exhibiting strong learning abilities on the new task.

FIG. 2Post-training with LSL. We evaluate forgetting on a Qwen2.5-7B-Instruct model for three diverse downstream tasks. The x-axis measures the new task’s performance and the y-axis measures capability retention, the average score across three pretraining benchmarks. The upper-right corner reflects stronger performance. LSL learns the downstream task while achieving near-optimal retention across all settings.

Next, we finetune on all three tasks sequentially (Fig. 3). LSL retains both its pretrained capabilities and the skills gained in earlier finetuning phases, while the baseline consistently forgets.

FIG. 3Post-training with multiple phases. Each plot tracks the evaluation score of one task throughout training; the phase whose training data matches the evaluated task is marked with a dashed line. The training order is (1) Igbo translation, (2) Chemistry, (3) Cybersecurity. In the top-left plot, LSL retains pretrained capabilities throughout, while the baseline drops by 44% already after one phase. On the three finetuning benchmarks, LSL retains its gains after each task’s phase ends, whereas the baseline’s performance declines steadily once the peak is reached.

Additional results

Mechanistic motivation

Let v = Wx be the input–output mapping of a single weight matrix. We study how a new gradient update affects this mapping, given a new training sample x(s). Consider a simple gradient-based update ΔW:

ΔW = −α · ∂𝓛∂v · x(s)⊤

where α is the learning rate. ΔW is itself a matrix: applied to any input x, it changes the output by Δv = ΔW x.

Notice that the update is global: ΔW affects any input x that is not orthogonal to it, including inputs from previous learning phases. This is demonstrated below during the “Finetune with GD” phase, where the classification of prior classes (blue) changes when training on new data (red). To prevent this, we constrain the update to act only on the support of the current training distribution, making it local to that region.

LSL realizes this with a gate that opens only on the current phase’s data. As shown in the “Finetune with LSL” phase (Fig. 4), the new classes are learned while the classification of prior classes stays intact.

Reproducing forgetting in a two-phase training run

Pretraining

Class 0Class 1Class 2Class 3

Finetuning

Class 4Class 5

Accuracy

Pretraining classes (0–3)Finetuning classes (4–5)

0 / 1200 steps

FIG. 4Interactive toy example. A linear classifier v = Wx, trained by gradient descent. Background colors show the classifier’s decision regions, and their borders are the decision boundaries. Pretrain on the blue classes (0–3), then finetune only on the red ones (4–5). Finetuning with GD redraws the boundaries everywhere and forgets old classes; with LSL, the update is applied to a region that is local to the current training distribution, avoiding interference with previously learned data.

Theoretical motivation

As before, we consider a single weight matrix. Let A* denote the support of the current data. We frame retention as a minimax problem: since nothing is known about prior data, an adversarial pretraining sample could lie anywhere outside A*. Under this framing, the ideal gate opens exactly on A*.1 A learned gate instead opens on some region Â. We show that against the worst-case test distribution, its error decomposes into two terms: the deficit, the part of A* the gate rejects, which limits learning, and the excess, the region outside A* it accepts, which causes forgetting.

A* Â (gate opens) deficit excess current-phase data prior-phase data
Deficit: the part of the current data’s support on which the gate stays closed. The gate is trained on this data, so its training objective optimizes this term directly, and we do not discuss it further.
Excess: all points outside the current data’s support on which the gate erroneously stays open. This term quantifies forgetting, because inputs from previous phases that fall here interact with the update and may therefore be altered. Crucially, since this region is off-data by definition, the gate’s training objective is blind to it, so this error component cannot be directly optimized.
No error (grey): the gate correctly opens inside A* and correctly stays closed outside it.

While no objective on the current data alone can act on the excess, we show that other factors, such as the hypothesis class, can control it under explicit conditions. Specifically, a gate based on a Gaussian Mixture Model (GMM) has a likelihood that decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We study several hypothesis families, and find that in general, consistent density estimators can recover optimal gates, at costs we analyze.

1 If prior-phase data also lies inside A*, that region is being optimized by the current phase, so it is valid for the gate to be open there. ↩

For full details and more experiments, see the paper!

BibTeX

@article{benkish2026lsl,
  title   = {Local Support Learning},
  author  = {Ben-Kish, Assaf and Kumar, Akarsh and
             Glass, James and Giryes, Raja},
  journal = {arXiv preprint},
  year    = {2026}
}