1Tel Aviv University 2MIT CSAIL
†Work done while visiting the Spoken Language Systems group at MIT CSAIL
Modern neural networks learn in phases: large-scale pretraining followed by smaller finetuning phases. Catastrophic forgetting hurts this process, as each new phase can overwrite capabilities acquired in previous ones. We propose Local Support Learning (LSL), which keeps weight updates local to their training data, minimizing interference with prior capabilities and overcoming forgetting in LLMs of up to 7B parameters.
We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.
In each new learning phase, LSL pairs a standard weight adapter (e.g. LoRA) and a GMM-based gate. The weight adapter is trained with gradient descent as usual. The gate fits two Gaussian mixtures: Φpos on the current phase’s input activations, and a wider Φneg, fit once on about one million generic pretraining tokens. A token passes through the adapter only where it is more likely under the current phase:
For any other token, the weight matrix behaves exactly as before the current training phase.
We finetune Qwen2.5-7B-Instruct on diverse tasks: chemistry, translation to a low-resource language (English→Igbo) and cybersecurity, and measure retention of pretraining capabilities such as math, coding, and instruction following (Fig. 2).
LSL achieves almost optimal retention while also exhibiting strong learning abilities on the new task.
Next, we finetune on all three tasks sequentially (Fig. 3). LSL retains both its pretrained capabilities and the skills gained in earlier finetuning phases, while the baseline consistently forgets.
Let v = Wx be the input–output mapping of a single weight matrix. We study how a new gradient update affects this mapping, given a new training sample x(s). Consider a simple gradient-based update ΔW:
where α is the learning rate. ΔW is itself a matrix: applied to any input x, it changes the output by Δv = ΔW x.
Notice that the update is global: ΔW affects any input x that is not orthogonal to it, including inputs from previous learning phases. This is demonstrated below during the “Finetune with GD” phase, where the classification of prior classes (blue) changes when training on new data (red). To prevent this, we constrain the update to act only on the support of the current training distribution, making it local to that region.
LSL realizes this with a gate that opens only on the current phase’s data. As shown in the “Finetune with LSL” phase (Fig. 4), the new classes are learned while the classification of prior classes stays intact.
Reproducing forgetting in a two-phase training run
Pretraining
Class 0Class 1Class 2Class 3Finetuning
Class 4Class 5Data
Accuracy
0 / 1200 steps
As before, we consider a single weight matrix. Let A* denote the support of the current data. We frame retention as a minimax problem: since nothing is known about prior data, an adversarial pretraining sample could lie anywhere outside A*. Under this framing, the ideal gate opens exactly on A*.1 A learned gate instead opens on some region Â. We show that against the worst-case test distribution, its error decomposes into two terms: the deficit, the part of A* the gate rejects, which limits learning, and the excess, the region outside A* it accepts, which causes forgetting.
While no objective on the current data alone can act on the excess, we show that other factors, such as the hypothesis class, can control it under explicit conditions. Specifically, a gate based on a Gaussian Mixture Model (GMM) has a likelihood that decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We study several hypothesis families, and find that in general, consistent density estimators can recover optimal gates, at costs we analyze.
1 If prior-phase data also lies inside A*, that region is being optimized by the current phase, so it is valid for the gate to be open there. ↩
For full details and more experiments, see the paper!
@article{benkish2026lsl,
title = {Local Support Learning},
author = {Ben-Kish, Assaf and Kumar, Akarsh and
Glass, James and Giryes, Raja},
journal = {arXiv preprint},
year = {2026}
}