← Research

Locas: Your Models are Principled Initializers of Locally-Supported Parametric Memories

Authors

Sidi Lu, Zhenwen Liang, Dongyang Ma, Yan Wang, Haitao Mi, and Dong Yu

Context

This work grew out of a close collaboration with Zhenwen Liang during my time at Tencent AI Lab Seattle.

Overview

Locas introduces a form of locally-supported parametric memory aligned with the feed-forward blocks of modern Transformers. A memory module can be attached to a model for efficient continual learning and later either offloaded or merged into the model’s existing parameters.

We study two variants: a conventional two-layer MLP with a clearer theoretical account, and a GLU-FFN design that matches the structure of current language models. Their initialization reuses information already present in the model—parameters, activations, and gradients—to improve convergence, generalization, and retention of prior knowledge.

Results

The method is evaluated on PG-19 whole-book language modeling and LoCoMo long-context dialogue question answering, with MMLU used to measure loss of general capability after memorization. In its most parameter-efficient setting, Locas-GLU uses 0.02% additional parameters while retaining information from past context with a much shorter active context window.

Paper

arXiv

Status

Tencent AI Lab Technical Report, 2026.

Citation

Sidi Lu, Zhenwen Liang, Dongyang Ma, Yan Wang, Haitao Mi, and Dong Yu. “Locas: Your Models are Principled Initializers of Locally-Supported Parametric Memories.” arXiv:2602.05085, 2026.