MMLA Stage B — learning from feedback with LARC

This package contains three small, trained learning modules for a frozen language model. Each module changes how the model chooses among four supplied programs after seeing their execution results. The model's backbone stays fixed.

The module is called LARC: Low-Rank Adaptive Residual Connection. Earlier code uses LRARC; those filenames are kept for compatibility. This package contains the Stage-B static-objective models, one for each of three training seeds.

中文:这里是 MMLA 的一组实验权重。模型先对四个给定程序打分,再根据程序执行反馈更新一个小型残差模块,改变下一次选择。目录包含三个训练种子的模块和必需的语义适配器;使用时还需下载下面指定版本的基础模型。

What is included?

Component What it does Included?
MiniCPM5-1B-SFT Provides the frozen language-model computation Download separately
omega.safetensors The original semantic adapter; keep it enabled Yes
static-2026092811.safetensors Trained LARC starting point, seed 2026092811 Yes
static-2026092812.safetensors Same training recipe, seed 2026092812 Yes
static-2026092813.safetensors Same training recipe, seed 2026092813 Yes

All weights are in weights/. These are components for a research implementation, not a ready-to-chat model or an automatic generate() pipeline. The three seeds are retained together; this is not a best-run selection. The later CI-trained version-926 models are separate and are not in this package.

How learning works

LARC changes the input hidden vectors just before they enter the frozen mixer:

updated = hidden + (hidden @ A.T) @ B.T

A reads four learned combinations of the hidden coordinates, and B writes a correction back into the full vector. There are 12,288 learned values: A has shape [4, 1536] and B has shape [1536, 4], both FP32. The residual has rank 4 and no extra scaling factor.

The saved pair of matrices is the trained starting point, called rho. For each episode, make a working copy called Phi. Feedback updates Phi with two SGD steps at learning rate 0.1. Resetting means copying the trained rho again, including its trained B. Before the original training, A was random and B was zero; resetting a trained model does not return to that zero residual.

LARC shares the low-rank factorization used by LoRA. This experiment places it at the mixer input and gives its factors an episode-specific update/reset lifetime. The separate omega is a frozen rank-16 LoRA adapter on query/value projections in layers 16–23, with alpha 32. It supplies the common semantic substrate and stays enabled while Phi learns.

Load the weights

Use Python, PyTorch and safetensors. Backbone loading also needs Transformers and its tokenizer dependencies; the recorded versions are in environment.json.

Run from this package directory:

from load_weights import verify_bundle, load_rho, begin_episode

verify_bundle()
rho = load_rho(seed=2026092811, device="cpu")
phi = begin_episode(rho)  # independent working copy, gradients enabled
# Call begin_episode(rho) again to reset.

Download openbmb/MiniCPM5-1B-SFT at revision a60b37f1fc409c54e1e337b0723aaac6f92dfec0. Then load the frozen BF16 substrate:

from load_weights import load_backbone, apply_input_residual

model, tokenizer = load_backbone("/path/to/exact/MiniCPM5-1B-SFT", device="cuda:0")
phi = begin_episode(load_rho(seed=2026092811, device="cuda:0"))
# In your scoring loop, apply_input_residual(embeddings, phi)
# before the mixer, preserving attention masks and position IDs.

load_backbone checks the original model identity and installs omega. Your scoring loop must apply the input residual explicitly. episodic_core.py provides the original fast-state update and reset operations. Gradients must pass through the frozen mixer to reach Phi; cache only representations before Phi, and give each episode its own writable state.

lineage.json identifies the original checkpoints and exact base files. manifest.json lists package hashes. Exported tensors were reloaded and compared exactly with the originals. Full optimizer/RNG continuation states remain in the original training checkpoints.

What the experiment found

The study trained two objectives at the same LARC location: 2 objectives × 3 paired seeds × 256 outer updates, with two episodes per batch. Static training evaluates the query at rho. Adapted training evaluates it after two support updates, using a first-order identity-Jacobian estimate. Only the selected static models are included here.

The development set has 16 held-out parameter groups and 64 episodes from familiar arithmetic/list program families. Each episode offers four complete programs. The primary metric is expected execution error on new query inputs after feedback adaptation; it weights each candidate's execution error by the probability the model assigns to it. This can improve even when the top-ranked program stays the same.

Training objective Error after adaptation, mean ± seed SD Improvement from real feedback Permutation-adjusted improvement
Static — included 29.30% ± 6.59 pp 24.65 pp 28.36 pp
Adapted — comparison 24.61% ± 15.66 pp 36.65 pp 35.77 pp

pp means percentage points. Feedback improvement compares the adapted state with its own trained reset. The permutation adjustment checks whether the correct pairing of programs and feedback matters. The static objective was selected by the fixed worst-seed feedback-stability rule. The adapted objective has lower mean error, but its advantage reverses for one seed. Training seeds are repeated runs, not additional independent task groups.

A direct rule that selects the program with the lowest support error reaches 0.78125% query error, outperforming both model groups. The result demonstrates feedback adaptation in this finite task. It does not establish the full PDSA system or strict recursive self-improvement. The reserved final audit remains unopened. See evaluation_summary.json for per-seed results and statistical contrasts.

Paper and citation

Junyi Zou and Avrova Donz. LARC: Low-Rank Adaptive Residual Connections for Learning in Frozen Models, arXiv:2609.40063, 2026.

This package contains the three static-objective program-selection initializations studied in the LARC technical report. LARC implements the numerical policy carrier of MMLA, a mixer-agnostic architecture; this implementation uses a Transformer mixer.

@misc{zou2026larclowrankadaptiveresidual,
      title={LARC: Low-Rank Adaptive Residual Connections for Learning in Frozen Models}, 
      author={Junyi Zou and Avrova Donz},
      year={2026},
      eprint={2609.40063},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.40063}, 
}

License

This package uses AGPL-3.0; see LICENSE and NOTICE. The separately downloaded MiniCPM base retains its upstream license and terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MMLA-ORG/MMLA-LRARC4

Adapter
(4)
this model

Paper for MMLA-ORG/MMLA-LRARC4