Quick Answer: A 34M logit-correction module sitting atop a fully frozen Gemma 4 E2B model corrects 53.3% of domain errors with zero capability degradation on standard benchmarks. A LoRA baseline gets 83.3% correction but suffers 30–75% capability loss. The frozen-base approach preserves what the base model knows.

The Core Question

When a language model makes domain errors, the standard fix is fine-tuning — which often causes catastrophic forgetting. We asked a deliberately narrow question: if you freeze the base model completely and only train a small correction module on top, how far can you get and what capabilities do you keep?

CRN v2 Architecture

CRN v2 (Cognitive Resonance Network v2) is a ~34M trainable parameter logit-level correction module — 0.73% of the 4.65B text module of Gemma 4 E2B. The base model is never updated. Training: supervised fine-tuning on 83,400 error-correction pairs, followed by reference-free DPO with a KL preservation term (lambda=0.1).

The KL Term is Critical

An ablation shows that lowering lambda from 0.1 to 0.01 degrades correction from 53.3% to 35.0%. The KL penalty is what preserves general capabilities during error correction training.

The Capability-Correction Tradeoff

  • CRN v2 (34M): 53.3% correction · MMLU 62.5% (unchanged) · BoolQ 72% (unchanged) ✅
  • LoRA (6.6M): 83.3% correction · MMLU 62.5%→32% · car-wash 75%→0% ❌
  • DPO at depth 7: Destroyed capabilities entirely (MMLU 13%) ❌

What We Couldn't Break Past

We tried 5 different approaches to break past ~53% correction: more rank, longer training (5k SFT + 2k DPO), multi-depth correction, and deep injection. None exceeded the rank-128 logit result. We published this as the finding rather than burying it.

Links

ArXiv: 2609.16145 · GitHub: eulogik/prajna · Weights: eulogik/Prajna-CRNv2

Want more technical insights?

Follow us on X for real-time updates on KARN and our other ventures.