Sencelium 30M
Sencelium is an independent research project exploring a non-Transformer language model architecture: gated-delta-rule blocks with a content-blind "nucleus" channel, a memory pool that keeps spawning, merging and reorganizing itself during training and at inference, live self-monitoring signals, and an ESN residual path.
This checkpoint (29.638M params, 15 blocks) is the canonical 30M-scale run. Trained on FineWebEdu, evaluated on WikiText-103.
Code, architecture write-up, full benchmarks: github.com/vitaliihalak/sencelium Architecture diagrams: sencelium.com/architecture
Honest framing
This is not a PPL win over a matched Transformer at short context β it isn't one, at this scale. What's real: the memory pool is measurably load-bearing (ablating it hurts), and perplexity scales differently with context length than a Transformer's does.
Results
| Sencelium 30M | Transformer 30M | |
|---|---|---|
| params | 29.638M | 29.403M |
| val_ppl (WikiText-103) | 272.46 | 229.49 |
| val_ppl (context 4096) | 238.19 | 640.93 |
At short context the Transformer baseline is ahead. Past ~1024 tokens that reverses β Sencelium's perplexity keeps falling as context grows, the Transformer's rises sharply past the length it was trained on, even with its RoPE cache extended.
Key findings
Full methodology and all numbers: benchmarks/results.md Β· sencelium.com/results.
Memory structures itself during training, unsupervised. Starting from zero slots, it spawns new ones and merges similar ones into categories as training goes β no labels, no fixed taxonomy. Usage stays concentrated in a small fraction of slots and near-duplicate content stays low, at both scales.
| Sencelium 30M (this checkpoint) | Sencelium 65M | |
|---|---|---|
| categories formed | 317 | 500 |
| spawn β merge β reassignments | 5,354 β 1,659 β 8,492 | 5,137 β 778 β 6,850 |
| slots holding 90% of usage | 76 (2.1%) | 154 (3.5%) |
| near-dup rate | 0.53% | 1.54% |
Memory is load-bearing, and more so at scale. Forcing the pool empty at
eval time (no_mem) costs real perplexity on the same checkpoint, same
data β and the cost grows with model size instead of shrinking.
| Sencelium 30M (this checkpoint) | Sencelium 65M | |
|---|---|---|
| val_ppl / no_mem | 272.46 / 276.68 | 129.39 / 154.36 |
| mem_delta | 4.22 (1.5%) | 24.97 (19.3%) |
Carrying state across a document helps, for free. No retraining, no extra parameters β just not resetting recurrent state between a document's own windows at eval time.
| Sencelium 30M (this checkpoint) | Sencelium 65M | |
|---|---|---|
| carry_gain | 9.6% | 14.4% |
The ESN residual and the inner voice loop are both doing real work.
esn_alpha's magnitude grows with scale instead of decaying toward zero;
the voice loop's per-round updates shrink but stay nonzero through the
third round β iterative refinement, not a no-op pass.
| Sencelium 30M (this checkpoint) | Sencelium 65M | |
|---|---|---|
| esn_alpha | β1.74 | β3.61 |
| voice-loop deltas (round 1β3) | 0.299 β 0.140 β 0.110 | 0.467 β 0.251 β 0.126 |
A fair ablation of the ESN branch at 65M β replacing its output with the dataset mean, or with another document's ESN output, rather than zeroing it (zeroing mostly measures distribution shock, not content value) β shows a real, measurable content-specific contribution:
| ESN ablation (65M) | val_ppl |
|---|---|
| trained (real) | 124.30 |
| output β dataset-mean | 166.29 |
| output β another document's ESN output | 194.27 |
How to use this checkpoint
This is a non-standard architecture β not AutoModel-compatible. Use the
training/eval code from the GitHub repo directly:
git clone https://github.com/vitaliihalak/sencelium
cd sencelium
pip install -r requirements.txt
python benchmarks/eval.py ppl --script scripts/train_sencelium.py \
--config configs/30M.yaml --checkpoint path/to/sencelium-30M.ckpt
python benchmarks/eval.py features --script scripts/train_sencelium.py \
--config configs/30M.yaml --checkpoint path/to/sencelium-30M.ckpt
License
Apache 2.0 (code and weights). See LICENSE.
Citation
See CITATION.cff.
- Downloads last month
- 11