๐Ÿ›๏ธ QU-SSM-60M-MoE: Continuous Quasi-Unitary State Space Model with Sparse Mixture-of-Experts

๐Ÿ“„ Official Research Paper (PDF) & Open Verification

Title: "Gated Quasi-Unitary Lie-Algebra Recurrent State Space Models"
Author: Prannessh K.V.A. (@prannesshkva)

๐Ÿ“ฅ Download Full Research Paper (PDF) | ๐Ÿ›๏ธ Zenodo DOI: 10.5281/zenodo.22283431 | ๐ŸŽฎ Live Interactive Space

Official Paper Zenodo DOI LinkedIn Profile License: CC BY-NC-ND 4.0 Model: qu_ssm-60Moe Flagship: qu_ssm-130Moe

QU-SSM-60M-MoE is the mid-tier foundation model of the QU-SSM family designed and invented by Prannessh K.V.A. (Sole Architect & Inventor). It combines continuous Lie-group unitary recurrence on SO(2) โ‰… U(1) with 4 SwiGLU Mixture-of-Experts (MoE) and Top-2 routing (44.64M active parameters per token).


๐Ÿงฌ Base Model Lineage & Technical Notes

  • Core Architecture: Continuous Quasi-Unitary State Space Model (SO(2) phase rotations) coupled with 4 SwiGLU Mixture-of-Experts and Top-2 routing.
  • Tokenizer Lineage: Standard GPT-2 Byte-Pair Encoding (BPE) vocabulary (50,257 tokens).
  • Pre-training & Calibration: Pre-trained on roneneldan/TinyStories (~20M+ tokens) demonstrating sub-millisecond step latency and exact norm preservation.
  • Parameter Footprint: 64.30M Total Parameters, 44.64M Active Parameters per token.
  • Inference Efficiency: Constant O(1) inference state RAM (0.19 MB) regardless of sequence length.
  • Official Research Contact: LinkedIn โ€” Prannesh K. V. A.

๐Ÿ” What is QU-SSM?

QU-SSM is a linear-time continuous sequence engine that replaces the monotonic dissipative decay of classical state space models with non-dissipative SO(2) unitary phase rotations (โ€–R(ฮธ)โ€–โ‚‚ โ‰ก 1.00000), delivering strictly constant O(1) inference memory and sub-millisecond step latency.


๐Ÿ“Š Architecture Specifications

Specification Value
Model Name QU-SSM-60M-MoE
Sole Architect & Inventor Prannessh K.V.A.
Total Parameters 64.30M
Active Parameters / Token 44.64M (Top-2 Sparse MoE)
Hidden Dimension (D) 384
Layers 6
SSM State Dimension 8
Expert Count 4 SwiGLU Experts
Vocabulary 50,257 (GPT-2 BPE)
Inference State RAM 0.19 MB (Constant O(1))

๐Ÿ”’ Intellectual Property & Citation


๐Ÿ”— Related QU-SSM Hub Repositories

Downloads last month
2,676
Safetensors
Model size
64.3M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Dataset used to train Prannesshkva/QU-SSM-60M-MoE

Space using Prannesshkva/QU-SSM-60M-MoE 1