Reactive LLM inference.
-
Updated
Jul 7, 2026 - Python
Reactive LLM inference.
CUDA-native residual-frame Delta memory for causal language models: full-matrix prefix geometry, rank-one recurrent updates, exact local similarity, Hugging Face integration, and CUDA Graph training.
Hyper-Flux Projection (HFP): A physics-inspired causal LM architecture exploring constant O(1) memory size and cubic-plateau retention laws.
NOVA: Neural Organic Vivid Architecture — a sub-quadratic sequence architecture with recurrent multi-timescale memory (FLUX), antisymmetric attention (IRIS), and role-separated representations (RSLA)
To associate your repository with the recurrent-memory topic, visit your repo's landing page and select "manage topics."