GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
-
Updated
Sep 6, 2026 - Python
GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.
Agent-assisted and full-agent reproducibility package for MLSys 2026 FlashInfer AI Kernel Generation Contest submissions: kernels, agent workflows, skills, configs, writeup, benchmark artifacts, and full optimization records.
NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.
Reproducible SGLang recipe + public prebuilt image (ghcr.io) for DeepSeek-V4-Flash-0731 on 4x RTX PRO 6000 Blackwell (SM120): TP4/DP4/EP4, 1M ctx, benchmarks, and the DSPARK draft-depth corruption boundary
180-226 tok/s single-stream decode for Qwen3.8-Flash-Next NVFP4 on one RTX PRO 6000 Blackwell, at full 262K context with unchanged quantization. Config, benchmark harness, and the FlashInfer autotune correctness bug that silently corrupts output.
Production runbook for Qwen3.5-122B hybrid INT4+FP8 on NVIDIA DGX Spark GB10 — optimization stack, PD firmware wedge diagnosis, bench results
vLLM + FlashInfer source integration for GLM-5.3-Flash on SM120 / RTX PRO 6000 Blackwell
🚀 Accelerate attention mechanisms with FlashMLA, featuring optimized kernels for DeepSeek models, enhancing performance through sparse and dense attention.
Run GLM-5.3 Flash NVFP4 on 2x NVIDIA DGX Spark with vLLM TP2, FlashInfer sparse MLA, FP8 KV and MTP3
🚀 从零手写的 Qwen3 高性能推理引擎 —— 1,500 行纯 Python 实现 Paged KV Cache · 连续批处理 · FlashAttention-2 · CUDA Graph,零框架依赖(不基于 vLLM/SGLang),单卡 256 并发 1,200+ tok/s
DeepSeek-V4-Flash-DSpark on 2× DGX/ASUS Spark (GB10, SM120) using the stock spark-vllm container - recipes, a faster SM120 topk fix, and one-command run scripts.
Single-GPU LLM decode research prototype: paged KV cache, Triton attention, CUDA append, scheduling, shared prefixes, and multi-layer transactions.
Community wishlist for reproducible kernel definitions, workloads, and validated optimized solutions.
Production-grade fine-tuning & LoRA toolkit for Chatterbox-Flash zero-shot TTS models. Combines parallel block diffusion and FlashInfer acceleration with smart placeholder vocabulary extension supporting languages. Features offline feature preprocessing, Silero VAD silence trimming, zero-padding leakage prevention, and fast voice cloning for custom
⚡ Optimize attention mechanisms with FlashMLA, a library of advanced sparse and dense kernels for DeepSeek models, improving performance and efficiency.
vLLM plugin: per-layer attention backend selection for Gemma 4 heterogeneous head dims (drag-and-drop backport of PR #38891), with a text-only fast path + MTP-safe head=512 Triton pin.
Single-RTX-5090 serving stack for Qwen3.8-27B: 262K context, native NVFP4 KV, MTP, GDN-aware caching, and vision.
To associate your repository with the flashinfer topic, visit your repo's landing page and select "manage topics."