Train and serve MoE models that do not fit in VRAM: fused 4-bit experts, QLoRA, CPU/NVMe offload, and fast inference on consumer NVIDIA GPUs.
cuda transformers pytorch triton moe lora quantization nvme fine-tuning mixture-of-experts int4 fp8 llm-training llm-inference qlora bitsandbytes nf4 consumer-gpu mxfp4 gpu-offloading
-
Updated
Sep 7, 2026 - Python