qpeft: quantization-aware training (QAT) and PEFT for LLMs in PyTorch. EfficientQAT, QA-LoRA and PEQA for Hugging Face models at 2/3/4 bits: train scales, zero-points or LoRA adapters, then merge into a low-bit integer model (GPTQ layout) that is tested to equal the trained one. Pure-torch and torchao backends.
machine-learning deep-learning transformers pytorch lora quantization model-compression qat fine-tuning peft huggingface low-bit quantization-aware-training large-language-models llm gptq llm-quantization torchao qa-lora efficientqat
-
Updated
Sep 29, 2026 - Python