Skip to content
LyoAIPublic

About

[ACL'26] The official implementation for "Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models"

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

16 Commits

Folders and files

Repository files navigation

Astra

Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models

Paper PEFT Integration PR License

A task-aware LoRA initialization that adapts large language models in the under-utilized tail eigenspace of output activations.


πŸ“° News

  • 2026-09-29 β€” Astra was merged into huggingface/peft as init_lora_weights="astra".
  • 2026-09-29 β€” Astra was integrated into Hugging Face PEFT through PR #3725.
  • 2026-02 β€” The Astra paper was released on arXiv.

✨ Overview

Astra is a data-driven LoRA initialization method for parameter-efficient fine-tuning. Instead of initializing an adapter only from pretrained weights, Astra first characterizes the activation space induced by the downstream task.

For each targeted linear layer, Astra:

  1. Collects output activations on a small task-specific calibration set;
  2. Estimates the output-activation covariance matrix;
  3. Performs one eigendecomposition;
  4. Initializes the adapter in the under-utilized tail eigenspace;
  5. Residualizes the frozen base weight so the model output is preserved at initialization.

This directs adaptation toward directions that are useful for the downstream task while avoiding the dominant pretrained representation.

Highlights

Property Description
Task-aware Uses downstream activations rather than only pretrained weights.
Activation-space Constructs adapters from output-activation covariance.
Lightweight preprocessing One eigendecomposition per target layer; no weight SVD or covariance inverse.
Low-rank efficient Strong performance at small ranks and reduced parameter budgets.
PEFT-native Officially integrated as init_lora_weights="astra".

Astra was evaluated across 16 NLU and NLG benchmarks covering mathematical reasoning, code generation, commonsense reasoning, and GLUE-style tasks. It consistently outperforms existing PEFT baselines at equal or lower rank and surpasses full fine-tuning on several tasks.

Astra overview
Astra results

πŸš€ Quick Start

Use Astra from Hugging Face PEFT

Install the PEFT revision that contains Astra:

pip install "peft @ git+https://github.com/huggingface/peft.git@73f9a1a901b03adf63ed6d042c22ac36033e8a12"

Astra is used as a LoRA initialization method:

import torch
from peft import LoraConfig, get_peft_model
from peft.tuners.lora import AstraConfig, preprocess_astra
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-2-7b-hf",
    dtype=torch.bfloat16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-2-7b-hf")

# Replace this with a loader over a small calibration split of your downstream task.
calibration_dataloader = ...

def run_model():
    model.eval()
    with torch.no_grad():
        for batch in calibration_dataloader:
            batch = {key: value.to(model.device) for key, value in batch.items()}
            model(**batch)

lora_config = LoraConfig(
    init_lora_weights="astra",
    r=128,
    lora_alpha=128,
    target_modules=[
        "q_proj", "k_proj", "v_proj", "o_proj",
        "gate_proj", "up_proj", "down_proj",
    ],
    astra_config=AstraConfig(),
)

# Preprocessing must happen before get_peft_model.
preprocess_astra(model, lora_config, run_model=run_model)
peft_model = get_peft_model(model, lora_config)

target_modules must be specified explicitly because preprocess_astra() runs before PEFT infers model-specific defaults.

Run the research repository

Clone this repository and install the full training environment:

git clone https://github.com/LyoAI/Astra.git
cd Astra
uv sync

The equivalent conda installation is:

conda create -n astra python=3.10
conda activate astra
pip install -r requirements.txt

πŸ’Ύ Saving and Conversion

Astra preprocessing creates a residual base model and saves the untrained adapter to astra_init. The initial adapter is required to convert a trained Astra adapter into a standard LoRA adapter.

# After preprocessing, before training:
peft_model.peft_config["default"].init_lora_weights = True
peft_model.save_pretrained(
    os.path.join(residual_model_path, "astra_init")
)

# After training:
peft_model.save_pretrained(
    lora_output_dir,
    path_initial_model_for_weight_conversion=os.path.join(
        residual_model_path, "astra_init"
    ),
)

The converted adapter can be loaded on top of the original base model and used with standard LoRA tooling.


πŸ“¦ Prepare Datasets

We use the processed datasets uploaded to the Hugging Face Hub by PiSSA.

from datasets import load_dataset

train_data = load_dataset("fxmeng/pissa-dataset", split="train")

# MetaMathQA
math_types = {
    "GSM_Rephrased", "GSM_AnsAug", "GSM_SV", "GSM_FOBAR",
    "MATH_Rephrased", "MATH_AnsAug", "MATH_SV", "MATH_FOBAR",
}
train_data = train_data.filter(lambda example: example["type"] in math_types)
train_data.to_json("dataset/metamath/train.json")

# CodeFeedback-Python
code_types = {"python"}
train_data = train_data.filter(lambda example: example["type"] in code_types)
train_data.to_json("dataset/python/train.json")

# Commonsense reasoning
train_data = load_dataset("zwhe99/commonsense_170k", split="train")
train_data.to_json("dataset/commonsense/train.json")

πŸ” Reproduce Results

Run the benchmark scripts in this repository:

# MetaMath
bash scripts/metamath/run.sh

# Code generation
bash scripts/code/run.sh

# Commonsense reasoning
bash scripts/commonsense/run.sh

🧭 Repository Map

Path Purpose
init_astra.py Build an Astra-initialized residual model using the official PEFT API.
train.py Fine-tune an Astra-initialized model.
scripts/metamath/ Mathematical reasoning experiments.
scripts/code/ Code generation experiments.
scripts/commonsense/ Commonsense reasoning experiments.
dataset/ Calibration data utilities.
utils/ Evaluation and merging utilities.

πŸ‘₯ Core Contributors

Contributor Role
Kainan Liu Β· GitHub Lead author and maintainer
Yong Zhang Β· GitHub Co-author and advisor

πŸ“ Citation

If you use Astra, please cite:

@inproceedings{liu2026astra,
  title = {Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models},
  author = {
    Liu, Kainan and Zhang, Yong and Cheng, Ning and Zhu, Yun and
    Wang, Yanmeng and Wang, Shaojun and Xiao, Jing
  },
  booktitle = {Findings of the Association for Computational Linguistics: ACL 2026},
  year = {2026}
}

πŸ“„ License

This project is released under the MIT License.

About

[ACL'26] The official implementation for "Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models"

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages