Skip to content
LucaJiangPublic

About

Repo for "traceCB: Trans-ancestry cell-type-specific eQTLs mapping by integrating scRNA-seq and bulk data"

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

traceCB

Python 3.10+ License: GPL-3 bioRxiv Open Tutorial In Colab

traceCB maps trans-ancestry cell-type-specific eQTL effects by integrating single-cell and bulk-tissue summary statistics. This repository contains the Python package, full-data workflows, simulations, and manuscript analyses.

traceCB workflow

Repository layout

  • src/traceCB/: installable traceCB package.
  • scripts/: full-data preprocessing, LD-score, traceCB, and colocalization workflows.
  • src/simulation/: main, supplementary, and chromosome 22 simulations.
  • src/enrichment/: reproducible ancestry-matched S-LDSC and pathway-enrichment analyses; see src/enrichment/README.md.
  • src/figures/: manuscript figure scripts and shared metadata.
  • src/preprocess/ and src/coloc/: workflow implementations used by scripts/.
  • data/toy_example/: small public inputs for the tutorial.

Large input data and generated results are intentionally excluded from Git. Full-data input paths are configured in scripts/config.sh; its data root currently points to /home/wjiang49/group/wjiang49/data. Generated pipeline outputs and figures default to results/. Edit the roots for your filesystem; all Python and R manuscript figures share this configuration.

Start with the path configuration guide for a copyable local configuration, required file layouts, separate input/output roots, and checks to run before full-data analysis or plotting.

Installation

Python 3.10 or newer is required. For repository-level reproduction, create the Python 3.12 reference environment for release 1.0 from the tracked specification:

git clone https://github.com/lucajiang/traceCB.git
cd traceCB
conda env create -f environment.yml
conda activate py312

For a library-only installation from the source directory, use pip install . in a supported Python environment, or pip install -e . for development. The wheel installs only the Python library. The source archive additionally includes the workflows, documentation, and public tutorial inputs.

When using a lightweight or custom environment, add the corresponding optional dependencies for the local notebook and manuscript analyses:

pip install -e '.[tutorial]'            # local Jupyter tutorial
pip install -e '.[enrichment,figures]'  # enrichment and figure scripts
pip install -e '.[simulation]'         # simulation and plotting scripts

Quick start

The tutorial notebooks are available at docs/tutorial/run_traceCB.ipynb and on Google Colab. They use the tracked files in data/toy_example/ and demonstrate the model on a single gene.

Full-data workflow

The shell workflows require a Unix-like environment, Conda, and the external PLINK, S-LDXR, and R tools listed in the pipeline guide. These tools and the full study datasets are separate from the Python wheel.

External inputs such as population-specific eQTLs, tissue eQTLs, and 1000 Genomes reference panels are not redistributed here. Set their locations in the environment or edit the defaults in scripts/config.sh. The BBJ cell-type eQTL data used for the EAS analysis are available from Human Database of Japan: hum0099-v1. See the pipeline guide for all data sources, expected input formats, and preprocessing commands.

# Optional examples
export TRACECB_DATA_ROOT=/path/to/input-data
export TRACECB_OUTPUT_ROOT=/path/to/results
export SLDXR_DIR=/path/to/s-ldxr
export PLINK_BIN=/path/to/plink

bash scripts/prepare_inputs.sh
bash scripts/run_ld_scores.sh
bash scripts/run_gmm.sh
bash scripts/run_colocalization.sh  # optional

See the simulation guide for manuscript simulation entry points and the enrichment guide for the S-LDSC and pathway-enrichment analyses. Figure-specific dependencies and invocation patterns are listed in the figure guide.

Before generating the case-study figures, edit TRACECB_STUDY_DIR, TRACECB_GTEX_GENE_ANNOTATION, and TRACECB_FIGURE_DIR in scripts/config.sh to match your data and output locations. The input defaults refer to the current machine and must be replaced on other systems. In a Python environment with .[figures] installed, run:

source scripts/config.sh
python -m figures.case_study

source scripts/config.sh sets the shared figure paths, PYTHONPATH, and the default headless Matplotlib backend in the current shell. Run it once per terminal session before invoking Python figure scripts.

Citation

If you use traceCB, please cite the paper:

@article{jiang2026tracecb,
  title={{traceCB}: Trans-ancestry cell-type-specific {eQTLs} mapping by integrating {scRNA-seq} and bulk data},
  author={Jiang, Wenxin and Xiao, Jiashun and Cai, Mingxuan},
  journal={bioRxiv},
  year={2026},
  doi={10.64898/2026.06.20.733502},
  url={https://doi.org/10.64898/2026.06.20.733502},
  publisher={Cold Spring Harbor Laboratory}
}

Machine-readable citation metadata are available in CITATION.cff.

Code and Data Availability

The source code, documentation, and processed tutorial example are maintained in this repository. Checksums for the tutorial inputs are stored in data/toy_example/SHA256SUMS.

The full study datasets are not redistributed. Their source repositories, reference builds, access restrictions, and preprocessing formats are documented in docs/pipeline.md. The EUR S-LDSC workflow additionally records reference checksums and run provenance in src/enrichment/README.md.

License

traceCB is distributed under the GPL-3.0 license.

About

Repo for "traceCB: Trans-ancestry cell-type-specific eQTLs mapping by integrating scRNA-seq and bulk data"

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages