traceCB maps trans-ancestry cell-type-specific eQTL effects by integrating single-cell and bulk-tissue summary statistics. This repository contains the Python package, full-data workflows, simulations, and manuscript analyses.
src/traceCB/: installable traceCB package.scripts/: full-data preprocessing, LD-score, traceCB, and colocalization workflows.src/simulation/: main, supplementary, and chromosome 22 simulations.src/enrichment/: reproducible ancestry-matched S-LDSC and pathway-enrichment analyses; seesrc/enrichment/README.md.src/figures/: manuscript figure scripts and shared metadata.src/preprocess/andsrc/coloc/: workflow implementations used byscripts/.data/toy_example/: small public inputs for the tutorial.
Large input data and generated results are intentionally excluded from Git.
Full-data input paths are configured in scripts/config.sh; its data root
currently points to /home/wjiang49/group/wjiang49/data. Generated pipeline
outputs and figures default to results/. Edit the roots for your filesystem;
all Python and R manuscript figures share this configuration.
Start with the path configuration guide for a copyable local configuration, required file layouts, separate input/output roots, and checks to run before full-data analysis or plotting.
Python 3.10 or newer is required. For repository-level reproduction, create the
Python 3.12 reference environment for release 1.0 from the tracked specification:
git clone https://github.com/lucajiang/traceCB.git
cd traceCB
conda env create -f environment.yml
conda activate py312For a library-only installation from the source directory, use pip install .
in a supported Python environment, or pip install -e . for development.
The wheel installs only the Python library. The source archive additionally
includes the workflows, documentation, and public tutorial inputs.
When using a lightweight or custom environment, add the corresponding optional dependencies for the local notebook and manuscript analyses:
pip install -e '.[tutorial]' # local Jupyter tutorial
pip install -e '.[enrichment,figures]' # enrichment and figure scripts
pip install -e '.[simulation]' # simulation and plotting scriptsThe tutorial notebooks are available at
docs/tutorial/run_traceCB.ipynb and on
Google Colab.
They use the tracked files in data/toy_example/ and demonstrate the model on a
single gene.
The shell workflows require a Unix-like environment, Conda, and the external PLINK, S-LDXR, and R tools listed in the pipeline guide. These tools and the full study datasets are separate from the Python wheel.
External inputs such as population-specific eQTLs, tissue eQTLs, and 1000
Genomes reference panels are not redistributed here. Set their locations in the
environment or edit the defaults in scripts/config.sh.
The BBJ cell-type eQTL data used for the EAS analysis are available from
Human Database of Japan: hum0099-v1.
See the pipeline guide for all
data sources, expected input formats, and preprocessing commands.
# Optional examples
export TRACECB_DATA_ROOT=/path/to/input-data
export TRACECB_OUTPUT_ROOT=/path/to/results
export SLDXR_DIR=/path/to/s-ldxr
export PLINK_BIN=/path/to/plink
bash scripts/prepare_inputs.sh
bash scripts/run_ld_scores.sh
bash scripts/run_gmm.sh
bash scripts/run_colocalization.sh # optionalSee the simulation guide for manuscript simulation entry points and the enrichment guide for the S-LDSC and pathway-enrichment analyses. Figure-specific dependencies and invocation patterns are listed in the figure guide.
Before generating the case-study figures, edit TRACECB_STUDY_DIR,
TRACECB_GTEX_GENE_ANNOTATION, and TRACECB_FIGURE_DIR in
scripts/config.sh to match your data and output locations.
The input defaults refer to the current machine and must be replaced on other
systems. In a Python environment with .[figures] installed, run:
source scripts/config.sh
python -m figures.case_studysource scripts/config.sh sets the shared figure paths, PYTHONPATH, and the
default headless Matplotlib backend in the current shell. Run it once per
terminal session before invoking Python figure scripts.
If you use traceCB, please cite the paper:
@article{jiang2026tracecb,
title={{traceCB}: Trans-ancestry cell-type-specific {eQTLs} mapping by integrating {scRNA-seq} and bulk data},
author={Jiang, Wenxin and Xiao, Jiashun and Cai, Mingxuan},
journal={bioRxiv},
year={2026},
doi={10.64898/2026.06.20.733502},
url={https://doi.org/10.64898/2026.06.20.733502},
publisher={Cold Spring Harbor Laboratory}
}Machine-readable citation metadata are available in CITATION.cff.
The source code, documentation, and processed tutorial example
are maintained in this repository. Checksums for the tutorial inputs are stored
in data/toy_example/SHA256SUMS.
The full study datasets are not redistributed. Their source repositories,
reference builds, access restrictions, and preprocessing formats are documented
in docs/pipeline.md. The EUR S-LDSC workflow additionally
records reference checksums and run provenance in
src/enrichment/README.md.
traceCB is distributed under the GPL-3.0 license.
