Quantitative inference / computational systems

Shucheng Cao 曹书诚

Also known and published as Bangli Cao

I work on quantitative problems where conclusions must survive confounding, distribution shift, and independent checks.

I am a Quantitative Life Sciences PhD candidate at McGill University. I build end-to-end computational pipelines across statistical inference, machine learning, simulation, and AI systems—taking raw, heterogeneous data through modelling, self-skeptical validation, and decision-ready output. My current applications are in genetics and biomedicine; the computational problems recur across data-intensive domains.

Portrait of Shucheng Cao
PhD candidate
McGill University
Montréal, Canada

Problems, systems, evidence

The projects below began with different practical questions: which targets are causal, which AI claims are supported, whether a benchmark is trustworthy, whether a model will generalize, and which molecular mechanism should guide design. Each required a different method, but the same standard: make the result testable and useful to the next decision.

01 Statistical inference
McGill

Causal targets for metabolic liver disease

Many protein–disease associations do not survive checks for confounding, multiple testing, or changes in measurement platform. I built an R/HPC pipeline across more than 790,000 samples that estimates causal effects, tests colocalization, and requires replication on two independent proteomic platforms. Thousands of candidates narrowed to four high-confidence targets. The result gives drug-discovery teams a shorter, human-genetics-supported list to investigate instead of prioritizing proteins from association alone.

02 Applied AI
CABS 2026

Evidence-grounded LLM for drug-target prioritization

Biomedical LLMs can produce fluent conclusions that are not supported by the evidence they retrieved. I architected and shipped OpenCausal, a Python/Streamlit system spanning eight public databases, with an append-only provenance ledger, deterministic report rendering, and claim validators. Sixty-two regression tests expose unsupported output before release. The same tool layer now serves 991 dossiers built from 101,543 published estimates, giving scientists an auditable starting point for drug-target prioritization.

An OpenCausal IL6R evidence card marked validation failed because the generated interpretation contradicted the retrieved estimate
A real failed run. The evidence card remains readable, but the validator exposes two unsupported claims before the page can pass.

03 Geometric deep learning
McGill

A reproducible GNN benchmark for neuron classification

Could a neuron's 3D shape identify its cell type, and could a new training-free method beat a strong deep-learning baseline? I reproduced MorphoGNN and built a PyTorch pipeline that converts raw SWC files into point clouds across three datasets. An audit uncovered a silent mapping bug affecting 97% of samples; after leakage checks, stratified splits, five seeds, and deterministic training, the corrected baseline reached 80.1 ± 1.0% balanced accuracy, making the method comparison defensible.

04 Model evaluation
IEEE TPAMI

Can medical AI generalize across hospitals?

Medical-imaging models often perform well on curated benchmarks but fail when hospitals, scanners, and protocols change. I helped build AbdomenCT-1K, a 1,112-case dataset from 12 hospitals, and used cross-hospital evaluation to measure how leading multi-organ segmentation models degraded outside their training distribution. The study made deployment risk measurable rather than assumed. It was published in IEEE TPAMI, and its open dataset and toolkit have since been used by more than 4,000 teams.

05 Computational simulation
KAUST

Simulation-guided design for an aggregation inhibitor

What physical interaction drives ultrashort peptides to self-assemble, and can that mechanism guide an inhibitor for Alzheimer’s-linked aggregation? I ran controlled molecular-dynamics experiments on Linux/HPC, changing hydrophobicity and aromaticity separately, then built a quantitative ranking of new peptide designs. The results challenged a 20-year field assumption and produced the first proof-of-concept simulation for this inhibitor class, replacing trial-and-error with a computable shortlist for experimental testing.

A working ruleI treat traceability as part of the result: another person should be able to reconstruct the inputs, assumptions, checks, and failure modes behind it.

Tools used across these projects: Python, R, PyTorch, Shell, and Linux/HPC.

Experimental grounding: Earlier, I built 3D tumour and humanized models to test whether drug responses seen in simplified assays survived in more realistic systems. That experience taught me to treat every computational metric as a proxy that must be checked against the real objective.

Papers behind the work

These are the durable outputs of the questions above: a human-genetics study for target prioritization, an open benchmark for clinical-AI robustness, a 3D platform for drug screening, and a simulation framework for molecular design.

  1. 2026 Disentangling osteoarthritis-specific genetic effects from obesity to identify novel therapeutic targetsCY. Su, M. Hasebe, D. Tan, S. Cao, et al., G. Butler-Laporte. Nature Communications — accepted; preprint available.Related work: 01 · causal inference ↑
  2. 2023 A PDA-Functionalized 3D Lung Scaffold Bioplatform to Construct Complicated Breast Tumor Microenvironment for Anticancer Drug Screening and ImmunotherapyW. Zhang, Y. Chen, M. Li, S. Cao, et al. Advanced Science.Experimental grounding · proxy versus real response
  3. 2022 AbdomenCT-1K: Is Abdominal Organ Segmentation a Solved Problem?J. Ma, S. Cao, et al. IEEE Transactions on Pattern Analysis and Machine Intelligence.Related work: 04 · OOD evaluation ↑
  4. 2022 MD simulations of amphiphilic ultrashort peptide self-assemblyS. Cao. MSc thesis, KAUST — ranked among the university's top ten theses of the year.Related work: 05 · controlled simulation ↑

See the full publication record on Google Scholar ↗

Writing and recent updates

  • Our osteoarthritis genetics study was accepted at Nature Communications.
  • Presented the MASLD proteome-wide causal-inference work at ESHG 2026 in Gothenburg.
  • Gave an oral presentation at McGill QLS Research Day.
  • Invited talk at the Cardiovascular & Metabolic Innovation Centre, Guangdong Medical University.
  • Awarded the FRQNT Doctoral Scholarship.

Get in touch

I am interested in quantitative research, applied AI, and research-engineering roles wherever model reliability and evidence quality matter—from finance and technology to bioinformatics and drug R&D. My current domain experience is strongest in biomedicine. Email is the best way to reach me.

shucheng.cao@mail.mcgill.ca