Research program

Biobanks Linked to Electronic Health Records

What can years of health records tell us that a genome or a single clinical visit cannot?

We link genomic data with years of clinical history to study disease patterns, treatment, and outcomes across large biobanks.

Why health-record data need careful handling

Electronic health records were created for care, not research. They cover long periods, but what gets recorded depends on who receives care, when they are seen, and how clinical practice changes. Phenotype definitions and comparisons across datasets have to account for that history.

Building phenotypes and testing findings

Define clinical phenotypes

Documented definitions and mappings make clinical phenotypes easier to inspect, reproduce, and revise.

Combine data sources

Genotypes, diagnoses, treatments, outcomes, and survey responses each add a different kind of evidence.

Account for how care is recorded

Who receives care, how often they are seen, and what clinicians record can shape both associations and model performance.

Replicate findings across biobanks

Comparisons between the Michigan Genomics Initiative, UK Biobank, and other datasets help identify results that depend on one setting.

Selected work

These papers describe how the Michigan Genomics Initiative (MGI) is assembled and how its records help researchers track the timing of diagnoses, account for selection bias, and link genotypes with prescribing histories.

Current work

Projects and collaborations

Public research resource

Current collection with archived releases

PRSweb Research Portals

PRSweb provides polygenic risk score evaluations, phenome-wide association results, and downloadable score files.

Related resources

Tools & data

Polygenic risk resource

Public collection

PRSweb Research Portals

Compare published polygenic risk scores evaluated in the Michigan Genomics Initiative (MGI) and UK Biobank, and explore phenome-wide associations and downloadable files.

View resource details

PheWAS software

Public

phewasFlow

An R package for reproducible phenome-wide association studies in both directions, with local and cluster workflows, multiple-testing correction, and Manhattan and volcano plots.

View resource details

Publications

Selected papers

All publications →
2023

The Michigan Genomics Initiative: A biobank linking genotypes and electronic clinical records in Michigan Medicine patients

Cell Genom

Why it matters: MGI illustrates how a health-system biobank can complement population-based cohorts. Recruiting through surgical care provides useful case counts for many clinical outcomes, even in a smaller cohort. That sampling design still needs to be considered when interpreting the results.

2021

Phenotype risk scores (PheRS) for pancreatic cancer using time-stamped electronic health record data: Discovery and validation in two large biobanks

J Biomed Inform

Why it matters: Adding timing to diagnosis codes made the phenotype more informative than a simple comorbidity list. A pancreatic cancer phenotype score built in the Michigan Genomics Initiative (MGI) from records five years before diagnosis also showed an association in UK Biobank and added information beyond a cancer polygenic risk score (PRS) and standard risk factors. This shows what longitudinal phenotyping and external validation can contribute.

2024

To weight or not to weight? The effect of selection bias in 3 large electronic health record-linked biobanks and recommendations for practice

J Am Med Inform Assoc

Why it matters: Weighting is not an automatic fix for a nonrepresentative biobank. Across three cohorts, it changed prevalence and targeted effect estimates more than broad discovery analyses. The practical lesson is to define the target population and estimand first, then decide whether weighting improves the analysis.

2023

Identifying the prevalence of clinically actionable drug-gene interactions in a health system biorepository to guide pharmacogenetics implementation services

Clin Transl Sci

Why it matters: Pharmacogenetics matters clinically when a patient carries a relevant genotype and receives a medication whose dosing or effects may depend on it. In MGI, guideline-defined drug–gene interactions appeared across specialties, and nearly a quarter of evaluable patients had more than one. That pattern supports panel-based, system-wide implementation rather than tackling one gene and drug at a time.

Updates

News and talks

· Video

MGI 2024: Lars Fritsche

In this U-M symposium talk, Lars Fritsche reviews a decade of EHR-linked genetics research, including PheWeb and PRSweb.