Research program
Biobanks Linked to Electronic Health Records
What can years of health records tell us that a genome or a single clinical visit cannot?
We link genomic data with years of clinical history to study disease patterns, treatment, and outcomes across large biobanks.
Why health-record data need careful handling
Electronic health records were created for care, not research. They cover long periods, but what gets recorded depends on who receives care, when they are seen, and how clinical practice changes. Phenotype definitions and comparisons across datasets have to account for that history.
Building phenotypes and testing findings
Define clinical phenotypes
Documented definitions and mappings make clinical phenotypes easier to inspect, reproduce, and revise.
Combine data sources
Genotypes, diagnoses, treatments, outcomes, and survey responses each add a different kind of evidence.
Account for how care is recorded
Who receives care, how often they are seen, and what clinicians record can shape both associations and model performance.
Replicate findings across biobanks
Comparisons between the Michigan Genomics Initiative, UK Biobank, and other datasets help identify results that depend on one setting.
Selected work
These papers describe how the Michigan Genomics Initiative (MGI) is assembled and how its records help researchers track the timing of diagnoses, account for selection bias, and link genotypes with prescribing histories.
Current work
Projects and collaborations
Public research resource
Current collection with archived releasesPRSweb Research Portals
PRSweb provides polygenic risk score evaluations, phenome-wide association results, and downloadable score files.
Related resources
Tools & data
Polygenic risk resource
Public collectionPRSweb Research Portals
Compare published polygenic risk scores evaluated in the Michigan Genomics Initiative (MGI) and UK Biobank, and explore phenome-wide associations and downloadable files.
PheWAS software
PublicphewasFlow
An R package for reproducible phenome-wide association studies in both directions, with local and cluster workflows, multiple-testing correction, and Manhattan and volcano plots.
Publications
Selected papers
The Michigan Genomics Initiative: A biobank linking genotypes and electronic clinical records in Michigan Medicine patients
Cell Genom
Why it matters: MGI illustrates how a health-system biobank can complement population-based cohorts. Recruiting through surgical care provides useful case counts for many clinical outcomes, even in a smaller cohort. That sampling design still needs to be considered when interpreting the results.
Phenotype risk scores (PheRS) for pancreatic cancer using time-stamped electronic health record data: Discovery and validation in two large biobanks
J Biomed Inform
Why it matters: Adding timing to diagnosis codes made the phenotype more informative than a simple comorbidity list. A pancreatic cancer phenotype score built in the Michigan Genomics Initiative (MGI) from records five years before diagnosis also showed an association in UK Biobank and added information beyond a cancer polygenic risk score (PRS) and standard risk factors. This shows what longitudinal phenotyping and external validation can contribute.
To weight or not to weight? The effect of selection bias in 3 large electronic health record-linked biobanks and recommendations for practice
J Am Med Inform Assoc
Why it matters: Weighting is not an automatic fix for a nonrepresentative biobank. Across three cohorts, it changed prevalence and targeted effect estimates more than broad discovery analyses. The practical lesson is to define the target population and estimand first, then decide whether weighting improves the analysis.
Identifying the prevalence of clinically actionable drug-gene interactions in a health system biorepository to guide pharmacogenetics implementation services
Clin Transl Sci
Why it matters: Pharmacogenetics matters clinically when a patient carries a relevant genotype and receives a medication whose dosing or effects may depend on it. In MGI, guideline-defined drug–gene interactions appeared across specialties, and nearly a quarter of evaluable patients had more than one. That pattern supports panel-based, system-wide implementation rather than tackling one gene and drug at a time.
Updates
News and talks
· Video
MGI 2024: Lars Fritsche
In this U-M symposium talk, Lars Fritsche reviews a decade of EHR-linked genetics research, including PheWeb and PRSweb.