Skip to Content

Search: {{$root.lsaSearchQuery.q}}, Page {{$root.page}}

Kai Shimagaki

Assistant Research Scientist

kaissh@umich.edu

Office Information:

Office Number: 1868 EH

Applied Mathematics; Mathematical Biology; Mathematics; Research Scientist

Education/Degree:

Sorbonne University Ph.D in Information Science Sep 2018 – Jul 2021 ◦ Thesis: Advanced statistical modeling and variable selection for protein sequences 2 ◦ Advisor: Martin Weigt
The University of Tokyo M.S. and Ph.D in Physics Sep 2015 – Jul 2018 ◦ Thesis: Performance analysis of parameter inference using unsteady relaxation and application to inverse Ising problem ◦ Advisor: Synge Todo 2 ◦ Completed in the first year of a Ph.D program and obtained all required credits
Ibaraki University B.S. in Physics and Mathematics Sep 2011 – Jul 2015 ◦ Thesis: Modeling and analyzing catalytic chemical reaction networks ◦ Advisor: Naoko Nakagawa 2 ◦ GPA: 3.96/4.0, Dean’s list ◦ Completed 180 credits (beyond the 120-credit requirement), primarily in physics, mathematics and information science, with additional coursework in biology and chemistry

Living organisms are built on multiple layers of proteins and molecules that are ultimatelyencoded in genetic and epigenetic information, and evolution operates on genotypes togenerate phenotypes under principled rules. A central problem is to quantitatively understandthe genotype–phenotype relationship and the mechanisms by which evolutionary processes create particular phenotypic traits. Recent advances in genetic sequencing technologies anddata accessibility have generated increasingly large, complex datasets, providing a uniqueopportunity to understand evolutionary processes across scales. These datasets are oftenstructured, and associations within/between them (e.g., correlations and autocorrelations) canreveal important biological insights, but they are frequently noisy, confounded by biological orexperimental factors, and difficult to interpret at scale. To address these issues, I developstatistical and computational methods for high-dimensional data, with a particular focus ongenetic and phenotypic datasets.
My research centers on three key areas: (1) addressing biologically motivated questions byintegrating high-dimensional data with close comparison to existing experimental and/or clinicalstudies; (2) developing statistically and computationally efficient algorithms that are applicable toreal biological data; and (3) clarifying the limitations of existing methods and the relationshipsamong different frameworks and theories. At MCAIM, I hope to develop and apply algorithmsthat tackle complex biological problems and deepen our understanding of evolutionaryprocesses and genotype–phenotype relationships. To this aim, I’ll employ statisticalinference/modeling and interpretable machine learning approaches, together with principledtheoretical frameworks (statistical physics, population genetics, quantitative trait).