Table of Contents
TL;DR: Biased memos from MLHC 2026 talks/posters.
Thu Aug 13 09:00 EDT 2026
-
Google AMIE team
-
Diagnosis: https://www.nature.com/articles/s41586-026-10764-5
-
Expert-level medical care https://www.nature.com/articles/s41586-026-10764-5
-
Had access to NEJM case studies and single-site EHR text
-
No access to researchers
-
-
How can we leverage RCT in observational data analysis or vice versa
-
The Illusion of Learning from Observational Data: An Empirical Bayes Perspective
-
Many people seem to (or would like to) train LLMs on PT-HCP interaction data
-
Having an agentic system for healthcare, are we opening up the possibility of over-triage, over-diagnosis, and over-treatment?
-
No physician will be judged numerically.
-
What will be the major obstacles for human-computer collaborations?
Thu Aug 13 14:05 EDT 2026
-
Ilya Shpitser
- "Recursive equations for imputation of missing not at random"
- Two approaches:
- Model missing status, e.g., (1) inverse propensity weighting (2) semi-parametric
- Model missing values
- Missing data = a counterfactual problem
- MCAR: $X^{(1)} \to X \gets R$,
- MAR: confounder $\to X^{(1)}$ and $R$
- MNAR: $X^{(1)} \to R \to X$ and $X^{(1)} \to X$, "self-censoring"
- Need to model $p(X|R, X^{(1)})$
- No self-censoring model, a conditional MRF
- IF a DAG is a submodel of NSC...
- Semi-parametric estimator available
- MICE relies on MAR
- goal: $p(X^{(1)},O|r)$ = extrapolation $\times$ interpolation
- extra: $p(X|X',O,r)$, Gibbs
- intra: $p(X'|O,r)$
- Estimable vs. estimation
-
Concept bottleneck model
- Can we prevent a large network from forgetting in continual learning?
- Just because we can choose not to learn a part of a model... we switch off gradient flows... (soft or hard)?
-
Reinforcement learning with data censorship
- Matt Engelhard, Duke
- POMDP; Survival data sets are clearly censored.
- Early alert, pseudo label, post-hoc adaptation
Fri Aug 14 09:50 EDT 2026
-
Anna Goldenberg
- Heterogeneity in large-scale ICU data
- Self-supervised learning representation learning
- Different sampling intervals
- Not high-dimensional?
- Hierarchical Dirichlet Process, flow model...
- Representation = latent states?
- How do we take into account patient specificity?
- DynaSub: adaptive subgrouping with encoders
-
Hoifung Poon
- Learning the language of patients across modalities
- Virtual patient
- Targeted therapy worked but hard to deal with relapse
- Immunotherapy, Keytruda (checkpoint blockade), not for everyone
- Need to learn tumour microenvironment grammar
- The other side... autoimmune disease
- Amara's Law
"We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run"
- Physical experiments are expensive
- Can we simulate a clinical trial based on real-world data? (NEJM AI)
- Patient journey is multimodal
- Precision health is a multimodal generation problem
- Dealing with missing information is a genuinely important challenge
- Can we model out missing information?
- Virtual patient, digital twin
- Unimodal: encoder $\to$ decoder
- Digital pathology: a transformer may not be the right architecture
- 16 x 16 = 56 million tokens
- A whole-slide foundation model
- Proposing "text" as an interlingua medium
- Spatial proteomics/transcriptomics is expensive
- GigaTIME
- Can we predict next medical event of a person?
- EPIC COSMOS database contains 115 Billion medical events