HPOQuest: A Rare-Disease Diagnostic Agent Using Active Phenotype Acquisition
A training-free agent ranks rare-disease hypotheses and asks the next most useful clinical questions when the first symptoms are thin.
The 30-second take
- What: HPOQuest starts from a few observed phenotypes, keeps a probabilistic disease ranking, and iteratively picks follow-up questions; confirmed findings update ranks, and every answer updates the question pool.
- Abundance angle: today, expert rare-disease judgment is scarce — more than 300 million people live with one of over 7,000 conditions, and first visits often show incomplete, mixed signs. Sequential phenotype acquisition is a step toward diagnostic support as a cheaper default layer, not a replacement clinician (mid-horizon: validation and regulation decide access).
- Who should care: Clinical geneticists, rare-disease centers, digital-health teams building intake tools, and patients' advocates watching whether sparse first exams can be turned into sharper next questions.
What the paper actually did
The authors present HPOQuest, a training-free framework for sequential phenotype acquisition in rare-disease diagnosis. They start from the public-health backdrop that more than 300 million people worldwide are affected by one of over 7,000 known rare diseases, while diagnosis stays hard because first presentations are incomplete and heterogeneous.
From a small set of observed patient phenotypes, HPOQuest maintains a probabilistic ranking over diseases and iteratively selects informative follow-up questions meant to support clinicians during assessment. When a phenotype is confirmed, the disease ranking updates. All responses — not only confirmations — update the candidate question set.
They evaluate across four benchmark cohorts and report large lifts from sparse initial phenotypes: up to 30 percentage points at Recall@1 and 45 percentage points at Recall@5. They conclude that asking for phenotypes in sequence can improve rare-disease diagnosis when the first clinical evidence is limited.
What makes this disruptive
The scarce capability under pressure is expert-led differential diagnosis for rare disease when the phenotype list is short. Many AI diagnostic tools want large training sets or a complete HPO dump. HPOQuest is framed as training-free and sequential: it treats questioning as an active loop, not a one-shot classifier.
If the Recall@1 and Recall@5 gains hold on real clinics, the next-question problem becomes a software primitive rather than a specialist's tacit skill. That does not retire geneticists; it changes what a first-pass workup can extract from an incomplete exam.
The disruptiveness is the combination of no training step plus explicit phenotype acquisition, scored on four cohorts — a roadmap signal, not a cleared diagnostic device.
Why it matters (outside the lab)
Abundance lens: rare-disease expertise and complete phenotyping are luxuries concentrated in a few centers. If a lightweight ranking-and-questioning loop reliably improves recall from sparse signs, more of that judgment can sit closer to ordinary intake — a cheaper default, not a magic diagnosis.
Near-term, use the preprint to update how teams think about active questioning versus one-shot phenotype-to-disease models. Medium-term, prospective trials, liability, and regulators decide whether anything here becomes a default clinical tool.
Do not read the 30- and 45-point recall lifts as a promise that missed diagnoses vanish on a calendar. Cost, workflow fit, and independent replication still sit in the way.
Limitations & open questions
Preprint ≠ product. Benchmark-cohort recall is not the same as time-to-diagnosis or harm reduction in clinic. The abstract does not specify how questions are phrased for patients versus clinicians, how noisy or contradictory answers are handled beyond updating ranks and the question set, or how the method behaves on diseases outside the four cohorts.
Training-free does not mean assumption-free; the probabilistic ranking still depends on whatever knowledge base and phenotype ontology sit underneath (not detailed in the abstract). Abundance is not automatic: software that asks better questions does not by itself democratize genetics care.
Explain ladder
Default article depth
Rare-disease visits often start with a handful of signs that fit many conditions. HPOQuest keeps a running probability list of diseases and chooses the next question that would most shrink that list. Confirmed signs reshuffle the ranking; every answer, including negatives, changes what it might ask next.
It is not a model that was trained on millions of charts, according to the authors — it is a questioning policy on top of a ranking. On four test collections, starting from thin phenotype lists, the right disease showed up in the top slot or top five much more often than without this loop.
For health-system readers: this is about structuring the interview, not replacing a genome report.
Key terms
- Phenotype
- An observable clinical feature (sign, symptom, or finding) used to describe a patient; rare-disease work often codes these in ontologies such as HPO.
- Active phenotype acquisition
- Choosing the next most informative clinical question instead of scoring a complete, static feature list.
- Recall@k
- How often the true disease appears among the system's top k ranked guesses.
- Training-free
- The authors do not fit a new supervised model on labeled cases; ranking and questioning use an existing structured setup.
Sources
Related explainers
Same topic and week first — keep exploring the scarcity → abundance map.
Can 4D Foundation Models Remember?
2026-W39 · score 89 · Artificial Intelligencesame weeksame topic
Quantifying Overclaiming Propensity in Frontier LLM Agents
2026-W39 · score 80 · Artificial Intelligencesame weeksame topic
Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics
2026-W39 · score 74 · Artificial Intelligencesame weeksame topic
Score Centering Stabilizes Off-policy Reinforcement Learning
2026-W39 · score 62 · Artificial Intelligencesame weeksame topic
Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
2026-W37 · score 93 · Artificial Intelligencesame topic
Disruptiveness
Editorial triage 0–100 · not peer review
- Novelty97
- Impact100
- Field heat83
- Practicality89
- Controversy54
