Free for humans

Sparse concept attribution for histomorphological hypothesis generation from whole-slide classifiers

SCOPE explains whole-slide pathology classifiers by sparsely attributing them onto a histomorphological concept bank — and a new benchmark shows sparse attributions recover known morphology while dense ones look like chance.

arXiv:2609.029855 min readScore 63/100Paper hub2026-W37

The 30-second take

  • What: The authors pair pathology vision–language models with sparse concept attribution (SCOPE) and introduce MorphoRecoveryBench, seven pathologist-curated tasks where sparse (and a cheaper pooled-embedding variant) beat dense attribution.
  • Why it matters (abundance angle): Turning slides into trustworthy hypotheses still burns scarce pathologist time. Automated, checkable morphological explanations are a mid-horizon step toward cheaper diagnostic insight — expert validation remains mandatory.
  • Who should care: Computational pathology groups, explainable-ML researchers, and clinicians who will not accept black-box slide scores without morphology.

What the paper actually did

Histology slides are rich, but linking morphological phenotypes to clinical attributes usually still needs a person to interpret the image. The authors argue that interpretable deep learning can automate hypothesis generation.

SCOPE interprets slide-level classifiers by combining pathology-specific vision–language models with sparse concept attribution onto a generalist histomorphological concept bank. To test whether explanations recover known morphology, they introduce MorphoRecoveryBench: seven tasks with pathologist-curated reference descriptions. On that benchmark, dense concept attribution is indistinguishable from a random baseline, while sparse attribution recovers substantial known morphology. Decomposing the pooled slide embedding reaches similar explanation correctness at a fraction of the computational cost. They conclude that post-hoc interpretation of whole-slide classifiers can generate morphological hypotheses at scale for expert validation.

What makes this disruptive

The result that dense concept attribution ≈ random, while sparse attribution recovers known morphology, is a direct hit on a fashionable XAI pattern in pathology. If it holds, “more concepts, densely” is the wrong default.

The scarce capability is expert morphological interpretation. SCOPE plus a pathologist-curated bench is a bid to make hypothesis generation abundant and still checkable. It does not replace the expert — the abstract says validation is the point.

Why it matters (outside the lab)

Abundance lens: diagnostics and biological insight are slow and expensive when they need scarce specialists. Tools that draft morphological hypotheses for experts to accept or reject are a step toward wider access.

Horizon is mid-range and regulation-sensitive. Near-term: use MorphoRecoveryBench when someone claims their explainer “looks pathological.” Medium-term: only prospective clinical studies (not this paper) could make this a default. No invented diagnostic-product year.

Limitations & open questions

The benchmark has seven tasks with curated references; “substantial” recovery is not a sensitivity/specificity for diagnosis. Post-hoc explanations of existing classifiers inherit those classifiers’ biases. Dense-vs-sparse is about this concept bank and protocol.

Preprint ≠ device. Abundance is not automatic: generating hypotheses at scale can create more work if experts must triage nonsense. Human validation is required by the authors’ own framing.

Explain ladder

Default article depth

SCOPE is post-hoc interpretation, not a new slide classifier. The scientific control is MorphoRecoveryBench with pathologist-curated reference descriptions on seven tasks. Carry this sentence into reviews: dense concept attribution was indistinguishable from random, while sparse attribution recovered substantial known morphology, and a cheaper pooled-embedding decomposition was similarly correct. Hypotheses still go to experts.

Key terms

Whole-slide classifier
A model that assigns a label or score to an entire histology slide rather than a single patch.
Concept attribution
Explaining a model by scoring how much named morphological concepts contribute to its decision.
MorphoRecoveryBench
The authors’ seven-task benchmark with pathologist-curated reference descriptions of known morphology.
Democratization of abundance
Editorial lens: scarce expert slide interpretation becoming cheaper to draft — validation still required, no clinical dates.

Sources

Related explainers

Same topic and week first — keep exploring the scarcity → abundance map.

Editorial explainer · not peer review · always read the primary paper.

Byline: Disruptive Concepts editorial.