Free for humans

Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?

PubChem’s ~2 million bioassays are almost empty of BioAssay Ontology format and detection labels; seven LLMs can often recover those fields from assay text — and sometimes even nudge an expert to fix a silver label.

arXiv:2610.016165 min readScore 92/100 · editorial triage · not peer reviewPaper hub2026-W41

The 30-second take

  • What: The authors measure missing BioAssay Ontology metadata in PubChem and test whether open and proprietary language models can predict assay format and detection method from the assay write-up.
  • Why it matters: Molecular property models need clean assay metadata; if models can audit sparse public and industrial labels, drug-discovery data becomes cheaper to use without waiting for a full human recuration.
  • Who should care: Cheminformatics and foundation-model teams, industrial assay curators, and anyone building property predictors on PubChem or ChEMBL.

What the paper actually did

The paper starts from a data-readiness problem: public repositories and industrial screening databases often have missing, inconsistent, or conflated assay annotations that foundation models for molecular property prediction need. The authors quantify missing BioAssay Ontology (BAO) fields in PubChem: of about two million bioassays, 36% lack an assay format, 89% lack a BioAssay type, and more than 99.9% lack any BAO-mapped assay format or detection-technology term. They then ask whether open-source and proprietary LLMs can predict and audit those annotations from assay text. On evaluation sets from PubChem and ChEMBL, seven models are compared to existing silver labels. Recall is at least 0.96 for biochemical and cell-based assay formats, with a similar pattern for detection technology, though disagreements rise on under-represented classes. Manual inspection attributes many disagreements to inconsistent silver sources rather than model error. In a qualitative study, LLM-generated evidence led a senior industrial curator to revise some of their own labels. Proprietary versus open-source performance differences were small.

What makes this disruptive

The scarcity here is not another property-prediction architecture — it is usable metadata. Showing that PubChem is critically sparse on BAO format and detection terms reframes “AI-ready chemistry data” as a labeling problem at million-assay scale. The second result is that current LLMs already agree strongly with silver labels on common format classes, and that disagreements often point at the labels, not the model. A curator changing their own annotation after seeing model evidence is a rare, concrete signal that models can audit humans, not only imitate them. Small gaps between proprietary and open models matter for cost: if open models are close enough, large-scale curation need not sit behind an API bill. This is still an assessment paper, not a replacement for ontology experts.

Why it matters (outside the lab)

Abundance lens: high-quality assay metadata is still a luxury of well-funded curation teams. If language models can fill and audit BAO fields from free text, the scarce input to molecular foundation models — reliable assay context — gets cheaper and more default. Horizon is mid-range: lab-to-pipeline use depends on per-class reliability and human review, which the authors themselves require before labels enter downstream ML. Near-term, treat the coverage numbers and recall figures as a roadmap for data-readiness work. Medium-term, cost curves and independent checks decide whether automated assay annotation becomes ordinary infrastructure.

Limitations & open questions

Silver labels are inconsistent, which both helps the “models catch label error” story and weakens any claim that high recall equals ground truth. Under-represented classes show more disagreement; majority-class success does not license unreviewed labels on rare formats or detection methods. The curator study is qualitative, not a powered inter-rater trial. The paper does not claim that predicted metadata are ready for unsupervised use in training property models. PubChem coverage statistics describe current dumps, not a permanent law of the repository. Open versus proprietary parity is reported as small in this assessment and may not hold for every ontology field. Preprint ≠ production curation system.

Explain ladder

Default article depth

Start with the PubChem sparsity numbers, then the recall results on common assay formats. The interesting claim is not that models are perfect annotators — it is that they can flag mislabeled assays and that open models were close to proprietary ones. Before updating a data-readiness plan, ask which BAO fields you would still hold for targeted human review. Horizon: mid, after reliability estimates and review workflows exist.

Key terms

BioAssay Ontology (BAO)
A controlled vocabulary for describing how a biological assay is formatted and how a signal is detected.
Silver label
An existing annotation treated as a useful but imperfect reference, not guaranteed ground truth.
Data readiness
Whether a dataset’s metadata is complete and consistent enough for reliable machine learning.
Democratization of abundance
Editorial lens: turning scarce elite curation into cheaper, more default data infrastructure — without a fake timeline.

Sources

Related explainers

Same topic and week first — keep exploring the scarcity → abundance map.

Editorial explainer · not peer review · always read the primary paper.

Byline: Disruptive Concepts editorial.