Cross-kingdom phenotype annotations for 35,856 species generated with a web-search-enabled language model
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Standardized organismal trait matrices support comparative biology, macroecology and conservation, but broad taxonomic coverage is difficult to obtain by manual curation or corpus-bound extraction. This Data Descriptor presents a cross-kingdom phenotype dataset generated by a web-search-enabled large language model embedded in a scripted annotation pipeline. The dataset contains 7,027,776 ordinal phenotype annotations for 35,856 animal, plant and fungal species across 196 questions spanning cellular biology, anatomy, physiology, ecology, behavior, life history, reproduction, environmental tolerance and human associations. Each annotation uses a five-level scale from Definitely No to Definitely Yes, including an Unknown or not-applicable category. The repository provides organism and question metadata, raw and processed annotation matrices, imputation outputs, derived phenotype scores, validation tables and analysis code. Technical validation includes comparisons with independent trait databases, a focused audit of 200 organisms, latent-structure analyses and phenotype-similarity checks. The data are intended for exploratory comparative analyses, hypothesis generation, benchmarking and targeted expert review, rather than as substitutes for primary measurements or curated species-level observations.