Utilizing large language models to construct a dataset of Württemberg’s 19th-century fauna from historical records

Maximilian Teich
Belen Escobari
Malte Rehbein

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Constructing datasets on past biodiversity from historical sources is crucial for understanding long-term ecological changes. Typically, compiling such datasets relies on prior knowledge of the sources’ composition and requires considerable manual effort. To overcome these challenges, we implement an automated approach based on prompted large language models (LLMs) to detect mentions of species in texts from 19th-century Württemberg and link these mentions to identifiers in the GBIF database. Based on our evaluation, we find that LLMs can reliably identify species in the texts with high recall (92.6%) and precision (95.3%), while providing estimates of the correct species identifier with considerable accuracy (83.0%).

Version published to 10.1101/2025.10.14.681982 on bioRxiv
Oct 16, 2025

Multi-Scale Computational Analysis of Wikipedia’s Telling of Global History

This article has 7 authors:
1. Steph Buongiorno
2. Jo Guldi
3. Marnie Hughes-Warrington
4. Nan Jiang
5. Rosie Larson
6. Sohan Bellam
7. Gregory J. Palermo
This article has no evaluationsLatest version Jan 19, 2026
Random forests in corpus research: A systematic review

This article has 1 author:
1. Lukas Sönning
This article has no evaluationsLatest version Jan 17, 2026
Random forests in corpus research: A systematic review

This article has 1 author:
1. Lukas Sönning
This article has no evaluationsLatest version Jan 17, 2026

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Multi-Scale Computational Analysis of Wikipedia’s Telling of Global History

Random forests in corpus research: A systematic review

Random forests in corpus research: A systematic review