Application of a methodological framework for the development and multicenter validation of reliable artificial intelligence in embryo evaluation

D. Gilboa
Akhil Garg
M. Shapiro
M. Meseguer
Y. Amar
N. Lustgarten
N. Desai
T. Shavit
V. Silva
A. Papatheodorou
A. Chatziparasidou
S. Angras
J. H. Lee
L. Thiel
C. L. Curchoe
Y. Tauber
D. S. Seidman

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Background

Artificial intelligence (AI) models analyzing embryo time-lapse images have been developed to predict the likelihood of pregnancy following in vitro fertilization (IVF). However, limited research exists on methods ensuring AI consistency and reliability in clinical settings during its development and validation process. We present a methodology for developing and validating an AI model across multiple datasets to demonstrate reliable performance in evaluating blastocyst-stage embryos.

Methods

This multicenter analysis utilizes time-lapse images, pregnancy outcomes, and morphologic annotations from embryos collected at 10 IVF clinics across 9 countries between 2018 and 2022. The four-step methodology for developing and evaluating the AI model include: (I) curating annotated datasets that represent the intended clinical use case; (II) developing and optimizing the AI model; (III) evaluating the AI’s performance by assessing its discriminative power and associations with pregnancy probability across variable data; and (IV) ensuring interpretability and explainability by correlating AI scores with relevant morphologic features of embryo quality. Three datasets were used: the training and validation dataset ( n = 16,935 embryos), the blind test dataset ( n = 1,708 embryos; 3 clinics), and the independent dataset ( n = 7,445 embryos; 7 clinics) derived from previously unseen clinic cohorts.

Results

The AI was designed as a deep learning classifier ranking embryos by score according to their likelihood of clinical pregnancy. Higher AI score brackets were associated with increased fetal heartbeat (FH) likelihood across all evaluated datasets, showing a trend of increasing odds ratios (OR). The highest OR was observed in the top G4 bracket (test dataset G4 score ≥ 7.5: OR 3.84; independent dataset G4 score ≥ 7.5: OR 4.01), while the lowest was in the G1 bracket (test dataset G1 score < 4.0: OR 0.40; independent dataset G1 score < 4.0: OR 0.45). AI score brackets G2, G3, and G4 displayed OR values above 1.0 ( P < 0.05), indicating linear associations with FH likelihood. Average AI scores were consistently higher for FH-positive than for FH-negative embryos within each age subgroup. Positive correlations were also observed between AI scores and key morphologic parameters used to predict embryo quality.

Conclusions

Strong AI performance across multiple datasets demonstrates the value of our four-step methodology in developing and validating the AI as a reliable adjunct to embryo evaluation.

Version published to 10.1186/s12958-025-01351-w
Jan 31, 2025
Version published to 10.21203/rs.3.rs-5438430/v1 on Research Square
Dec 17, 2024

Embryo selection tools in IVF can favour XY embryos: implications for equitable reproductive AI

This article has 7 authors:
1. Teodora Popa
2. Chloe He
3. Christian S. Ottolini
4. Colin Davis
5. Matthew Lau
6. Francisco Vasconcelos
7. Helen C. O'Neill
This article has no evaluationsLatest version Dec 22, 2025
Impact of paternal age in donor oocyte cycles: An Artificial Intelligence-Based Analysis

This article has 10 authors:
1. Cufré Barbieri Magali
2. Josefina Ormachea
3. María Florencia Veiga
4. Macarena Felici
5. Ana María Biagini
6. Daniela Chavez
7. Santiago Giordana
8. Mariana Weigel Muñoz
9. Fernando Neuspiller
10. Natalia Basile
This article has no evaluationsLatest version Dec 16, 2025
An Overview of Existing Applications of Artificial Intelligence in Histopathological Diagnostics of Lymphoma: A Scoping Review

This article has 7 authors:
1. Mieszko Czapliński
2. Grzegorz Redlarski
3. Mateusz Wieczorek
4. Paweł Kowalski
5. Piotr Mateusz Tojza
6. Adam Sikorski
7. Arkadiusz Żak
This article has no evaluationsLatest version Jan 16, 2026

Discuss this preprint

Listed in

Abstract

Background

Methods

Results

Conclusions

Article activity feed

Related articles

Embryo selection tools in IVF can favour XY embryos: implications for equitable reproductive AI

Impact of paternal age in donor oocyte cycles: An Artificial Intelligence-Based Analysis

An Overview of Existing Applications of Artificial Intelligence in Histopathological Diagnostics of Lymphoma: A Scoping Review