Back to list
Multicenter IVF clinic study linking Gardner grading harmonization by consensus-trained AI with improved pregnancy outcomes
Jul 08, 2026
원문 보기
Authors
S. Ramakrishnan , R. Pandian , R. K , S. Doddamaneni , R. Gutgutia , S. Malik , H.M. Kim , J. Kang , H.J. Lee
Conferences
ESHRE
ABSTRACT


Study question

Does clinical integration of a consensus-trained AI system improve grading concordance and pregnancy outcomes over time in a multicenter IVF setting?



Summary answer

Progressive alignment between embryologists and AI grading was observed after implementation and was accompanied by an increase in fetal heart tone pregnancy rates.



What is known already

Gardner grading is the most widely used system for blastocyst assessment but is associated with substantial inter- and intra-observer variability, particularly for inner cell mass (ICM) and trophectoderm (TE). Artificial intelligence systems can generate consistent embryo assessments; however, prior studies have largely focused on predictive performance rather than how AI integration influences embryologist grading behaviour in routine clinical practice. Evidence remains limited on whether AI-driven standardization of grading, particularly using static embryo images combined with clinical variables, translates into improved downstream pregnancy outcomes across clinics.



Study design, size, duration

This multicenter interrupted time-series study was conducted in 2025 and included 11,604 Day 5 blastocyst evaluations collected between July and December. In October, AI-generated outputs were intentionally integrated into the clinical grading interface, allowing embryologists to view AI assessments during routine embryo evaluation and enabling comparison of grading behaviour and pregnancy outcomes before and after this intervention.



Participants/materials, setting, methods

Blastocysts were graded by embryologists using conventional Gardner criteria for developmental stage, inner cell mass (ICM), and trophectoderm (TE) based on Day 5 static microscope images. The AI system uses Day 5 static embryo images combined with maternal age to generate Gardner grades and predicted probabilities of fetal heart tone (FHT) pregnancy. The AI grading model was trained on consensus Gardner grades established by embryologists with over 20 years of experience.



Main results and the role of chance

Overall concordance between embryologists and AI increased after clinical implementation of AI-based grading. Cohen’s kappa improved from 0.29 to 0.59 for developmental stage, from 0.17 to 0.47 for inner cell mass (ICM), and from 0.12 to 0.39 for trophectoderm (TE) between the pre-implementation (July–September) and post-implementation (October–December) periods.

Generalized linear models showed an immediate post-implementation decrease in agreement, consistent with an adaptation phase (Stage = −3.63; ICM = −1.63; TE = −2.41; all p < 0.001), with no significant pre-intervention trends. Post-implementation time interactions were positive across all components (Stage = 1.27; ICM = 0.53; TE = 0.75; all p ≤ 0.001), indicating progressive harmonization of grading behaviour over time.

To minimize potential bias from clinics with very small sample sizes, pregnancy outcome analyses were restricted to IVF clinics that had submitted at least 20 cycles to the Vita Embryo system. Accordingly, eight clinics comprising a total of 217 cycles were included in the final analysis. In these clinics, fetal heart tone (FHT) pregnancy rates increased from 52.9% (73/138) during the pre-implementation period to 59.5% (47/79) after AI integration, corresponding to an absolute increase of 6.6 percentage points.



Limitations, reasons for caution

Although not a randomized trial, the study incorporated a clearly defined intervention and an interrupted time-series design. Residual confounding and concurrent temporal changes cannot be fully excluded, and causality between grading harmonization and improved pregnancy outcomes cannot be definitively established.



Wider implications of the findings

Clinical integration of an AI system may promote grading harmonization across clinics and be accompanied by improved pregnancy outcomes, supporting AI as a standardized reference that complements, rather than replaces, expert embryologist judgement.


원문 보기