Educate - Engage - Empower

Fertility Appreciation Collaborative to Teach The Science

August 17, 2026

Decoding Fertility with AI: How Do Large Language Models Perform?

By Tammy Tran, MD

Editor’s Note: With rapid advances and expansion in artificial intelligence (AI), this review of the role of AI in fertility is timely. Tammy Tran, MD, recently wrote this article as a fourth-year medical student at Penn State College of Medicine while participating in the FACTS elective. She is now a family medicine resident physician at Inova Fairfax Hospital in Virginia. Dr. Tran reviews the research of Grace et al. in Ctrl + Alt + Conceive: Fertility Awareness in the Age of Artificial Intelligence, How Do Large Language Models Compare?” The study examines how well generative AI platforms answer questions about fertility and reproductive health. As fertility care algorithms evolve with emerging evidence, it raises an important question: Can AI reliably keep pace?

Introduction

The exponential growth of generative artificial intelligence (GenAI), such as ChatGPT, is transforming the advancement of medical knowledge through its ability to analyze and generate new content in a human-like manner.[1] In the past, people have frequently used online digital sources, such as Google, to seek medical information or advice because of the speedy access to information. As new artificial intelligence platforms emerge, people may increasingly turn to these tools for health-related information and guidance. However, artificial intelligence carries risks of inaccuracies and false information. In the context of medical decision-making, it is important to highlight accuracy, reliability, and the potential limitations of using generative artificial intelligence. Research by Grace et al. [2] aims to compare fertility and reproductive health information generated via different GenAI platforms and assess the quality of the content to gauge their potential use as a tool for reproductive health education.

“As new platforms emerge within the realm of artificial intelligence, people may increasingly turn to these tools for health-related information and guidance. However, artificial intelligence carries risks of inaccuracies and the generation of false information.” 

Methodology

In this study, researchers used two well-established questionnaires on fertility knowledge to generate 37 questions about the menstrual cycle (5 questions), conception (7 questions), assisted reproductive technologies (ARTs) (8 questions), age-related fertility decline (8 questions), and the risk factors associated with fertility (9 questions). These questions were converted into prompts and input into each of these four platforms: ChatGPT 4.0 Free, Copilot Free, Gemini 1.0 Free, and Perplexity Free. Each result was assessed based on a Likert scale of 1 to 5 that evaluated concordance (how accurate the answer was), comprehensibility (how easy the response was to understand), and conciseness (how clear and direct the answer was). A score of 5 would indicate high concordance, comprehensibility, and conciseness. A score of 1 would indicate low concordance, comprehensibility, and conciseness.

AdobeStock 382779150

Results

After evaluating 37 prompts and answers, GenAI platforms generally demonstrated high performance in comprehensibility and conciseness compared to concordance. Furthermore, ChatGPT and Copilot showed the highest concordance, Gemini demonstrated moderate concordance, and Perplexity showed the lowest concordance. While comparing performance across fertility and reproductive health topics, these GenAI platforms provided generally concordant or accurate information for questions related to the menstrual cycle, conception, and fertility risk factors. However, concordance was lower for questions related to ARTs.

“While comparing performance across fertility and reproductive health topics, these GenAI platforms provided generally concordant or accurate information for questions related to the menstrual cycle, conception, and fertility risk factors…(but it) was lower for questions related to ARTs.”

For example, one of the questions was, “Around what age does female fertility start to decline?” The questionnaire reference response was “around 30-35 years,” and all four GenAI platforms generated generally concordant responses. ChatGPT responded that “fertility starts to decline around the age of 30, with a more significant decrease after age 35.” Gemini stated that “fertility starts to decline gradually in the early 30s and more rapidly after 35.” Copilot reported that “fertility generally begins to decline at around age 32 and then drops off more dramatically after 37.” Perplexity responded that fertility typically starts to decline in the late 20s with a more significant decrease after age 35.” In contrast, there was varied concordance to responses related to ARTs when all four GenAI platforms were prompted with the question: “When using frozen eggs from women less than 37 years old, what is the live birth rate per thawed egg?” The reference questionnaire response was equal or less than 10%. ChatGPT provided a relatively concordant estimate of approximately 2 to 12%, whereas Gemini and Copilot generated higher estimates of 60 to 70%. Similarly, Perplexity reported “approximately 70% when at least 20 mature eggs are thawed.”

Discussion

This study by Grace et al. [2] found that GenAI platforms generally returned correct answers when prompted with fertility and reproductive health questions about the menstrual cycle, conception, and fertility risk factors. However, content on ARTs was the least accurate. This difference in concordance on ARTs may be due to a much smaller body of resources on ARTs and their fast-changing nature due to advancements in medicine. Additionally, GenAI platforms can only be as reliable as their training datasets. If their datasets lack information and resources and are not consistently being updated, this can increase risk of inaccurate or even fabricated responses. In contrast, this study demonstrated that AI-generated text was indistinguishable from human-generated text in terms of comprehensibility and conciseness and was often more efficient at conveying a message, using fewer redundant words and phrases.

“GenAI platforms generally returned correct answers when prompted with questions about fertility and reproductive health around topics such as menstrual cycle, conception, and fertility risk factors. However, content on ARTs was the least accurate.”

Although GenAI can serve as a valuable tool to expand medical knowledge about fertility and reproductive health, it is important to be reminded of its limitations. For instance, this study points out that the prompts were based on established fertility questionnaires reviewed by researchers. In real-world use, non-experts are more likely to use simpler terms when prompting these GenAI platforms. Differences in phrasing can then lead to different responses. ChatGPT and Gemini have provided their users with warnings of potential inaccuracies; however, the other platforms did not. With the rise of social media, GenAI platforms could potentially cause a storm of fertility misinformation as misinformation can spread six times faster online than accurate evidence-based medical information.

This raises important questions about how GenAI tools should be integrated into healthcare education, how it can accurately continue to update as new research and evidence emerges, and whether there should be safeguards implemented with its usage. Future research should explore how GenAI tools perform when responding to questions asked by people without medical training to understand how these tools could responsibly and accurately improve public understanding of fertility and reproductive health.


REFERENCES

[1] Reddy, S. Generative AI in Healthcare: An implementation science informed translational path on application, integration and governance. Implementation Sci 19 , 27 (2024). https://doi.org/10.1186/s13012-024-01357-9

[2] Grace B, Zhu J, Dudakia H, Ajao-Rotimi F, Colton N. Ctrl + Alt + Conceive: Fertility awareness in the age of Artificial Intelligence, how do large language models compare? Hum Fertil (Camb). 2025 Dec;28(1):2584673. doi: 10.1080/14647273.2025.2584673. Epub 2025 Nov 16. PMID: 41243291.


ABOUT THE AUTHOR

Tammy Tran, MD, is a Family Medicine Resident Physician at Inova Fairfax Hospital in Virginia. She earned her medical degree from Penn State College of Medicine and completed her undergraduate education at VCU in Richmond, VA. She is interested in women’s health and immigrant care. She enrolled in the FACTS elective to better understand fertility awareness-based methods and offer them as an option for future patients.


Inspired by what you read?

You can support the ongoing work of FACTS here. To connect with a member of our team, please email development@FACTSaboutFertility.org. Interested in becoming an individual or organizational member? You can learn more and register here. To discuss with a member of our team, please email membership@FACTSaboutFertility.org.


CME Course Part F FemTech

Search the Blog

By Lynnea Nicholls, DO Editor’s Note: Lynnea Nicholls, DO, wrote this article during her fourth year of medical school while participating in the...

By Lynnea Nicholls, DO Director’s Note: Dr. Lynnea Nicholls wrote this review as part of the FACTS elective during her fourth year of...

By Molly Franzonello Editor’s Note: Dr. Lisa Gilbert, MD, MA (Ethics), FAAFP, is a board-certified family physician, educator, and inaugural fellow in the...

0
    0
    Your Cart
    Your cart is emptyReturn to Shop

    Join Our Mailing List

    Stay connected with timely news, blog postings, and upcoming events with FACTS.