For app developers
- Read the task families and evaluation pitfalls in the Field Guide.
- Then go to Data & Code for starter paths and dataset tradeoffs.
- Finish with Current Research to understand what you should not overpromise.
Open Educational Resource
This OER is built for cross-disciplinary readers. Instead of dropping every visitor into the same wall of papers and terminology, it first helps you locate yourself: are you building an app, approaching the field from AI, or entering from medicine?
I need to know what should actually be predicted, which route is realistic first, what public data can support an MVP, and where product claims become scientifically unsafe.
Persona 02I know how to build models, but I need the physiological constraints, label realities, task definitions, and the reasons papers in this field are often not directly comparable.
Persona 03I understand the cycle clinically, but I need to quickly understand data processing, model families, subject-wise splits, and what AI papers are actually comparing.
Real-world cycle length, ovulation timing, and variability differ widely. Fixed templates will hide biology inside apparent model error.
In this field, label definition, splits, and deployment constraints often matter more than architecture choice.
Temperature, heart rate, and HRV are valuable proxies, but they vary with device, wear site, and cohort.
Regular cycles, irregular cycles, long cycles, and highly variable cycles may be fundamentally different prediction problems.
Scalable and app-friendly with low user burden, but fragile for irregular cycles and sparse history.
Can expose biological structure beyond dates alone, but carries noise, device differences, and weak-label problems.
Closest to biological ground truth, but costly, smaller in scale, and harder to map onto daily deployment.
Many studies mix wearable and clinical routes: clinical signals define labels, but those labels are unavailable at deployment time.
The goal is no longer only to improve next-date prediction by a small average margin. The stronger goal is to understand different cycle dynamics, identify under-served groups, and support more specific forms of personalization.
Real-world cycles vary substantially in length, follicular timing, and ovulation timing, so one summary score can hide who the model actually works for [1] [2].
Short-stable, long, and highly variable cycles are not the same prediction problem. Subgroup-aware modelling can be more meaningful than a single pooled model [2].
Temperature, heart rate, and HRV matter not only for prediction gains, but because they can reveal structure that date history alone cannot [5].
A stronger method is not just one with lower average error, but one that explains where failure happens and for whom [4].
Many methods are developed or reported on relatively regular cohorts, so performance often looks better than it is for the people most likely to need support.
Studies define ovulation differently, report different windows, and use different metrics, which makes cross-paper comparison weak.
The most influential app and device datasets are often private, so public data are still mainly used for baselines and method validation.
A model that works retrospectively may still fail when limited to information that would exist in real product use.