中文

Research Case Study

Case Study: Wearables and Irregularity

This page gives a detailed, teaching-oriented summary of the workshop paper Menstrual Cycle Prediction with Wearable and Personal Historical Information: Heterogeneous Effects Across Irregularity Profiles: research question, protocol, compared methods, subgroup findings, and limitations.

What was the core question?

For daily menstrual-event prediction, can wearable physiology provide meaningful value beyond calendar and personal history baselines, and does that value depend on irregularity profile rather than a single regular-versus-irregular split?

Dataset and protocol at a glance

Data scope

The study uses mcPHASES, a multimodal menstrual-health dataset with wearable, hormonal, and self-reported observations. The processed cohort includes 111 cycles from 40 users; 79 labelled cycles from 37 users are used for labelled ovulation and next-menses analyses.

Operational constraint

Prediction is daily and prospective: on cycle day d, only information available up to day d can be used. Population wearable models are evaluated in a leave-one-subject-out setting.

Irregularity is analyzed along two axes: mean cycle length (short / typical / long) and cycle-to-cycle variability (low / medium / high), plus stable-length intersections (for example, long-stable). Subgroup labels are used for analysis only, not as model inputs.

What exactly was compared?

Method Overall MAE (days) Within +/- 3 days
Calendar 3.843 0.557
History-only 4.492 0.471
Wearable reference 1.850 0.757
Wearable + History ACL 2.152 0.686
Wearable + history-physiology refinement 2.216 0.714

Headline result: the fixed wearable reference clearly outperforms both static references at aggregate level, while added countdown-stage history is not a uniform overall gain.

What did subgroup analysis show?

Large gains in long and high-variability profiles

Against History-only, MAE drops from 10.286 to 1.325 days in the long subgroup and from 10.430 to 1.087 days in the high-variability subgroup.

Strong benefit also appears in long-stable subgroup

In long-stable cycles, MAE drops from 8.000 to 1.019 days, supporting the value of current-cycle physiology in this profile.

Counterexample: short-stable profile

History-only remains better than wearable reference (2.750 versus 4.202 days), showing that some shifted-but-stable regimes can still favor simple history.

History extensions are selective, not universal

Both countdown variants improve the high-variability subgroup (to 0.945 and 0.872 days), but do not deliver broad improvement across all stable or typical profiles.

What still needs stronger evidence?

Use this case as a concrete anchor for the Field Guide's open-question framework: irregularity taxonomy, wearable value boundaries, label transparency, and deployment realism.

After this page