From Prediction to Accountable Action: Designing Uncertainty in Femtech AI

from-prediction-to-accountable-action:-designing-uncertainty-in-femtech-ai

Source: Unite.AI

In women’s health, accuracy is necessary, but it is only the first layer of product safety. It’s important for the system to know when a prediction is strong enough to act on and is built to say so when it isn’t.

Open most fertility or cycle-tracking apps, and you’ll see a clean result: ovulation on day 15, a high-fertility label, a confidence score of 78%. The number looks precise. The biology underneath it is not.

A period date is something the user observed. Ovulation is a latent event no consumer device measures directly. It is inferred from proxies. Skin temperature reflects not only progesterone but also sleep, alcohol, illness, ambient conditions, and where the sensor sat that night. A symptom entry blends physiology with perception, memory, and the user’s decision to log it at all. By the time all of that resolves into “Day 15, 78%,” several different kinds of uncertainty have been quietly compressed into one confident-sounding sentence.

In the products I’ve worked on across women’s health, the hard problem isn’t accuracy. The pattern I keep returning to is decision integrity. Accuracy tells you how well a model predicts. It says nothing about whether the product knows when a prediction is strong enough to act on. In a domain where the same output can inform casual planning or a contraceptive decision, that gap is where trust is won or lost.

So my argument is that femtech AI doesn’t primarily need better prediction. It needs better uncertainty design. And uncertainty design isn’t a disclaimer bolted on before launch; it’s an architecture. I structure it into six layers that carry a signal from raw input to an action the system can justify and audit. I apply a six-layer version of the Calibrated Decision-to-Action Product Method, designed to govern how uncertain health signals are converted into permitted and accountable actions. The method includes data qualification, inference calibration, consequence mapping, control policy, workflow execution, and an accountability loop. Its governing path starts with health data, which leads to decision architecture, and culminates in accountable action.

1. Qualify the Data Before You Trust It

The first layer decides what the system actually knows. Femtech products pull from very different sources. It can include things the user directly observed, such as period dates or subjective reports like pain or mood, as well as wearable features like skin temperature and heart rate variability, and variables the model generated itself. Treating these as interchangeable inputs is the original sin.

Each signal needs provenance the system can read: where it came from, when, how often it’s sampled, what confounds it, and how it relates to the thing you’re actually predicting. Ovulation is the event. Temperature, cervical mucus, LH, and cycle dates are evidence about that event, each with its own lag and its own error.

Missing data deserves particular attention because, in health tracking, it is rarely random. People log more when they’re worried and stop when they feel fine, so a gap can carry as much information as an entry. In a real-world analysis of more than 600,000 ovulatory cycles, ovulation could not be detected in 665,603 of the 1.4 million cycles first considered. Three-quarters of those had valid temperature readings on fewer than half the days in the cycle. Data sufficiency was a major limiting factor in what the algorithm could infer. A product that emits a crisp fertility label in that situation isn’t being confident. It’s manufacturing precision it doesn’t have.

2. Calibrate the Inference to the Evidence

A single confidence score can’t represent what’s really going on, because these systems face several distinct sources of uncertainty at once. These include biological variability in the process itself, measurement quality, missingness from self-tracking, model uncertainty from thin training data, and distribution shift when the current user doesn’t resemble the population the model was validated on.

In that same analysis, only 13% of cycles were exactly 28 days long, and the average follicular phase ran 16.9 days across a range wide enough to make any fixed “day 14” assumption misleading. The Apple Women’s Health Study (AAPL ) found cycle length and within-person variability were associated with age, self-reported ethnicity, and body mass index. Personalization, then, can’t mean swapping a population average for a single personal point. It means producing a distribution that tightens as evidence accumulates.

Consider two AI-based predictions that both read “75%.” One rests on a year of history and dense measurements, and its uncertainty is mostly real biology. The other rests on two cycles, five temperature readings, recent travel, and unknown medication. Here, the uncertainty is mostly due to missing data. Same number. The product should not be allowed to respond to them the same way. This is also why calibration has to be checked within subgroups — aggregate performance can hide a model that’s overconfident, particularly for irregular cycles or perimenopausal users. The clinical-AI literature now separates discrimination, calibration, and decision utility for exactly this reason. A model can rank users well and still produce incorrect probabilities for real decisions.

3. Map the Consequence of Being Wrong

The third layer asks what happens when the system is wrong, and ties the permitted response to that cost. A cycle summary, a fertile-window estimate, a low-fertility label read as contraceptive guidance, and a symptom interpretation that decides whether someone seeks care are not the same product, even when the model behind them is identical.

I sort outputs into four tiers by consequence: informational, behavioral, reproductive, and clinical.  Each gets its own evidence threshold, permitted language, and escalation path. A 70% probability might be fine for guessing when a period will start and nowhere near enough to reassure someone trying not to get pregnant.

This is where design meets regulation. Whether software crosses into medical-device territory turns on its intended use and the role its output plays in a health decision. The FDA’s current Clinical Decision Support Software guidance, updated in January 2026, is explicit that functions intended for patients and caregivers can meet the definition of a device. Regulatory positioning isn’t a legal review at the end. It’s encoded in your thresholds, your wording, and your escalation logic from day one.

4. Turn Uncertainty Into a Control Policy

A blanket “results may be inaccurate” hands the entire interpretation problem back to the user. A control policy does the opposite. For a given evidence state, it decides whether the system shows an observation, offers a bounded range, asks for another measurement, points to a clinician, or declines to answer. Abstention is a product capability, not a failure. “We’re not sure yet” should be a designed state with a next step, not an error screen.

Concretely, instead of “Ovulation: Day 15, 78%,” a governed response reads closer to this: ovulation is most likely within a four-day window; confidence is limited by missing temperature data and disrupted sleep; a few more readings would sharpen the estimate, and an LH test could add prospective evidence and narrow the likely window. The user gets an estimate, the reason it’s uncertain, and the single action that would change it.

5. Support the Action the Output Implies

Even a consumer app creates a workflow: log another reading, repeat a test, keep observing, export data, call a clinician. The system should know which action each response points to, and whether it can support that action safely.

The boundary between product guidance and medical advice lives here, and it isn’t a fixed line. Educational content explaining what a signal generally means sits safely on the guidance side. The product moves into higher-risk territory the moment it interprets one person’s data as disease, directs treatment, or offers reassurance that carries real clinical weight. That boundary has to hold at every interaction, not only in the terms of service.

6. Close the Accountability Loop

The final layer judges the whole system, not just the model. Standard model metrics still matter and focus on discrimination, calibration, subgroup performance, external and temporal validation. But product-level metrics matter just as much. We should ask how often the system had enough input to act, how often it abstained. Plus the one I would insist on is a false-reassurance rate, meaning how frequently it told someone everything looked fine on evidence that couldn’t support the claim.

Every consequential response should be reconstructable after the fact. Regulators, standards bodies, and global health-governance institutions are increasingly adopting this lifecycle view. The IMDRF’s good machine learning practice principles treat trustworthy medical AI as a total-lifecycle responsibility spanning design, deployment, and monitoring, echoing the WHO’s guidance on AI for health.

What the Architecture Looks Like in Practice

Say a product has three months of period dates, six nights of temperature, one positive LH test, a scatter of symptom entries, and a week of bad sleep. The model’s highest ovulation probability lands on day 15. A point-estimate app shows “Ovulation: Day 15, Confidence 78%.”

The architecture produces something else. It marks the temperature data as sparse and sleep-confounded. It generates a distribution across several candidate days rather than one. It applies a gentle threshold for cycle awareness but a stricter one if the output touches pregnancy prevention. It offers a range plus the next useful measurement. And it stores the whole evidence state so the decision can be reconstructed later. Same underlying model and a very different product.

The Right to Act

Femtech systems are only getting more data-rich, with apps, wearables, home tests, medical records, and conversational layers stacked on top. More data can sharpen inference, but it also widens the surface for false precision. The teams that earn trust won’t be the ones with the cleanest-looking confidence score. They’ll be the ones whose products know what they don’t know, and are built to say so.

Accuracy describes the quality of a prediction. Decision integrity decides whether the product has earned the right to act on it.