Skip to content

Body

I Wore Two Sleep Trackers for a Month and Ignored the Grades

An Oura Ring, a Fitbit and a paper diary agreed on broad patterns but fought over individual nights. The scores were most useful when treated as delayed evidence, not morning instructions.

Nour HaddadBody — Drugs & Harm Reduction

August 11, 2026 · 8 min read

An Oura Ring, a Fitbit and handwritten sleep diary cards arranged on a bedside table.
An Oura Ring, a Fitbit and handwritten sleep diary cards arranged on a bedside table.

For a month, I slept with an Oura Ring on one hand and a Fitbit on my wrist. The more important device sat beside the bed: a white index card and a black pen.

Each morning, before opening either app, I wrote down when I thought I had fallen asleep, whether I remembered waking, how rested I felt and anything obvious about the previous evening. The card took less than two minutes. Only after that did I look at the competing grades generated overnight.

The rule was blunt. I could observe the scores, but I could not alter the next day around them. No canceling plans after a low grade. No earlier bedtime to satisfy an app.

No extra caffeine because a dashboard said recovery was poor. I kept my ordinary routine, including its inconsistencies, and let the trackers watch.

That refusal mattered more than wearing two devices. Consumer sleep tracking is sold as measurement, but its interface is built for behavioral management. A score appears in the morning, colored and ranked, followed by suggestions that turn an estimate about last night into a set of obligations for today. The product does not merely describe sleep.

It attempts to govern the person who woke up.

This is reporting from a hands-on evaluation, not medical or mental health advice. Consumer wearables cannot diagnose a sleep disorder or rule one out.

The grade hides the guesswork

Neither device watched me sleep in the clinical sense. Polysomnography, the testing used in sleep laboratories, records signals including brain activity, eye movement and muscle activity. A ring or wrist tracker works with a narrower set of inputs, such as motion, heart rate and changes in the timing between heartbeats, then uses a proprietary model to infer when sleep began and which stage followed.

“Infer” is doing serious work there.

The Oura Ring and Fitbit often recognized the broad shape of a night. They could usually distinguish a long sleep opportunity from a short one, and both were better at identifying schedule drift across several nights than at explaining one difficult morning. Their agreement weakened around the details. A wakeful stretch recorded on the index card might become light sleep in one app, wake time in the other, or disappear into a smooth chart that looked more decisive than the night felt.

Sleep staging was even less useful as a daily judgment. Sleep staging means dividing the night into categories such as light, deep and rapid eye movement sleep. The apps displayed those categories with clinical-looking precision, but two devices attached to the same body could distribute the night differently. That disagreement did not prove either device worthless.

It showed the limit of the claim.

The final score conceals those limits. It compresses several estimated signals into one clean number, applies thresholds chosen by the company and presents the result with the authority of a test grade. The user sees certainty. Underneath sits a chain of sensor noise, missing context and algorithmic judgment.

This compression is useful for a product. A chart requires interpretation. A grade tells you how to feel.

The index card got there first

The manual diary changed the order of authority. Before the apps could label the night, I had to describe it.

That small sequence exposed a recurring mismatch. Some mornings felt ordinary until a low score supplied a reason to inspect the body for fatigue. On others, I felt depleted while one tracker offered a reassuring result. The score did not create the sensation, but it could recruit attention toward it, giving a passing heaviness the status of evidence or asking me to distrust an unpleasant morning because the wearable had approved the night.

This feedback loop has a name in sleep medicine discussions: orthosomnia, an unhealthy fixation on optimizing wearable sleep data that can itself interfere with sleep. The mechanism is not mysterious. Sleep requires reduced vigilance, while tracking can make vigilance feel responsible. A person starts monitoring bedtime, checking the clock during awakenings and reviewing the morning chart for errors.

The attempt to control sleep adds another task to the bed.

Refusing to react nightly interrupted that loop. A poor grade could not earn an emergency bedtime. A favorable one could not overrule the index card. I still saw the score, which meant I was not immune to its mood, but delaying the app until after the diary preserved one account of the night that had not yet been edited by software.

The cards also accumulated without trying to entertain me. Laid together, they made recurring patterns easier to see: schedule changes mattered more than the apps’ stage-by-stage drama, and several similar mornings carried more weight than one unusually colored dashboard. The trackers became more credible when their claims repeated across time and matched something outside their own systems.

That is the useful scale of consumer sleep tracking. It is better at revealing tendencies than issuing verdicts.

A score needs a return visit

The wellness industry has learned that measurement creates its own demand. Hardware brings the first payment. Memberships, premium interpretation and replacement devices can extend the relationship, while regular notifications keep the product present after the novelty of the sensor fades.

Oura builds recurring membership into much of its experience. Fitbit, owned by Google, places its tracker inside a broader account and service ecosystem, with additional analysis available through premium features. The exact commercial arrangement differs, but the incentive is shared: a user who checks every morning is more valuable than one who reviews a month of data, understands a pattern and closes the app.

Nightly scoring supports that return visit. It refreshes constantly, cannot be completed and carries enough variation to remain interesting. Even a stable sleeper receives a new result. The score therefore operates like a tiny content feed produced by the body, with the company ranking the material and the user supplying attention before getting out of bed.

The recommendations complete the loop. A low score prompts an earlier night, a gentler day or another behavior framed as recovery. If the next score improves, the system appears to have worked. If it falls, the user has more reason to keep monitoring.

The device rarely has to admit that ordinary biological variation, an imperfect fit or an incorrect estimate may explain the movement.

This does not make every recommendation bad. Regular sleep timing and enough opportunity to sleep are hardly corporate inventions. The problem is the transfer of authority from a broad, familiar principle to a proprietary nightly grade whose calculation the wearer cannot inspect.

Measurement clarified less than repetition did

By the end of the month, neither tracker had won. That was the point.

The useful findings appeared where the index cards and both devices converged over repeated nights. Bedtime drift showed up. Short sleep opportunities remained short regardless of how the apps divided their stages. A single reading mattered less once it sat beside several mornings that did not repeat it.

The least useful material was the most visually polished. Precise stage charts invited interpretation beyond what the sensors could support. Readiness language encouraged a decision about the day before the day had supplied any evidence of its own. The nightly grade made uncertainty legible by deleting it.

There is also a basic harm-reduction issue here. A reassuring wearable result should not cancel persistent symptoms such as severe daytime sleepiness, repeated breathing interruptions noticed by someone else or ongoing difficulty functioning. A bad score should not be treated as a diagnosis. The devices occupy an awkward middle ground: intimate enough to influence behavior, limited enough that their confidence needs resisting.

The best alternative was already on the nightstand. The index card stored less information, produced no graph and made no claim about deep sleep. It also did not congratulate me, warn me or ask for another month of attention. Its value came from recording experience before a platform ranked it.

I kept the cards. The scores stayed in their apps.

Questions people ask

Are consumer sleep trackers accurate enough to be useful?

They can be useful for broad patterns such as sleep timing, duration estimates and changes that repeat across many nights. Their estimates of awakenings and sleep stages are less authoritative than the interfaces suggest, particularly when two devices disagree. A consumer tracker is not the same as a clinical sleep study.

Can checking a sleep score make sleep anxiety worse?

Yes. A nightly grade can direct attention toward fatigue, encourage repeated checking and make sleep feel like a performance target. People already prone to anxiety or compulsive tracking may find the feedback loop especially difficult. Turning off notifications or stopping tracking remains a legitimate response, not a failure of discipline.

Why keep a manual sleep diary with a tracker?

A diary records your account before the app frames the night. Writing down perceived sleep, remembered waking and morning function creates an independent reference, making it easier to notice when the device identifies a repeated pattern and when its score conflicts with lived experience.

Should a low sleep score change the next day’s plans?

One consumer score carries too much uncertainty to function as an order. Repeated patterns may justify closer attention, while persistent or severe symptoms belong in a conversation with a qualified professional rather than an app. This experiment did not test medical treatment and should not be read as professional advice.

Was this worth your time?
ShareFacebook
wellness industrymental healthsleep trackingwearable techmental healthwellness industry

One update a day

Today's story, in your inbox

One story each morning — no hype, no filler, no algorithm deciding for you.

Read next