A wrist tracker displays steps, heart rate, calories and sleep stages as though all four were measured. Only some are measured; the rest are inferred by models, and knowing which is which explains most of the disagreement between devices.

Motion is measured, steps are inferred

The primary sensor is an accelerometer measuring acceleration along three axes many times per second. That raw signal is genuine measurement.

A step is not directly present in that signal. Software identifies a repeating pattern that matches walking and counts the peaks.

Detection thresholds are choices. Set them sensitively and pushing a trolley registers steps; set them conservatively and slow walking is missed.

This is why two trackers on the same person disagree, sometimes considerably, and why counts differ between wrist and pocket placement for the same walk.

Manufacturers tune thresholds against reference recordings of people walking, so the algorithm performs best on the gait patterns present in that reference set.

Optical heart rate reads blood volume, not beats

Wrist heart rate uses green light emitted into the skin and a photodetector measuring how much returns. Blood absorbs the light, so reflection varies with each pulse.

The device extracts a periodic component from that fluctuating signal and reports its frequency as heart rate.

The signal is small and easily disturbed. Wrist movement, loose fit, cold hands and tattoo ink all degrade it.

During steady activity the method works well. During activity with rapid changes and heavy wrist movement, error rises and the device may lock onto the cadence of motion instead of the pulse.

Calorie figures are model outputs

No consumer wearable measures energy expenditure. The figure shown is produced by an equation using height, weight, age, sex, heart rate and movement.

The equation estimates a resting rate from body characteristics and adds an activity component derived from the sensor data.

Individual variation around such estimates is wide, because two people with identical inputs can differ substantially in actual expenditure.

The number is therefore most useful as a relative signal. Comparing today with last Tuesday on the same device is informative; comparing the absolute figure between devices is not.

GPS traces are smoothed before display

Satellite positioning gives a series of position fixes with error attached to each one. Raw fixes scatter around the true path.

Software filters the sequence, discarding implausible jumps and fitting a smoother line, then computes distance from the filtered track.

Smoothing reduces the distance slightly compared with summing raw fixes, which would inflate it by counting jitter as travel.

In dense cities and under tree cover, reflections off buildings cause larger errors, and different manufacturers handle this differently, which produces divergent distances for the same route.

Sleep staging is an inference from two signals

Trackers infer sleep from movement and heart rate variation. They do not measure brain activity, which is what clinical staging is based on.

Distinguishing sleep from lying still is reasonably tractable. Distinguishing between stages of sleep from the same two signals is considerably harder.

Manufacturers train classifiers against clinical recordings, and agreement is better for total sleep time than for stage breakdown.

The practical consequence is that nightly stage percentages should be read as a rough model output rather than a measurement, particularly when comparing across devices.

Total time asleep and consistency of timing are the more dependable outputs, and they are also the ones least often emphasised in the interface.

Firmware updates change historical comparisons

Algorithms are updated over the life of a device, and an update can change how the same underlying data is interpreted.

A user may see a step in their long-term chart on the date of an update, reflecting a change in the model rather than in their behaviour.

Most manufacturers do not reprocess historical data, so a multi-year chart can contain several different algorithms stitched together.

This is worth knowing before drawing conclusions from long time series, particularly where the trend is small relative to the discontinuity.

Why devices disagree with each other so visibly

Each manufacturer chooses its own sensors, sampling rates, thresholds and models, and none of these are standardised across the industry.

There is no shared reference that devices are calibrated against, unlike a bathroom scale that can be checked with a known mass.

Independent comparisons consistently find better agreement on heart rate during steady effort than on energy expenditure or sleep staging, which follows directly from what is measured versus modelled.

Users encountering this generally assume one device is faulty. More often both are working as designed and reporting the output of different assumptions.

What the numbers are actually good for

Trend and consistency are the durable outputs. A device that is systematically wrong in the same direction still shows change reliably.

Prompts based on the device's own history, such as noticing an unusually inactive day, work regardless of absolute accuracy.

Absolute figures are the weakest use, and they are also the ones the interface presents most prominently because they are easy to display.

Reading a tracker as a consistent but uncalibrated instrument, rather than a precise one, aligns expectations with what the hardware can actually do.