What actually breaks it?
Two earlier recordings disagreed about whether movement breaks the measurement, because nobody had measured how much movement there was. So the test was rebuilt: five conditions, each recorded as still, then the condition, then still again, so every one carries its own baseline — and the movement was measured, not described.
Error while sitting perfectly still - nothing is moving.
Error while sitting perfectly still - nothing is moving.
PHASE-Net during slow turns. About twice the resting error, and nothing gets worse than this.
Triple the movement, and the error changes by 0.1 BPM.
Movement doubles the error, then stops mattering
Each dot is one condition. Movement is how much the picture inside the face crop changes between frames, above what it changes while still.
The curve lifts off the resting floor to roughly double it, then flattens. Going from slow to fast head turns triples the measured movement and changes nothing, so how much someone moves is not what sets the ceiling.
Which conditions differ from doing nothing?
The same numbers as bars. Each dashed line is that method's error while still, so a bar level with its own line added nothing measurable.
Talking sits on the line. Only turning the head lifts either method clearly above its own floor — and both rise together, so neither is the robust one.
Every condition, side by side
Resting rate comes from the still half of the same recording. Drift is how far the reading strayed from that baseline.
| Condition | Movement | Face moved | of face | PHASE-Net rest | during | drift | POS rest | during | drift |
|---|---|---|---|---|---|---|---|---|---|
still nothing at all - the control | 0.00 | 8 px | 6 % | 80 | 77 | 8.0 | 80 | 75 | 6.8 |
talking speak, head kept still | 0.29 | 9 px | 7 % | 82 | 82 | 7.8 | 66 | 67 | 5.9 |
mixed turn, nod and lean at once | 0.55 | 28 px | 22 % | 79 | 81 | 10.5 | 78 | 63 | 13.7 |
slow turns turn left and right, one cycle per 4 s | 1.04 | 48 px | 38 % | 80 | 96 | 15.5 | 77 | 65 | 12.9 |
fast turns the same turn, one cycle per 1.5 s | 2.88 | 37 px | 29 % | 81 | 72 | 15.4 | 71 | 64 | 14.1 |
The face never travels far: the largest excursion anywhere is 38 % of a face width, so it stays inside the crop throughout. And even in the control, single readings from PHASE-Net span 68–108 BPM while nothing at all is happening.
Three explanations, all rejected
each was turned into a test that could kill itThe face leaves the fixed crop under motion
One recording processed twice - fixed box versus re-detecting the face every second - so the crop strategy is the only variable.
- Furthest the face ever moved
- 21 % of a face width
- Does error track the face moving?
- 0.07 out of 1.00 - no link
Tracking made it worse: a re-detected box jitters between frames and injects motion of its own. A slightly misaligned but stable crop beats a well-centred jittery one.
Facial appearance change (speech, expression) breaks it
A 'talk' condition - the face changes continuously but barely moves.
- Face moved while talking
- 8.8 px
Talking is indistinguishable from doing nothing.
One method is inherently more motion-robust
Dose-response across five measured motion levels.
Both degrade by a similar factor and the effect saturates - tripling the motion from slow to fast changes nothing.
What survived: the twenty-fold gap
The same model is about 0.39 BPM off on controlled recordings and about 8 BPM off on this webcam before anyone moves — roughly 21× worse. Movement adds a further doubling on top of that, and then stops.
So the bottleneck is the camera, the lighting, the distance and how much of the frame the face fills — not movement, and not the model. A better capture setup is worth more than a more motion-tolerant architecture. The controlled-data figure →