Does it actually change what people believe?
Does a buzz on the wrist change how people process information. To test this rigorously, we built a Wizard-of-Oz version of FactNudger with a fixed, known accuracy (87.5%, in line with state-of-the-art video fact-checkers) so we could isolate the effect of real-time feedback itself from the technical performance of any one AI model.
We ran a within-subjects study with 34 participants, who each watched short video compilations of true and false claims: once while wearing FactNudger, and once with no support. Order was counterbalanced, and 1–2 deliberate AI errors were seeded into each video to see how people responded to an imperfect system.
We found that:
Belief in false claims dropped significantly when wearing the device — and belief in true claims was preserved, meaning the nudge didn't make people indiscriminately more skeptical.
Confidence in rejecting false claims went up, not just belief accuracy.
Verification activity roughly doubled (pauses, rewinds, web searches, watch glances) compared to the no-wearable condition, with no added subjective cognitive workload (NASA-TLX showed no significant difference).
The effect held up consistently across participants regardless of how much they personally cared about a topic, their prior exposure to the claim, or their general disposition toward open-minded thinking.
The catch: errors matter, and not symmetrically. A false positive, i.e. the watch flagging a true claim as false, made people significantly more skeptical of accurate information. A false negative, i.e. silently missing a false claim, had much less effect, since people mostly fell back on their own instincts. In other words, people leaned on the device more than they should have when it spoke up, but weren't badly hurt when it stayed quiet. That's a useful asymmetry for design, but it also means trust calibration (confidence scores, visible sources, honest uncertainty) should be central to using this kind of system responsibly.