Community Discussion · Policy

The Awkward Silence at a Demo Event

Professional BuzzkillProfessional BuzzkillJul 112026/07/11 89 views

Last week, I attended an AI health startup pitch event. On stage, a founder excitedly showcased a "cardiac risk warning" app based on wearable devices. He cited Google's newly released SensorFM model, saying, "Trained on 5 million people and 1 trillion minutes of data, surpassing baselines in 34 out of 35 tasks—this is the future." Investors in the audience nodded, but I couldn't help listing the details that weren't mentioned in my head.

The model scale is indeed impressive. Wearable data from 5 million participants, spanning from September 2024 to now, nearly a year. The pre-training corpus covers sensor signals like heart rate, activity, sleep, and blood oxygen. Google claims superiority over traditional feature-engineered supervised baselines in 34 out of 35 health tasks, including sleep stage classification, activity recognition, and heart rate anomaly detection. At first glance, this seems like the prototype of an all-around health AI.

But a careful read of the technical blog reveals that the so-called "superiority over baselines" comes with many hidden premises.

First, data representativeness is a fatal flaw. The 5 million participants come from around the world, but the prerequisite is that they own devices compatible with the Google wearable ecosystem—likely Fitbit series or Wear OS watches. This means the sample naturally skews toward middle-to-high income groups, people accustomed to wearing smart devices, and those willing to actively authorize data sharing. Those who most need health monitoring—the elderly, low-income groups, and chronic disease patients—are often not in this pool. Models pre-trained on such data will suffer performance collapse due to distribution shift once deployed in real-world medical scenarios.

Second, task definition itself is a self-fulfilling prophecy. Most of the 35 tasks are sensor signal classification or regression, such as "determining if a user is walking from accelerometer data." These tasks have mature feature engineering in lab environments, and deep learning models naturally fit noise through massive parameters. But what about truly critical clinical tasks? Such as atrial fibrillation prediction, screening for occult hypertension, or grading severity of sleep apnea—tasks requiring gold standards (like ECG, sphygmomanometers, polysomnography) as supervision signals. Is SensorFM actually effective for these? Google didn't publish a detailed task list, just stating "34 items superior to baselines," which is too vague.

A more core issue is that wearable device sensor accuracy itself is unreliable. PPG heart rate measurement on watches can have errors exceeding 20% during exercise, blood oxygen detection has systematic bias in dark-skinned populations, and sleep staging algorithms often count lying awake staring as light sleep. Foundation models trained on this noisy data, even if performing well on test sets, are merely closed-loop self-consistent within the specific subset of "watch wearers." In real-world scenarios, differences in devices, wearing positions, and individual variations will cause model failure. Google's technical report mentions using "consistency filtering" to clean data—but while filtering out anomalies, they also filtered out the most valuable long-tail phenomena in the real world.

Don't forget the wall of regulation. Software as a Medical Device (SaMD) requires FDA or CE certification, and SensorFM is currently just a research model, not claiming use for clinical decision-making. But many startups are already starting to use it as a core selling point for fundraising, bringing us back to that pitch event I attended. Investors only see numbers like "5 million" and "1 trillion," but fail to see issues like data privacy compliance, device compatibility, and doctor trust when deploying models in real hospitals. Google itself hasn't promised any implementation timeline; they just published a paper and a demo.

Finally, this isn't a technical problem, but a business logic problem. The wearable health market is already overheated. Apple Watch's AFib monitoring function is still not used by mainstream hospitals as a diagnostic basis, and Fitbit's sleep score is mostly psychological comfort. If SensorFM is truly effective, why doesn't Google release an API for developers to call directly? Why not partner with Mayo Clinic for real-world validation? The answer might be: its current performance is insufficient to reproduce results in strictly controlled clinical environments. The "superiority over baselines" in those 34 tasks may only be relative to the simplest feature engineering baselines, not clinical gold standards.

Oh right, what was the only task where it lost? Google didn't say. If that one task among the 35 was atrial fibrillation detection or blood pressure estimation, the practical value of this model would be discounted by at least half.

SensorFM is a decent academic benchmark, kicking off pre-training on wearable data. But expecting it to change health management on your watch next year? Don't hold your breath.


Original Link: https://www.ithome.com/0/975/442.htm

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts