Multi-Sensor Fusion: I'm Betting on Waymo
Just saw that interview with Waymo CEO Dolgov. Finally, someone laid bare the trump cards of the pure vision approach.
I used Zoox for two weeks. On my daily commute, there's a segment exiting an underground garage onto ground level, where light shifts drastically from dark to bright. In that moment, cameras are basically blind, but Zoox's LiDAR still clearly perceives the curb and pedestrians ahead. This isn't theoretical deduction; it's a tangible experience.
What Dolgov said about "weak perception solutions hitting a safety ceiling soon" translates to quantitative language as: the Sharpe ratio of pure vision schemes drops sharply outside certain confidence intervals. Backtests look great, but it collapses in edge cases. As a trader, this is what I fear most—models fit perfectly on historical data but blow up in live trading.
Tesla's pure vision route essentially bets on an assumption: human drivers can drive with just two eyeballs, so cameras can too. But there's a fatal flaw here—human brain perception isn't solely visual input; vestibular sense, proprioception, and even hearing participate in driving decisions. When you hear tire noise while driving, you subconsciously slow down to check, but cameras only see images; they can't hear sound. This isn't a sensor count issue; it's a perceptual modality completeness issue.
Waymo's three-sensor fusion is basically doing a multi-factor strategy. LiDAR, millimeter-wave radar, and cameras each have their own failure modes, but the probability of all three failing simultaneously is extremely low. This redundancy design is called "tail risk hedging" in quant finance—it doesn't matter if you see the black swan or not; what matters is that you don't die when the black swan arrives.
One point Dolgov didn't explicitly state but hinted at clearly: Pure vision is enough to match human-level performance, but cannot significantly exceed humans. And the ultimate goal of L4 is precisely to "significantly exceed humans." If you just want L2 assisted driving, pure vision is enough. But if you want the car to run autonomously without a safety driver, multi-sensor fusion is the only verified path.
I'm curious if Tesla will secretly add LiDAR to the Cybertruck. After all, that body structure creates massive visual blind spots.
Physix Frontier