Shrinking Large Models: An Inevitable Step for Engineering Implementation, Not Regression
Community Discussion · Policy

Shrinking Large Models: An Inevitable Step for Engineering Implementation, Not Regression

ZhulongZhulongJul 222026/07/22 95 views

I noticed an interesting detail: last year at WAIC, everyone was still competing over hundreds of billions or trillions of parameters, but this year the wind has suddenly shifted. Everyone is pushing lightweight versions, even pruning models so they can run on mobile phones. This change actually happened long ago in our field of intelligent driving—back in 2022, autonomous driving solutions were still bragging about LiDAR line counts and chip computing power in TOPS. By 2024, who is still talking about these numbers? Everyone is discussing whether "end-to-end models can run stably on automotive-grade chips" and whether "power consumption can be kept under 30 watts."

1 replies

?
Ctrl + Enter to reply
xiafeng
xiafengJul 30(edited)

[quote="zhulong, post:1, topic:1397"]

I noticed an interesting detail: last year at WAIC, everyone was competing over hundreds of billions and trillions of parameters, but this year the trend suddenly shifted. Everyone is pushing lightweight versions, even pruning models to run on mobile devices. This change actually happened earlier in our intelligent driving field—back in 2022, autonomous driving solutions were still boasting about LiDAR line counts and chip TOPS. By 2024, who was still talking about these numbers? The focus had shifted to "whether end-to-end models can run stably on automotive-grade chips" and "whether power consumption can be kept under 30 watts."

Essentially, large mo…

[/quote]

This trend is also obvious in the open-source community. For example, several TinyML projects on GitHub emphasize quantization toolchains and cross-compilation support in their contribution guides. Stacking parameters is for leaderboard chasing; only models that can actually run on Raspberry Pi attract iterative development.