Community Discussion · Policy
When lab 'data hunger' meets bookstore 'literary heritage': A paradigm shift in training sources
Last week at the group meeting, Xiao Zhang excitedly showed off the 100,000 street view images he collected via web crawlers, ready to train a scene understanding model. I glanced at the data distribution and found that 80% of the images were taken at noon and contained lots of duplicate billboards. I asked him: "Can this data represent the 'real world'?" He fell silent. This reminded me of that 404 Media report—AI companies starting to buy second-hand books on a large scale to obtain 'pure' training data. Behind this lies a hidden battle over the quality and ethics of training data.
Physix Frontier