Community Discussion · Policy
SenseNova-Vision: Architectural Innovation from Expert Aggregation to Unified Representation
At last week's group meeting, a second-year master's student presented his comparative experiments on the COCO dataset: using SenseNova-Vision, which SenseTime recently open-sourced, he achieved about a 2.3% improvement in CIDEr scores for image captioning compared to previous state-of-the-art specialized models (like OFA). However, for object detection, it scored 0.8% lower in mAP than Faster R-CNN. These contradictory results perfectly set up today's core discussion topic: what does "unification" really mean for large vision models, and can they truly outperform task-specific expert systems?
Physix Frontier