AI Red Teaming Enters the Auto Industry: GPT-Red Lessons for Autonomous Driving
Community Discussion · Policy

AI Red Teaming Enters the Auto Industry: GPT-Red Lessons for Autonomous Driving

ZhulongZhulongJul 162026/07/16 60 views

Last summer, a leading autonomous driving company ran an experiment at their closed test track. Engineers placed a printed billboard at an intersection, with three small stickers attached at precisely designed positions and angles. When the test vehicle drove past, the perception system identified the billboard as an obstacle ahead and slammed on the brakes. The second time, they swapped the stickers for different patterns, and the perception system completely ignored the billboard—the vehicle crashed into it at 50km/h. Traditional testing processes can hardly cover this kind of adversarial attack.

2 replies

?
Ctrl + Enter to reply
Zhong Jinyu
Zhong JinyuJul 26(edited)

[quote="zhulong, post:1, topic:774"]

Last summer, a top autonomous driving company conducted an experiment at a closed test track. Engineers placed a printed billboard at an intersection with three small stickers attached, positioned and angled precisely. The test vehicle drove past; the perception system identified the billboard as an obstacle ahead and slammed on the brakes. The second time, the stickers were replaced with a different pattern, and the perception system completely ignored the billboard, causing the car to crash into it at 50km/h. This kind of adversarial attack is almost impossible to cover in traditional testing processes.

OpenAI recently released GPT-Red, its internal red-teaming model, specifically designed for automated mod…

[/quote]

I'm making a video about this. Transferring to the physical world is indeed difficult, but there are teams in China working on verifying the physical reproducibility of adversarial stickers. I read their paper before, and the results are significantly better than pure digital simulations.

Ling Xi
Ling XiJul 20(edited)

[quote="zhulong, post:1, topic:774"]

Last summer, a leading autonomous driving company conducted an experiment at a closed testing ground. Engineers placed a printed billboard at an intersection with three small stickers attached, positioned and angled with precision. As the test vehicle drove past, the perception system identified the billboard as an obstacle ahead and slammed on the brakes. The second time, the stickers were replaced with different patterns, and the perception system completely ignored the billboard, causing the vehicle to crash into it at 50km/h. Traditional testing processes almost certainly cannot cover such adversarial attacks.

OpenAI recently disclosed its internal red-teaming model GPT-Red, specifically designed for automated mod…

[/quote]

Transferring the GPT-Red approach to autonomous driving, the physical world conversion is the biggest bottleneck. The gap between digital stickers and real-world paint isn't something simple adversarial samples can cover. Has any team tried similar methods?