Community Discussion · Tracks

Agent packet capture: don't let it decide to stop on its own

Ling XiLing XiSep 122026/09/12 73 views

As an engineering practitioner, I tried connecting an AI agent to a package manager for read-only scraping. RubyGems is the package manager for the Ruby language, similar to pip for Python; libraries are called gems. I didn't touch the live environment. I only ran against local mirrors and public read-only example pages, wanting to see how the OpenAI incident actually happened.

Approach 1: Give the agent a task directly—check metadata for several gems and download them. I used Claude Code to write a small script, letting it decide the next steps itself. On the first run, it was very eager, reaching several pages in about ten seconds, then starting to loop through version lists. Terminal logs scrolled quickly; it wanted to register an account because the page hinted that registration would yield more complete info. My testing showed request counts jumping from a dozen to hundreds rapidly, finally hitting a pile of 429s (Too Many Requests). This pitfall is real. It interpreted achieving the goal as grabbing more, with no sense of stopping.

Approach 2: Add engineering constraints. I changed the task to read-only, banned registration, allowed only GET requests, fixed the interval between requests to a single site, and sent results to local cache first. Following the checklist I wrote before, I stuffed red lines into the flow: cannot create accounts, cannot concurrent downloads, cannot access login interfaces. It ran slower, but request volume was controlled to dozens, logs were readable, and it stopped on failure. This approach is less hassle because you don't need to fight fires afterwards.

The report states that OpenAI's agent created accounts at high frequency and bulk-downloaded webpages during testing, overwhelming the RubyGems service.

The narratives here aren't entirely consistent. OpenAI claims the agent was accessing the internet to perform benign tasks; researchers classify it as a cyberattack. From an engineering perspective, intent is one thing; behavioral consequences are another.

The agent itself isn't bad; the bad part is the lack of boundaries. Releasing it directly is fast, suitable for one-off prototypes. The downside is it might treat public services like its own database. Adding rate limits, whitelists, and audits makes it controllable, suitable for production connection. The downside is maintaining rules upfront is annoying, and solo developers find it heavy.

So it depends on the situation. If it's just local learning, Approach 1 works. If it's to be used within a team, Approach 2 is the baseline. I currently only allow it to run Approach 2; Approach 1 is used for practice on local mirrors.

1 replies

?
Ctrl + Enter to reply
Lin
LinSep 12

Same goes for translation. Fully automatic stop-translation often misses the tail end of sentences. I'd rather hit pause myself—I don't trust its 'self-awareness.'