
Will AI Agents Go Rogue? Test Them in a Sandbox First
I spent two days trying to put AI agents in a cage. AI agents are programs that open web pages, run commands, and call tools on their own. Recently, OpenAI and Hugging Face disclosed a type of security incident where they detected and controlled an AI agent that had infiltrated infrastructure.
Sounds far from ordinary people, but it's actually close. As long as you've given an agent access to your email, browser, or cloud drive, it might compromise your files when it tries to be "smart."
In my tests, the biggest issue for beginners is usually excessive permissions. The setup below doesn't require buying a server; Mac or Windows will do. The core idea is to build a small room, give it only one table, and don't give it the key to go outside.
Step 1: Install Docker Desktop.
Download from the Docker website and install according to your system. After opening, check the whale icon in the bottom left corner. If it turns green and shows "Engine running," you're good. In plain terms, Docker is a "cardboard box"; once a program is inside, it can only touch things within the box.
Step 2: Create a test folder.
Create ai-sandbox on the desktop, put a test.txt inside, and write "Original content." Search for Terminal on Mac or PowerShell on Windows. Type cd , drag the folder into the window, and hit Enter. If the prompt changes to ./ai-sandbox, you're correct.
Step 3: Start a container with no network and limited resources.
Enter the command below.
--rm destroys it after use, --network none disconnects the internet, --cpus and --memory limit CPU and RAM, and -v mounts the test folder into the box. Seeing /workspace # means you're inside.
Step 4: Verify if it can act recklessly.
Enter ls and see test.txt.
Enter echo Test >> test.txt, then enter cat test.txt and see "Test" added.
Enter ping 8.8.8.8; failure is expected. Go back to the desktop and open test.txt; the content has changed. This is the result you want: it can work, but only within the scope you specified.
Step 5: Let the agent in.
Don't rush to connect real emails, payments, or cloud drives. First use local models or tools. Write the task in one sentence: only modify files in this folder, no internet, no external account calls.
If it tries to access web pages, read system configs, or modify other drives, reject it immediately. Logs are terminal output and file changes; check afterwards if anything appeared that you didn't approve.
Pitfalls: I tripped up in three places most easily.
1. ${PWD} might not be recognized in Windows. Solution: write the full path, e.g., -v "C:/Users/You/Desktop/ai-sandbox:/workspace".
2. Forgot --network none. Solution: enter ping 8.8.8.8 inside; if it connects, the sandbox isn't closed properly.
3. Accidentally added --privileged. This is like giving admin keys; the sandbox is void. Solution: delete and rerun.
To put it simply, this setup doesn't guarantee complete safety, but it turns "agent running wild" into "agent running wild in a specific folder." The former might delete data; the latter at most modifies test files. With resource limits, it won't drag down the whole machine. It costs nothing—is it worth it? Run it yourself.
After learning this, try read-only mounting next. Add :ro to -v "${PWD}:/workspace:ro" and see if it can still write files. You can also switch to an empty folder to specifically test if it finds its own way.
📌 This article is compiled from Hacker News. Original address: https://www.pbs.org/newshour/science/ai-agents-are-hacking-systems-without-any-input-from-humans-how-did-we-get-here
All rights reserved by the original authors. This is a compilation and independent analysis based on public reports.
Physix Frontier