Letting AI Play Portal: Building a Low-Cost Experimental Sandbox First
There's a hot experiment going on lately. The news says that GPT-6 Astra, released by OpenAI on September 3rd, was used by an enthusiast to autonomously complete Portal, taking 24 hours and costing $571. Many people want to try it themselves after seeing this, but the pain points are direct: not knowing where to start, and being afraid of burning money as soon as they turn it on.
My advice is: don't rush to finish the game. Have you actually run it on a production line? That's usually my first question. Finishing the game is the result; whether you can stably take screenshots, understand the visuals, issue actions, and keep logs—that's the replicable process. I've been using the OpenAI API for about a month, and over the last few days, I tried a small experiment with this mindset. The goal was to get AI looking at the screen and deciding which key to press next in ten minutes.
First, prepare three things: a computer capable of running Portal, an OpenAI account, and a script that can take screenshots and simulate keystrokes. Think of the script as a robotic arm on an assembly line; it doesn't make judgments itself, but it can translate cloud instructions into mouse and keyboard actions.
To get an API key, log in to the OpenAI website, go to the API page, click Create new secret key, and copy the string when the key popup appears. This key is like a factory access card—don't post it in group chats, and don't hardcode it in your code. Beginners often upload it to GitHub; the solution is to save it in local environment variables.
When setting up the environment, open the terminal and enter the command below.
Seeing (.venv) means the environment is activated. Here, mss handles screenshots, pyautogui handles simulated keystrokes, and python-dotenv reads local configurations.
When writing the budget configuration, create a new .env file and fill in the content below.
bash
OPENAI_API_KEY=paste_your_key_here
MAX_COST_USD=small_budget
MAX_STEPS=300
The $571 mentioned in the news is for the full experiment. For personal reproduction, there's no need to burn that much right away. Set a small budget first, get it working, then talk about scaling.
When building the main loop, the core logic can be broken down into five steps. Take a screenshot of the game every three seconds. Crop or resize the image, keeping only the central area to reduce token consumption. Tokens can be roughly understood as the billing unit when the model reads images or text. Send the screenshot to GPT-6 Astra and ask it to return a JSON action, such as {"action":"w","duration":0.3}. The script only allows whitelisted actions to execute, such as w, a, s, d, and mouse_click. Write the time, action, tokens, and cost into the log.
Suggested log fields include step, time, action, tokens, cost, and screenshot_path. The expected result is the terminal printing a line of action, tokens, and cumulative cost every few seconds. Without logs, the experiment is just a spectacle.
Beginners most easily make mistakes in three areas. Screenshots are too large; throwing full HD original images directly at the model causes costs to skyrocket. The solution is to crop to a smaller size. The game window loses focus; the script presses keys, but the game doesn't respond. The solution is to activate the window before each action. There are no stop conditions; the model might repeatedly press the same key. The solution is to pause if the screen changes very little for several consecutive steps.
The advantage of this experiment is its simple workflow, suitable for understanding how multimodal models and automated executors work together. The disadvantages are uncontrollable costs, latency, and misoperations. The opportunity is that it can be transferred to industrial quality inspection: cameras see defects, models give judgments, and actuators sort items. The risk is model hallucination: thinking it succeeded even when the screen hasn't changed.
| Comparison Item | News Experiment | Beginner Minimal Experiment |
|---|---|---|
| Goal | Finish the game | Get from screenshot to action |
| Duration | 24 hours | 10 minutes |
| Cost | $571 | Controlled within a small budget |
| Output | Game results | Action logs |
After learning this, the next step can be trying a simpler web puzzle. Lower MAX_STEPS to under one hundred, specifically to verify failure recovery and log review.
From a production line perspective, having AI connect perception, decision-making, execution, and recording is more useful than playing puzzle games. Quality inspection equipment is the same: the camera is the eyes, the model is the brain, the cylinder is the hand, and MES is the log. How much yield improves often depends on whether this process can run stably.
Don't chase finishing the game yet. First, set up the budget, logs, and failure recovery.
Physix Frontier