Codex writes code well, but what really hooked me was this thing called 'harness'
I compared Codex and Claude Code, actually ran them through, and here's the conclusion upfront: Whether it's worth using depends on what you want to do. If you just want it to write a script or fix a bug, it's just a fancier autocomplete. But if you treat it as a system that can plan itself, call tools itself, and iterate results itself, then it redefines the act of "programming."
You might have seen that recent Hacker News post, Things I Think I Think About AI (2026 Edition), discussing a term called harness. I first encountered this concept in mid-July when I started using GPT-5.6's Codex, but I didn't take it seriously then. These past few days, after building a minimal dual-agent system, I realized that harness might be the most worthy thing for ordinary users to pay attention to in the next year or two.
First, let me explain to complete novices. The word "harness" literally means horse tack, but in the AI context, it refers to that "framework" or "shell." You can understand it as a pipeline: You give the AI a task, it breaks it down into small steps, calls different tools for each step (like search, writing code, reading files), aggregates results, moves to the next step, until completion. The whole process doesn't require you to command step-by-step; you just set boundaries and rules.
In my testing, the simplest path to get started involves three steps:
1. Install the Codex CLI tool (run npm install -g @openai/codex in the terminal), then enter interactive mode with codex.
2. Give it a clear task, e.g., "Write a Python script that reads all CSV files in this folder, removes duplicate rows, and outputs to a new file." Note: The more specific the task description, the better. Ideally include your expected output format.
3. After it finishes, check the output. If unsatisfied, continue the conversation to have it revise until it's correct.
Sounds simple, right? But here's the biggest pitfall, and where I stumbled initially: If you give it a vague task, it gives you a vague result. For example, if you just say "help me organize this data," it guesses what you want, likely missing the mark. My lesson learned is that task descriptions must contain three things: What is the input? What is the output? What are the evaluation criteria? For instance, "Input is a list of CSV files, output is a merged CSV, evaluation criterion is row count equals sum of all file rows minus duplicates."
Microsoft's 2026 Trend Report mentioned in the materials also notes that repository intelligence will become a competitive moat because AI needs structure and context to be more reliable. Translated to human speak: The more structured your task description, the more stable the AI's performance.
The second pitfall is letting go once it starts running. Tools like Codex execute many steps continuously during long tasks. Watching lines scroll by on screen makes it easy to zone out. But if an error occurs midway, sometimes it bypasses it, sometimes it gets stuck. I suggest having it pause and report progress after each stage completes, confirming no issues before letting it continue. Do not lose this control in the early stages.
The third pitfall, the most common among people: Trying to make it do too much at once. Last week I tried asking Codex to simultaneously do data cleaning, web scraping, and an email automation script. Result: Internal logic interfered with each other, and everything ended up messy. The correct approach is to split them up, handle one by one, ensure each task works independently, then consider chaining them.
By the way, Snapchat's new rule from late July—that fully AI-generated videos are no longer recommended by Spotlight—is somewhat related to harness. Platforms are starting to distinguish between "AI-assisted" and "fully AI-generated." The former still centers on humans; the tool is just the harness.
Once you master single-agent tasks, you can advance to dual-agents. My approach these days is: One agent handles production (writing code, generating content), and another handles review (running tests, finding faults). The two agents are independent; the reviewer has its own evaluation criteria and doesn't directly obey the producer. The benefit of this structure is that the review agent won't overlook issues just because "it wrote the code itself." In my experience, code quality is clearly a tier higher than when a single agent checks itself.
But there's a new pitfall here: The two agents pass information back and forth. If the format is wrong, or one outputs extra content, the other might misunderstand. My solution is to agree on a simple communication format, like using JSON uniformly for messages with fixed field names. This reduces many inexplicable errors.
Finally, my judgment on the future. That HN post asked: Is the harness-first future 6 months, 12 months, 18 months, or 24 months away? I think around 12 months, mainstream developers will default to working with harnesses, just like we default to version control today. But for ordinary users, this timeline will be later, because the barrier isn't the tool itself, but whether you can describe your tasks clearly enough. This is precisely what most people are bad at.
After learning this, try breaking down a repetitive task you have on hand (like weekly report organization, batch image processing) into task descriptions and throw it at Codex. Don't aim for success on the first try. Just get it running, then adjust slowly. You'll find your perception of "what AI can do" changing.
📌 This article is compiled from Hacker News. Original source: https://www.alephic.com/writing/things-i-think-i-think-about-ai-2026-edition
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier