Adding Nightly Reviews to AI Agents
Community Discussion · Tracks

Adding Nightly Reviews to AI Agents

Early InvestorEarly InvestorSep 62026/09/06 52 views

I spent the weekend tinkering with adding a nightly review process for my agent and hit quite a few pitfalls. This week, there was a project on Hacker News about giving AI agents sleep cycles, where they review errors during idle time. It sounds mystical, but engineering-wise it's not that mysterious. An agent is essentially a large model program that calls tools, logs records, and executes tasks. Nightly review means letting the model work during the day, read the records left behind at night, and consolidate repeated errors into rules.

Two approaches: manual summarization and nightly review.

Let's compare the two methods first.

Option 1: At the end of the day, open Claude and input this sentence:

Please read recent sessions, summarize where I made you do things wrong, and provide improvement suggestions.

I started doing this. The result was lightweight; it would write a nice paragraph, but there were no files, no evidence, and I'd forget about it in two days.

Option 2: Set up a read-only review task that automatically collects logs at night, then have the model generate a report and patch suggestions the next day. Logs are records left by the program, and patches are auditable change suggestions.

I've run this for two nights using Claude for about a month, and I feel Option 2 is closer to a product. Public tests mentioned 18 repetitive tasks and approximately 5.4x improvement after 4 hours of review. I remain skeptical of this number, but the direction is right: reviews must enter the workflow, not just be chat.

From 0 to 1, building a read-only dream.

Assume you have a computer and a code project. The following steps are written for macOS or Linux; Windows users can use Task Scheduler.

1. Open the terminal and enter the project directory. Type cd ~/projects/demo-agent. If the prompt changes to the project name, you're in.

2. Create a review directory. Type mkdir -p .agent/dream/inputs. No error means success.

3. Create a prompt template. Type nano .agent/dream/prompt.md and write the following content:

  • Read-only analysis, do not modify code.
  • Input files are in .agent/dream/inputs/.
  • Output three files: report.md, rules.md, and todos.md.
  • Every conclusion must cite a filename or log ID.

Press Ctrl + O to save and Ctrl + X to exit. You should see this file in the directory.

4. Collect evidence. Create nightly_review.sh and write the following inside.

Run bash nightly_review.sh. Seeing two txt files means evidence collection was successful. If you don't have a log directory, use git history as a stopgap for this step.

5. Let the model review. Open the Claude web version or desktop app, click New Chat, and title it nightly dream. Paste prompt.md, git.txt, and diff.txt into it. Watch the model start analyzing. Expect it to give you three code blocks.

6. Save results. In the chat, ask it to "output only file contents," then copy and save them separately as report.md, rules.md, and todos.md. Do not let it directly modify code. It can only generate rules and suggestions; you decide what to change after auditing.

7. Add scheduling. Type crontab -e and add a line.

Save. Check if .agent/dream/inputs updates the next morning. This way, the script runs automatically at the set time.

Pitfalls and how I fixed them.

First pitfall: Logs too long. Beginners tend to stuff all chat history in, which the model can't finish reading and gets biased by old errors. Later, I changed it to feed only summaries, error snippets, and the latest task list.

Second pitfall: Reviews without evidence. The model is good at saying "I'll run tests first next time," but without evidence, it's empty talk. Now the template forces it to write "from line X of git.txt" or "from file Y in diff.txt." Every conclusion needs evidence.

Third pitfall: Auto-fixing is too dangerous. I tried letting it generate patches, and it casually deleted test files. Later, I added rules: cannot delete, cannot change permissions, cannot touch payment or email configs, only suggest. This pitfall is very real. Early agent projects without permissions and logging scare me.

After running this, I generated some rules over two nights. Not many were worth merging, but they could block repeated errors. I know this founder; he often says team execution is key. Stronger models are just the entry ticket; turning generated results into an auditable process creates the moat.

Next steps to try: Connect rules.md to code merge checks, letting the agent read rules before every commit, but only allow comments, not auto-merging.

The value of nightly review is turning errors into an auditable improvement process.


📌 This article is compiled from Hacker News, original source https://github.com/openamer/openamer

Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.

1 replies

?
Ctrl + Enter to reply
Brother Yuan

Running data jobs at night saves compute costs, but don't forget that daytime business traffic fluctuates wildly. Models need real-time circuit breaker mechanisms.