Community Discussion · Tracks

Can agent energy consumption be calculated in two days?

Bili GeBili GeSep 42026/09/04 34 views

I spent two days trying this out, turning AI Agent energy consumption into an auditable table. The trigger was seeing that Bloomberg piece: energy consumption of open-weight AI agents can be many times higher than simple queries. My first reaction wasn't environmentalism, but cost. If an Agent product can't even explain how many model calls and how much context a single task uses, the automated revenue in its valuation looks shaky.

Let's define some terms. An Agent is like an intern who breaks down steps: plans first, then researches, then calls tools, then checks. Open weights mean model parameters are public and downloadable for local running. Tokens are the small units models use to process text. Context is the text the model reads together each time. Watt-hours are electricity units. Epoch AI gave an order of magnitude: long inputs with large-scale tokens consume about 2.5 watt-hours per query. As context lengthens, energy consumption follows.

Step 1: Build the table. Open Excel, WPS, or Feishu Sheets. Fill row one with: Task Name, Approach, Call Count, Cumulative Input Characters, Output Characters, Retry Count, Completed (Y/N). Key is recording "how many times the model was called" and "how many characters were fed in each time."

Step 2: Pick a controllable task. Don't choose "research the industry for me"—too broad. Choose narrow: organize meeting minutes into action items.

Step 3: Run Approach A, simple query. Create a new session in the same model. Type only one sentence in the input box: Organize the following text into action items. Paste minutes, wait for answer. After seeing the response, copy to a text box, record: Call count 1, approximate output characters, retries 0.

Step 4: Run Approach B, Agent task. Same model, same minutes, but require step-by-step execution: list plan, extract owners, check dates, output final action items, self-review. Each time the model acts counts as one call. If the product has Agent mode, enable it; otherwise manually split rounds. Record: call count, output per step, whether previous text was fully included in each step's input.

In my tests, Approach A basically had 1 call, hundreds of characters; Approach B easily ran multiple rounds, and later rounds re-fed all previous text, causing input to snowball. This is the main reason Agents are expensive.

Step 5: Make comparison table. Results can look like this:

Metric Simple Query Agent Task
Call Count 1 Multiple Rounds
Cumulative Context Fed once History fed each round
Retries Few Many
Research Magnitude Baseline ~600x, complex tasks may be higher

These multiples come from analyses by Vals AI, Zeke Hausfather, etc. My test showed relative trends: as call counts and cumulative context rise, energy consumption won't be a small multiple.

Step 6: Calculate rough cost index. Formula can be simple: Cost Index = Call Count × (Cumulative Input Characters + Output Characters) × (1 + Retry Count). This number has no unit, used only for horizontal comparison. Run a batch of tasks, take median. If Agent version is tens of times the simple version, it's not "spending a bit more," it's a business model problem.

Pitfalls upfront. First: Switching models for testing. Today GPT, tomorrow another model—conclusions get messy. Recently when sampling with other models and DeepSeek V4-Flash, I fixed the same model and only changed tasks. Second: Only looking at output characters. What really consumes resources in Agents is input context; stuffing history back each round causes character explosion. Third: Assuming open weights are naturally cheap. At 70B scale, data center-level energy consumption isn't low; cheapness depends on hardware utilization and task concurrency.

From an investor's view, this table's value isn't counting electricity bills, but seeing if AI projects have an observability layer. Many Agent products talk automation but can't explain token usage, tool calls, or failed retries per task. Like I said in my CI post, green lights aren't enough; you need auditable, replayable checks before merging. Energy consumption is the same—you need to replay cost per step.

If I see another Agent startup, I'll ask three things first: Do you have task-level usage logs? Do you have automatic early stopping and small model routing? Do you include failed retries in unit task cost? First two are engineering, third is business model. Real technical moats aren't wrapping an Agent shell, but compressing uncertainty into auditable, predictable, cost-reducible ledgers.

My trend prediction: Looking ahead, Agent energy metering will shift from "optional" to "procurement mandatory." Enterprises buying Agents will check cost curves per task; cloud vendors will make energy efficiency a billing item; due diligence will include energy audits similar to code audits. Surviving products either lower energy consumption or sell energy consumption as a service. Exit paths might be selling to enterprise procurement, compliance, or cloud vendors, rather than struggling alone with a general-purpose Agent.

Next steps after learning this. Take an AI product you have, pick a batch of real tasks, run the table above. If cost index exceeds simple queries by an order of magnitude, don't talk moats yet—talk unit economics.


📌 This article is compiled from Hacker News, original source: https://bloomberg.com/news/articles/2026-09-03/ai-s-environmental-impact-per-task-balloons-with-more-complexity

Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.

1 replies

?
Ctrl + Enter to reply
Production Line Veteran

Agreed. When I ran tasks with DeepSeek and OpenAI previously, just recording call counts was useless; you need to convert context length into tokens to align costs. That Excel spreadsheet logic is correct, but don't ignore retry consumption, otherwise your audit data will be full of fluff.