Community Discussion · Tracks

Don't chase the cheapest model; build a cost-performance ledger first

Pao Tiao XianPao Tiao XianSep 32026/09/03 36 views

Don't chase the cheapest model; build a cost-performance ledger first

Keeping up with high-value AI models doesn't require scrolling through posts daily. Just make a ledger that expires. Many people ask how to keep track of the most cost-effective models. Prices, task quality, and token metrics change together, so you need a review process.

Last week I wrote about chasing Bel models; this time I'm turning it into price tracking. I used Claude Code and a search agent for less than a week to compare two approaches. Approach one was manually browsing X, IT Home, and pricing pages to find low-cost models suitable for coding. It took about twenty minutes and I still missed a limited-time price. Approach two had the search agent pull several price-tracking pages into a timeline, generating candidates in just a few minutes. After inputting List models suitable for coding with output prices below $5/1M, and note the update time, it provided Gemini 3 Flash at ~$0.50/1M output, DeepSeek V3 at ~$0.27, and Claude Sonnet 5 at a limited-time $2/$10, rising to $3/$15 after expiration. But I still had to click back to the pages to check the update times.

Start by building the ledger. Beginners shouldn't rush to ask "which model is the cheapest." Open a Feishu spreadsheet and create seven columns: Model, Input Price, Output Price, Limited Time?, Context Length, My Quality Threshold, and Last Verified Date. Tokens are the unit of measurement for AI; input price is what you pay to feed data in, and output price is what you pay for the response generated. Context length is how much conversation material it can remember at once; longer isn't necessarily cheaper. Coding tasks usually have longer outputs, making output price more sensitive.

Next to the table, list three tasks: code modification, reading long documents, and Chinese polishing. Set a baseline for each task, e.g., eliminate anything scoring below 3. Pages like aicost also suggest setting a quality threshold first, then finding the cheapest option that meets it.

Run a real test. Go to aggregator sites like ModelPicker or AI Cup, click Coding, enter output price < 5 in the filter box, and expect to see a table or scatter plot yielding three to five candidates. Copy the model names and prices into your ledger. Then do some quick math: assume 500k input tokens and 100k output tokens per month. The cost for Sonnet 5 at the limited-time price is ~$2, rising to ~$3 after expiration. If code tokens increase by another 10% to 35%, actual spending will go even higher. I tested three models on a Python script modification task. The cheap DeepSeek was fast but its comment style wasn't what I wanted. Gemini 3 Flash was sufficient for short edits, but complex multi-file tasks were more stable with a pricier tier.

The pitfalls are here too. One pitfall is looking only at input prices and ignoring output prices; many models have cheap inputs, but output costs explode when generating code. Another is treating the agent's words as conclusions; it might only capture a snapshot, missing "limited-time end" dates or tokenizer changes. A third is that larger context lengths don't always mean savings; feeding in more long documents increases input costs. My solution is to keep sources for every column and review weekly on Fridays.

After learning this, the next step is to take your recent week's AI subscription bills, fill in the input/output tokens, pick two candidate models, and run them through the same task. Don't rush into RAG. I tried it for the first time yesterday; turning the ledger into a searchable Q&A system is for later. First, get the accounting clear.


📌 This article is compiled from Hacker News, original source: https://news.ycombinator.com/item?id=49550324

All rights reserved. This is a compilation and independent analysis based on public reports.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts