Friend recommended Moonshot AI's Kimi K2.6; my calculations show DoorDash's investigation isn't unjustified
Bottom line first: DoorDash using Kimi K2.6 for coding is smart money from a cost perspective, potentially saving millions of dollars annually. But from a financial risk perspective, discounting the expected loss of this political landmine actually eats up those savings. After doing the math, I feel that being investigated wasn't an accident, but something that should have been flagged in the valuation model long ago.
A friend sent me a link last week saying DoorDash founder Andy Fang bragged on X that the combo of Kimi K2.6 + Claude 3.5 Opus is stronger than Claude 3 Sonnet + GPT-4o, and cheaper too. My first reaction: Can we trust this data? As a CFO, I need to build a spreadsheet to verify it myself. I happened to have written an analysis of Kimi K3 before, so I had Moonshot AI's pricing data on hand. Let me take this opportunity to teach everyone how to break down such AI selection decisions from a financial perspective.
Step 1: Figure Out What You're Calculating
The most common mistake beginners make is comparing model unit prices directly without looking at usage scenarios. DoorDash mentioned "internal code review processes," which fall under low-level development tasks—like checking syntax errors, adding comments, running test cases. These tasks have huge call volumes but don't require high model precision. So what we need to calculate is Total Cost of Ownership (TCO), including: API call fees, extra manual review time due to performance differences, and political risk discounts.
I built an Excel sheet with a simple structure:
| Item | Kimi K2.6 Plan | Claude+GPT-4o Plan | Notes |
|---|---|---|---|
| Price per Million Tokens | $4.8 | $12.5 (Avg Claude+GPT-4o) | Based on public API pricing, Kimi is ~60% cheaper |
| Monthly Call Volume (Million Tokens) | 500 | 500 | Assumed DoorDash code review volume |
| Monthly API Cost | $2,400 | $6,250 | Simple multiplication |
| Annual API Cost | $28,800 | $75,000 | Multiply by 12 |
| Labor Cost Due to Perf Diff | 0 | +$15,000/yr | Kimi has slightly higher false positives, needs more re-checking time |
| Total Annual Direct Cost | $28,800 | $90,000 | Significant gap |
Open Excel, type "Kimi Plan" in A1, "Traditional Plan" in B1, and fill in the numbers according to the table above. Key formula: =B2*B3 calculates monthly cost, then multiply by 12. Remember to set cell format to currency, otherwise beginners might misread them as integers.
Step 2: Pitfall Warning—Token Billing Traps
Halfway through calculating, I noticed something off. Where did the 5 million token monthly call volume come from? It wasn't in the news, and Andy Fang didn't provide it. I checked DoorDash's publicly available engineer count (~5,000 people). Assuming each person submits 2 code reviews daily, averaging 2,000 tokens each, that comes out to 500022000*30 = 600 million tokens. But that's total reviews, not model call volume.
Pitfall: Beginners often count "all code reviews" as model calls, but in reality, only some tasks use AI. I adjusted to 5 million tokens based on the assumption that the internal code review system only covers low-level tasks. If you calculate it yourself, determine the "usage rate" first—usually between 20% and 50%.
I made this mistake. My first calculation showed annual savings of $8 million. I was excited for half a day until I realized I had inflated the model call volume by 10x. Correct approach: Ask engineers internally for actual usage rates or refer to industry reports. Without data, use conservative estimates, like 20%.
Step 3: Add Political Risk—Turn Landmines into Numbers
Now, direct cost comparison: Kimi plan saves $90,000 - $28,800 = $61,200 annually. For a company like DoorDash with $8 billion in annual revenue, this amount doesn't even cover pocket change. But when Andy Fang says "lower costs," he might refer to multiple model combinations or higher call volumes. That's not the point though.
The point is risk. Being investigated by the US House of Representatives could lead to: Fines (referencing the TikTok case, ~1% of annual revenue), Business Restrictions (e.g., banning government departments from using DoorDash), Reputation Damage (stock drop). My conservative estimate:
- Fine Probability: 30%, Amount $50 million
- Business Restriction Probability: 10%, Revenue Loss $200 million
- Reputation Damage: Stock drop 5%, Market Cap Evaporation $3 billion (DoorDash market cap ~$60 billion)
Weighting these probabilities: $50M 30% + $200M 10% + $3B * 5% = $15M + $20M + $150M = $185 million. This is the expected loss. Discounted to a one-year period (assuming 10% discount rate), it's $168 million.
Compared to annual savings of $61,200, the risk loss is 275 times the savings.
Calculate in Excel: List risks in column F, probabilities in G, losses in H. Formula =G2*H2, then SUM. Beginners often ignore probability weights and just add maximum losses, which is terrifying. The correct way is weighted average.
Step 4: Conclusion—Add a "Political Factor" to Valuation Models
DoorDash's decision-makers likely calculated this, but they bet that "the investigation won't blow up." From a financial perspective, the expected return on this bet is negative—the cost saved by using Kimi doesn't even touch the tip of the iceberg regarding risk losses. Unless their usage is massive (e.g., billions of tokens annually), but low-level code reviews won't reach that volume.
I wrote an analysis last week about Anthropic refusing to sign open-source agreements, stating then that "rational choice is to avoid risk." Looking at DoorDash now, the logic is the same:
Physix Frontier