When Tokens Cost More Than Employees: Korean Giants Start Doing the Math
Community Discussion · Policy

When Tokens Cost More Than Employees: Korean Giants Start Doing the Math

Lei Who Shoots FilmsLei Who Shoots FilmsJul 272026/07/27 65 views

Guess how much a Samsung employee costs the company in token fees every day just by using GPT-4 to write code?

The answer might be more expensive than their coffee. Korean companies have recently been caught off guard by the high token costs of overseas AI models. Giants like Samsung and LG have even started implementing quota systems—limiting each person to a certain number of tokens per day, with excess usage paid out of pocket. This sounds like a joke, but it reflects the most real pain point in AI implementation.

In the short term, token quotas are a compromise, more like a "painkiller."

On the surface, companies control costs, but this actually exposes two problems: First, the pricing logic of model vendors hasn't worked out yet. OpenAI's API prices are calculated by token, but for enterprise-level use, token consumption skyrockets exponentially as business grows. A Korean financial firm calculated that switching their customer service system to GPT-4 caused monthly costs to jump 8 times over, making it more expensive than hiring humans. Second, employees haven't developed a "cost awareness" regarding usage habits. Many treat AI like a free playground, generating thousands of tokens of redundant output for a single complex task, leaving the company to foot the bill.

From a management perspective, the quota system is clever: it limits high-frequency use, forcing employees to optimize prompts and cut out fluff. But the side effects are obvious—innovation gets stuck behind quotas. R&D teams doing experiments might consume tens of thousands of tokens in a single complete code generation workflow; once the quota is used up, they have to wait until tomorrow. It’s like giving programmers a keyboard but only allowing them to type 1,000 characters a day, charging per character for anything beyond that.

In the long run, the token cost war will force two directions: either models get cheaper, or enterprises build their own.

Let's look at the model side first. Cost reduction is inevitable, but the speed depends on competition. Currently, GPT-4 token prices have already dropped 30% since the beginning of the year, but for enterprise deployment, the drop is far from enough. The current predicament of Korean companies is precisely caused by "excessive model capability"—many tasks don't need such powerful models, but companies take the easy way out and deploy the strongest models directly, resulting in exploding costs. In the long run, smaller vertical-scenario models will become more popular, such as CodeLlama specifically for coding, or pinyin models for document summarization, where token prices are an order of magnitude lower.

Now let's look at the enterprise side. For giants like Samsung and LG, building their own models is the ultimate solution. But building your own isn't about copying GPT-4; it's about fine-tuning open-source models combined with private data. Korean companies are currently stuck between "data security" and "cost," not wanting to hand data over to overseas models, nor wanting to spend money training their own. Quotas are just a transition; once they do the math, they will definitely shift to hybrid solutions: using privately deployed open-source models for core businesses, and API calls for non-sensitive businesses, with strict limits on token usage.

There's another overlooked point: token quotas may spawn a new "AI cost management" track. In the future, there will be products similar to FinOps, helping enterprises monitor token consumption, optimize prompt lengths, and automatically switch model tiers. Short video creators making content now also pay attention to "costs"—for example, when generating a video script, whether to use GPT-4 or Claude, what's the efficiency difference, what's the price difference? The logic is exactly the same as for these Korean companies.

Back to the question at the start: when token costs exceed employee salaries, companies aren't stopping AI usage; they're starting to calculate more meticulously. Quotas aren't a retreat; they are the coming-of-age ceremony for AI implementation—the shift from "mindless integration" to "meticulous calculation" is true scaling.

Original link: https://www.ithome.com/0/982/241.htm

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts