AI Extraction Controversy: Stealing Recipes or Copying Homework?
This situation is like a restaurant kitchen. You smell a dish that tastes like the Michelin star place next door—is it stealing the recipe, or can anyone produce it with the same batch of ingredients?
Recent Bloomberg news has been intense. Multiple US institutions have named DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, claiming they systematically copy US models. Anthropic previously accused Alibaba of using thousands of fraudulent accounts to extract Claude's capabilities. Alibaba denied this and sued the US government, demanding removal from the Pentagon blacklist. OpenAI also stated they collected some evidence and are reviewing whether DeepSeek improperly distilled their models. More interestingly, reports mention that although DeepSeek is considered a national security risk by some cross-departmental assessments, the Commerce Department hasn't yet placed it on the trade blacklist. This shows that accusations, risk judgments, and actual regulations are three layers, not the same net.
But here are two terms often mixed up: distillation and extraction. Distillation is like a master chef compressing experience into a recipe usable by an apprentice; it's common in the industry. CSIS articles also note that US labs use distillation; DeepSeek even released R1 distilled models based on Qwen and Llama. Extraction is more like using someone else's API as a free question bank, asking in batches, collecting answers in batches, and feeding them into your own training pipeline. The issue isn't similarity, but how the answers were obtained.
The closed-source model route in the US is like making dishes in a central kitchen. Model capabilities are hidden in hosted services; outsiders only see the output. Pros: convenient, with safety evaluations, logs, and abuse detection held by the platform. Cons: users can't access full training data and weights; evaluation is superficial. Over the past three weeks, I ran GLM-5.2 on long-text summaries, structured tables, and small coding tasks, comparing it with open-source models and APIs I've used this month. Testing showed that output style, refusal responses, and error types can be mimicked, but the source fundamentally cannot be confirmed from a single benchmark run. Running a scorecard only tells you if it can do it, not where it came from. I also tested several open-source models on the same topics and found similarities in Chinese factual errors and over-compliance, but such similarities cannot prove extraction.
The open-weight route is like spreading the recipe on the table. Anyone can download, modify, and run it. Pros: transparent, allowing at least audits of weights and licenses, with community testing. Cons: low cost pressure, fast catch-up, and easier suspicion of standing on others' shoulders. For models like GLM-5.2 that publish weights, commercial value depends on providing traceable training sources, data licenses, and evaluation records. Closed source has benefits, open source has drawbacks; you can't just criticize one side.
What really needs to be calculated separately are three accounts. The technical account: is distillation common? Yes. The contractual account: did they bypass terms, use fraudulent accounts, or scrape in bulk? This is the core of the accusation. The regulatory account: sending controlled technology via hosted APIs vs. publicly releasing weights have different boundaries. A realistic point in the reports: public model weights don't necessarily constitute export, but putting controlled technology into hosted APIs might require licensing.
Export controls aren't metaphysics. For evaluators, the scariest thing is lumping all models into the "theft" basket. That only causes compliance costs to hurt the open-source community. Conversely, industry norms can't be used as absolution. If a company really uses thousands of accounts to exploit APIs and trains competitors with that data, the computing power they saved might come from others bearing safety and abuse costs. Whether the price is worth it isn't about tough talk, but the chain of evidence.
I usually write tool reviews and like to keep the books straight. Same here. Don't rush to ask if Chinese models copied, nor rush to say US institutions are labeling. Look at logs, account patterns, API terms, data deduplication, training pipelines, and third-party audits. If Anthropic's accusation relies only on output similarity, it's not persuasive enough. If Chinese companies only say "we trained independently," they need verifiable materials.
However, these recent days I've been trying AI agents, Agentic engineering, and Vibe coding, and I've just touched Docker for four days. The realization is direct: the more agents can automate work, the more they need to leave traces. Model extraction is essentially a tracing problem too. Without process records, no matter how pretty the result, it looks like a black box. I wrote about post-market daily reports before; the view hasn't changed—without real-time traceability, subsequent explanations are weak. Looking at the source of model capabilities is the same.
Ordinary users shouldn't treat any model as an absolutely original source. Enterprises should first do data lineage and API usage audits. Regulators must separate distillation, terms violations, and export controls. Otherwise, it ends up with both sides feeling aggrieved: one side thinks others stole answers, the other thinks others are choking them with industry norms.
Going forward, it depends on whether there are more detailed audits and litigation materials.
📌 This article is compiled from Bloomberg Tech. Original text: https://www.bloomberg.com/news/articles/2026-09-09/us-says-alibaba-deepseek-have-systematically-siphoned-ai-models
Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.
Physix Frontier