US Accuses China's Moonshot of Stealing From Anthropic's Fable AI Model | Vantage on Firstpost | 4K

美国指控中国月之暗面窃取Anthropic旗下 Fable 人工智能模型相关技术

China's AI start-up Moonshot AI is facing allegations from the United States that its latest AI model, Kimi K3, was built using proprietary American technology developed by Anthropic. Washington claims Moonshot used large-scale AI distillation without permission, while Beiing has rejected the accusations as baseless. The dispute follows earlier allegations by Anthropic and OpenAI that Chinese firms misused their AI models.

中国人工智能初创企业月之暗面(Moonshot AI)遭到美国指控,美方称其最新人工智能模型 Kimi K3,依托美国公司 Anthropic 研发的专有技术打造。美国政府表示,月之暗面在未经许可的情况下开展大规模人工智能模型蒸馏;中国方面驳斥该指控毫无依据。在此之前,Anthropic 与 OpenAI 就已经指控多家中国企业不当使用它们的人工智能模型,本次争端正是这一系列风波的延续。

以下是印度网友的评论:

glenwjohnson809
Anyone who wasn't expecting the US to bad mouth anything good from China was just born.

谁要是没料到美国会抹黑中国取得的任何成就,那只能说他阅历太浅。

 

Danny-o2v
US reaction is so predictable. When you lose, the opponent cheated. lol.

美国这套反应早就见怪不怪。竞争输了,就说对手作弊,哈哈。

 

Sooraj09073
2.8 Trillon parameter kimi k3 needs atleast 6 months of training. Fable 5 released last week.

拥有2.8万亿参数的Kimi K3至少需要6个月训练时间,而Fable 5上周才发布。

 

PratimaDahal-rc2gw
what does indian know about AI? China finishes a project in days what India finishes in a decade
印度人懂什么人工智能?中国几天就能完成的项目,印度往往要耗费十年。

NCM-xy8ow
@PratimaDahal-rc2gw you did not understand his point

你没读懂他想表达的意思。

PratimaDahal-rc2gw
@NCM-xy8ow he was saying Kimi k3 needs 6 months to train and then a company can launch this model if it was made from scratch

他想说Kimi K3从零训练至少要六个月,企业才能推出这款模型。

NCM-xy8ow
@PratimaDahal-rc2gw he is saying there is no way Kimi could have leveraged much on Fable cause Fable only came out a week or so before Kimi was released

他的意思是,Kimi基本不可能大量借鉴Fable,因为Fable仅仅比Kimi提前一周左右发布。

kriswright4739
@PratimaDahal-rc2gw He is saying you cannot steal from Fable 5 (released a week ago) with that timeline in training

 他想说,按照训练周期来看,根本不可能从一周前刚推出的Fable 5窃取技术。

 

qkstwpzmt
Just as Anthropic steal Internet data and all other AI companies distilled from each other. But to distill from another AI you need to pay, unlike Internet data that's been taken for free.

Anthropic本身就在抓取互联网数据,各家人工智能企业也一直在互相借鉴蒸馏。但从别的模型做蒸馏按理需要付费,这和免费抓取网络数据不是一回事。

 

openyourmind2546
- When US firms distill from each other, use open-source models, and scrape global data : it's called innovation!
- When Chinese firms use distillation, which is a standard academic and industrial practice : it's framed as theft!

美国企业互相蒸馏、使用开源模型、爬取全球数据,就被称作创新。

中国企业采用业界通用的蒸馏技术,却被定性成窃取。

 

hugocavalaristudartpereira1214
Nah, the Americans are afraid because Chinese AI is so good and cheap that even American companies prefer using it, it’s the fear of losing the AI race. All that’s left is to make false accusations against China and smear their AI with stupid claims like saying the PLA is spying on you through it even when it runs entirely locally on your machine LoL.

真相是美国人感到恐慌。中国人工智能产品性能强、价格低,就连美国企业都更愿意选用。他们害怕输掉人工智能赛道的竞争。于是只能捏造罪名抹黑中国人工智能,甚至编造离谱说法,声称解放军会利用模型窃取信息,即便模型完全在本地设备运行,简直可笑。

 

Dan-d1s8k
Stolen technology?
Hey, sour grapes, where is your PROOF, where is your PATENT ?

窃取技术?哼,纯属吃不到葡萄说葡萄酸,证据在哪?相关专利又在哪?

 

openyourmind2546
- When Western firms distill from each other, use open-source models, and scrape global data : it's called innovation!
- When Chinese firms use distillation, which is a standard academic and industrial practice : it's framed as theft!

西方企业互相蒸馏、使用开源模型、爬取全球数据,就叫做创新。中国企业采用学术界和行业通用的蒸馏手段,却被说成偷窃。

 

vision9275
Losers always complain—and even whine—when jealousy strikes them at their most vulnerable point, while meanwhile China rapidly elevates its AI model to the very highest level, leaving the US far behind.

竞争落败的人心生嫉妒,就会不停抱怨。与此同时,中国人工智能模型快速跻身顶尖行列,把美国远远甩在身后。

 

prakashlalwani3466
Dogs and bitches have started to bark but Russia and China are elephant

这群人开始叫嚣,而中俄如同大象一般沉稳。

 

KalimbeziJohn
The west is desperate to maintain its declining dominance in anything that matters.

西方国家拼命想要守住自己不断衰退的核心领域主导地位。

 

jameskam473
Typical reaction of sore losers

这就是输不起的人典型的反应。

 

LoveYaJSR
chyna also stole stealth bomber tech

中国当年也窃取了隐形轰炸机技术。

 

omPaine-tt2rx
Do consider that AI exsts at all because it stole all Human knowledge totally stole all human work including my own books so you know kind of Rich to complain.

仔细想想,人工智能本身就是依托人类所有知识发展而来,相当于吸纳所有人的劳动成果,其中也包括我写的书籍。所以他们现在跑来指责别人,实在没有说服力。

 

Parangaripotirrimirruaro
Trump is pathetic.

特朗普这个人很差劲。

 

lamencoguy3000
Do you believe everything the US claims?

美国说的所有话,你都全盘相信吗?

 

byronwang7050
If China holds off on releasing similar models until Anthropic or OpenAI go public via IPO, the US criticism would be much milder.

如果中国等到Anthropic或者OpenAI完成上市之后,再推出同类模型,来自美国的指责会温和很多。

 

Java123-y8t
I think distillation in the fable 5, by any company is stupid .The cost of that is very very really huge. And distillation means producing the data from the model. Just think about how many billions of tokens. Think about the charge. Any company will not do that.
I think instead of distillation it is very very cost effective to build a new advanced system.that is the China builded.

我认为任何企业想用Fable 5做模型蒸馏都不现实,成本高到难以承受。蒸馏需要依托目标模型生成海量数据,动辄数十亿token,开销十分惊人,没有公司会这么操作。
相比蒸馏,从头搭建全新的先进模型性价比高得多,中国正是选择了这条路。

 

daShadoSage
Not the data. The harness, reasoning, routing, etc.

问题不在于训练数据,而是模型的推理框架、逻辑链路、调度机制这类核心设计。

 

frankyalferado6093

Why other non-chinese model unable to do distillation and produce something like Kimi K3?

为什么其他非中国企业没法通过蒸馏技术做出类似Kimi K3水平的模型?

 

daShadoSage
They do. GLM 5.2 and all of DeepSeek did the same thing. All of the reports are out there.

其实有企业这么做过。智谱GLM 5.2、深度求索的模型都采用过蒸馏方案,相关报道都可以查到。

 

Sooraj09073
If an Indian company were to create a frontier AI model and the US made similar allegations and imposed sanctions, does any Indian govt support the company and researchers, i doubt. Could any government on Earth other than China protect and support? I think not. Only China has the resources and power to challenge the US, and that means a lot.

如果印度企业研发出顶尖人工智能模型,遭到美国类似指控和制裁,我怀疑印度政府未必会全力扶持企业和科研人员。全世界除了中国,还有哪个国家政府愿意提供保护与支持?我认为没有。只有中国拥有足够资源和实力和美国抗衡,这一点至关重要。

 

daShadoSage
I forgot the name of the model.. Sarvam?
It actually did distill from Mistral. But did so openly, honestly, and with auth rian with a license. So an Indian company that did it right. And India has partnerships with all major US labs who are spending billions in investments. Nvidia even gave GPUs for free to develop sovereign Indian models. That's partnership.
China on the other hand, bans US models, bans US chips, and illegally distills from US models while using those same models to compete against them. So you should be asking, what if China did the same to India? Well, India did ban TikTok and a number of Chinese tech for a reason.

我想不起模型名字了,是不是Sarvam?

这家印度企业确实基于Mistral做了蒸馏,但全程公开透明,取得了官方授权,做法合规。印度和美国各大人工智能实验室建立合作,美方投入巨额资金,英伟达甚至免费提供显卡,助力印度自研本土大模型,这才是合作模式。

反观中国,一方面限制美国模型、美国芯片,另一方面非法利用美国模型蒸馏,再拿出自研产品和美方竞争。不妨试想,如果中国用同样的手段针对印度会怎么样?印度当初封禁TikTok等多款中国科技产品,不是没有缘由。

 

leisana4097
The tight timeline, training , testing and release of these two models overlaps and access barriers of FABLE, remember FABLE was closed for public access and a full scale Fable distillation is technically dubious for a full frontier model of K3's caliber. Never mind the allegations k3 achieved much more than what the Americans thought. How many samples might have been distilled in a short span of time and moreover for a giant model like K3 distilling a few samples and training again doesn't make sense at all.

两款模型的训练、测试、发布时间高度接近,而且Fable并未对外开放。要依靠Fable完整蒸馏出K3这种顶级大模型,从技术层面来说很难实现。抛开指控不谈,K3展现出的能力远超美方预估。短时间内能够获取多少蒸馏样本?更何况K3这种超大规模模型,只依靠少量蒸馏样本重新训练根本行不通。

 

PlungerBumpersTreason
LOLOLLOLOLOL who's stealing? Ohh yeah AMERICAN AI IS:
"US judge approves Anthropic's $1.5 billion settlement of copyright lawsuit"

哈哈哈哈,到底谁在窃取技术?别忘了美国人工智能企业Anthropic刚和原告达成15亿美元版权诉讼和解,美国法官已经批准这份协议。

 

Surendra_58
The question is: why they let them steal

问题来了,他们为什么放任对方窃取?

 

daShadoSage
Do ask the same of every hack, robbery, or espionage scheme? Why did they LET them?

遭遇黑客攻击、盗窃、情报窃取事件时,你也会问受害者为什么放任坏事发生吗?

 

reckoner-nash
You forgot to mention what Jensen huwang said about K3 and he said that Moonshot Kimi 3 was perfectly safe to use and that opensource models should be encouraged. He didn't blink an eye nor was he worried about Distillation

你漏掉了黄仁勋对Kimi 3的评价。他表示月之暗面的Kimi 3可以放心使用,应当鼓励开源模型发展。对于模型蒸馏这件事,他丝毫没有表现出担忧。

 

daShadoSage
Because he sells GPUs lol. He's the gas company encouraging everyone to buy their own cars and drive more because of course, he makes more money. It's why he has open weight models for everything too. It's not because he has a heart of gold lol.
No matter what models are used, he makes money.

道理很简单,他是卖显卡的,哈哈。就好比燃油企业鼓励大家买车、多开车,目的就是赚取更多利润。他大力推动开源权重模型也是这个逻辑,并不是心地善良。
无论大家使用哪款模型,他都能赚到钱。

 

coolwalk11
Arrest modi now!! He’s a war criminal and there no de ocracy in India no free press!!!

立刻逮捕莫迪!他是战争罪犯,印度没有民诸,新闻也不自由!

 

duoduo1885
student A scoring 80 tells teacher student B scoring 100 copies his answers

考80分的学生,向老师告状,说考100分的同学抄袭自己的答案。

 

daShadoSage
Lol complete reverse. I'm guessing you got nothing higher than all C in school with your level of analysis and attention span. The video wasn't that long guy. Pay attention and you will here that Student A with 100 is copied by Student B with a 75... while also taking longer to get that 75. Go watch and listen again.

哈哈,事实刚好反过来。看你分析问题的水平,上学时成绩估计一直很差。视频并不长,认真听完就能明白:考100分的学生,被考75分的学生抄袭,而且后者花了更长时间才拿到75分。

 

Bheemlogic
Can we do the same

我们能不能照搬这套手段?

 

roro-v3z
I think it's not at all significant. It's just Anthropic whining because they fear open weight models.
Fable and Kimi were 2 weeks apart, it doesn't make any sense to claim it was distilled. Secondly, the token cost would be so high if they distilled that the company would just not be able to sustain.
Final point, when Anthropic is scra internet for training, there is no moral authority here!

我觉得这件事根本站不住脚。Anthropic只是害怕开源权重模型带来竞争,才不断抱怨。

Fable和Kimi发布时间仅仅相隔两周,说Kimi依靠蒸馏Fable得来完全不合逻辑。其次,大规模蒸馏产生的token成本极高,企业根本负担不起。

还有一点,Anthropic自身就在抓取网络数据训练模型,它根本没有资格站在道德高地指责别人!

 

Sooraj09073
Fable 5 came out last week, and a 2.8T parameter frontier model needs 6 months of training. So yeah, the US has a point. China only knows how to copy. Now they catch up and and soon moving forward like in everthing else.

Fable 5上周才推出,而2.8万亿参数的顶尖大模型至少需要六个月训练周期。这么看美国的说法有一定道理。中国向来只会模仿抄袭,如今好不容易追赶上,接下来各个领域都会继续走这条老路。