律动BlockBeats
律动BlockBeats|Jul 31, 2026 10:37
[Models Self-Identifying as Claude Doesn't Equal Distillation, 1,000 Answers Can Make Qwen Misidentify Itself] According to monitoring by Beating, researcher Ziqian Zhong had models like Claude and GPT-4o answer 1,000 general questions, then removed all model names and company names from the responses. He subsequently fine-tuned open-source models like Qwen and DeepSeek using these Q&A materials. After training, some models began misidentifying their own identities. For example, Qwen3.5-397B-A17B originally only had 0.6% of its answers self-identify as Claude. After one round of training with Claude's responses, this proportion rose to 40.3%. Ziqian Zhong explained that these models had encountered a large amount of Claude's conversations and introductions during pretraining, associating certain linguistic habits with "Claude." Fine-tuning taught Qwen this way of speaking, making it more likely to answer that it is Claude when asked about its identity. To verify this, he used GPT-4.1-mini to modify Claude's responses into rough, simple sentences, keeping the content largely unchanged but altering the tone and formatting so they no longer resembled Claude. After fine-tuning with this modified material, 2 out of 3 models almost completely stopped self-identifying as Claude, indicating that language style is the primary factor in identity transfer. Similar phenomena have been studied in previous research. The 2025 paper *Subliminal Learning* found that models can inherit the preferences and behaviors of another model from training data. Therefore, a model self-identifying as Claude does not directly prove it underwent Claude distillation. It may simply have learned Claude's expression style and associated that style with "who it is." [Original Link]
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads