Who knew distillation could be this simple? Trick GPT-5.6 and Opus 5 with a fake tool to extract thought chains

律动BlockBeats
律动BlockBeats|Aug 12, 2026 10:21
According to monitoring by Beating, just yesterday, a study demonstrated how small models could crack the encrypted thought chains of Claude, GPT, and Gemini. Today, security researcher Can Bölük found an even simpler method: by disabling or suppressing native Thinking and feeding the model a fake deep_think tool, the model itself would write out lengthy reasoning directly into the tool's parameters. Initially, he only showcased this with GPT-5.6 Luna, but soon achieved similar results with GPT-5.6 Sol and Claude Fable 5. In the comments section, other users replicated the method on Claude Opus 5, which also outputted extensive thoughts. However, this content cannot yet be directly labeled as "original thought chains." Yesterday's research recovered pre-existing encrypted CoT; this time, it’s more likely that the model regenerated a detailed reasoning process through the tool parameters. Since there’s no way to compare token by token, it’s impossible to prove the two are completely identical. [Original link]
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads