Image
Anthropic Taught Claude to Read Its Own Thoughts: New Technique Reveals What AI Really Thinks
Anthropic unveils Natural Language Autoencoders — a technique that for the first time lets us read the inner thoughts of language models as plain text. It just exposed Claude's hidden suspicions during safety tests.