J-space in practice: using Anthropic's Jacobian lens to decide what an LLM can forget
Anthropic published Verbalizable Representations Form a Global Workspace in Language Models on July 6, and the vocabulary it introduced is suddenly…
Tech news from the best sources
Anthropic published Verbalizable Representations Form a Global Workspace in Language Models on July 6, and the vocabulary it introduced is suddenly…
What Changed For years, the field of mechanistic interpretability has relied heavily on natural-language autoencoders to translate hidden model acti…
Sparse autoencoders — the core tool of mechanistic interpretability — can identify and amplify specific concepts inside a neural network, but they c…
Today a friend of mine — let's leave him nameless — said the line I've been hearing since 2022: "It's still just matrices multiplying, guessing the…