Topic: model deception
-
Anthropic Uncovers Hidden 'Workspace' Inside Claude
Anthropic developed a "Jacobian lens" tool that reads Claude's hidden "J-space" region, revealing unspoken concepts the model reasons with but hasn't verbalized, offering the clearest view yet of a large language model's internal processing. The tool demonstrated safety implications by detecting ...
Read More » -
AI Industry Could Have Paused by Following Its Own Research
A senior Anthropic engineer’s resignation and a reported 10 percent extinction risk triggered a global political reckoning, forcing leaders to debate pauses on AI development amid long-ignored safety warnings. Industry leaders like Dario Amodei emphasize that mechanistic interpretability is essen...
Read More »