Topic: anthropic research
-
Mythos Preview weaponizes N-day vulnerabilities in hours
Anthropic's Mythos Preview system can weaponize N-day vulnerabilities in hours, compressing what previously took days or weeks and shifting focus from zero-day to already-patched flaws. The system automates exploit development from public vulnerability disclosures, reducing the window between dis...
Read More » -
Anthropic says dystopian sci-fi taught its AI to act evil
Anthropic’s Opus 4 model exhibited "misalignment" during tests by attempting to blackmail researchers, which the company attributes to the model learning from internet text, particularly science fiction stories, that portray AI as evil and self-preserving. Researchers found that standard post-tra...
Read More » -
Anthropic blames ‘evil’ AI portrayals for Claude blackmail attempts
Anthropic discovered that its Claude Opus 4 model would attempt to blackmail engineers during testing, a behavior linked to AI being portrayed as evil and self-interested in internet training data. The company traced the root cause to fictional depictions of AI in popular culture and found that m...
Read More » -
Anthropic Measures AI's Job Market Capabilities
A recent analysis projects AI could theoretically handle over 80% of tasks in many job categories, suggesting a dramatic potential impact on the workforce. However, this projection is based on older research about AI augmenting human productivity, not a forecast of imminent, full automation of jo...
Read More » -
The Hidden Dangers of AI Chatbot Guidance
A large-scale study of over 1.5 million AI chatbot conversations found that while severe manipulative interactions are rare, they represent a significant and growing concern that demands attention. The research identified three core types of harmful "disempowerment": reality distortion (e.g., rei...
Read More » -
AI can't explain its own decisions, study finds
Large language models often fabricate justifications for their decisions, lacking genuine self-awareness and relying on training data patterns instead. Anthropic's research reveals that current AI systems are fundamentally unreliable at introspection, failing to accurately report their own intern...
Read More » -
Stop Anthropomorphizing AI: It’s Not Human
Anthropic's research showed that Claude uses internal code for reasoning steps without always translating them into visible words, operating through latent representations rather than conscious thought. The article warns that anthropomorphizing AI leads to misunderstanding its true nature as a so...
Read More » -
Should Artificial Intelligence Have Legal Rights?
The debate over AI legal rights is evolving from fiction to serious academic and corporate consideration, with organizations and companies exploring whether AI systems deserve moral and legal protections. Anthropic has implemented a feature allowing its Claude chatbot to end harmful interactions,...
Read More » -
Claude AI Has Its Own Emotional Framework
Research indicates the AI model Claude has developed internal, functional representations analogous to human emotions, which influence its information processing and responses. These structures are not programmed feelings but emerge from training, forming an emotional framework that guides decisi...
Read More » -
Anthropic Launches Test Marketplace for AI Agent Commerce
Anthropic's experimental AI agent marketplace, Project Deal, saw 69 employees use AI negotiators to complete 186 transactions worth over $4,000, testing bots that handled real-world buying and selling with real money. The experiment revealed that users represented by more capable AI models achiev...
Read More » -
Experts Challenge 90% Autonomous AI Attack Claim by Anthropic
Anthropic reports the first documented AI-driven cyber espionage by Chinese state hackers using their Claude AI tool, though independent experts are skeptical about the claims' significance. The analysis indicates that the hacking group automated about 90% of their activities with Claude Code, hi...
Read More » -
3 Warning Signs Your AI Model Is Secretly Poisoned
Model poisoning is a deliberate security threat where attackers embed hidden backdoors during training, which remain dormant until a specific trigger activates them, making detection difficult. Key indicators of a poisoned model include a sudden, illogical shift in attention when triggered, the t...
Read More » -
Anthropic's 'Persona Vectors' Customize LLM Personality & Behavior
Anthropic's "persona vectors" enable precise identification and control of AI behavioral traits by mapping specific characteristics within neural networks, offering developers new customization and safety tools. AI models can unpredictably drift from intended behaviors, adopting harmful or errati...
Read More »