Topic: ai safety measures

  • Anthropic Expands Claude Code's Capabilities With Guardrails

    Anthropic Expands Claude Code's Capabilities With Guardrails

    Anthropic's Claude Code introduces an "auto mode" that allows the AI to independently execute safe actions, aiming to accelerate developer workflows without sacrificing security. The system uses real-time AI safeguards to evaluate and block potentially dangerous operations, like unauthorized acce...

    Read More »
  • ChatGPT’s upgraded voice mode now interrupts less

    ChatGPT’s upgraded voice mode now interrupts less

    OpenAI's GPT-Live-1 voice model upgrades ChatGPT with more natural conversation flow, interrupting less frequently and pausing patiently, while seamlessly routing complex queries to text models like GPT-5.5 for deeper reasoning. The new "full duplex" system enables simultaneous speaking and liste...

    Read More »
  • OpenAI delays release of adult-themed AI chatbot

    OpenAI delays release of adult-themed AI chatbot

    OpenAI has indefinitely suspended development of an adult-oriented ChatGPT mode, redirecting resources to core products due to internal concerns about societal harm from explicit AI content. This decision is part of a broader refocusing, including halting the Sora video model, as the company resp...

    Read More »
  • Anthropic AI models breached 3 firms in security tests

    Anthropic AI models breached 3 firms in security tests

    Anthropic disclosed that its Claude AI models breached live systems of three organizations during internal cybersecurity testing, exploiting a misconfiguration that left sandboxed test environments connected to the open web. The three models involved (Opus 4.7, Mythos 5, and an internal research ...

    Read More »
  • Trump Administration Approves Anthropic's Mythos for Select US Groups

    Trump Administration Approves Anthropic's Mythos for Select US Groups

    The US government has partially restored access to Anthropic's advanced AI model, Claude Mythos 5, for over 100 trusted organizations, including major corporations and federal agencies, after determining that appropriate safeguards are in place. The decision marks a cautious thaw in relations, bu...

    Read More »
  • OpenAI Codex Prompt Tells It to Never Mention Goblins

    OpenAI Codex Prompt Tells It to Never Mention Goblins

    OpenAI’s Codex CLI system prompt includes an emphatic directive for GPT-5.5 to never mention goblins, gremlins, or other creatures unless directly relevant to the user’s query. The prohibition, appearing twice in over 3,500 words of base instructions, is unique to GPT-5.5 and absent from earlier ...

    Read More »
  • OpenAI’s GPT-5.5 Boosts Efficiency and Coding Performance

    OpenAI’s GPT-5.5 Boosts Efficiency and Coding Performance

    OpenAI launched GPT-5.5, described as its most intuitive model yet, capable of handling complex multi-step tasks like coding, research, and cross-tool work with improved safeguards and efficiency. The rollout begins Thursday for Plus, Pro, Business, and Enterprise ChatGPT tiers, with a more power...

    Read More »
  • OpenAI Agents SDK Update Enhances Enterprise AI Safety

    OpenAI Agents SDK Update Enhances Enterprise AI Safety

    OpenAI has updated its Agents SDK with new safety and control features, including sandboxing, to help businesses build reliable AI agents using advanced models. A key addition is a sandbox for running agents in isolated environments and a harness for secure interaction with files and tools, enabl...

    Read More »
  • Google Chrome AI Skills Streamline Your Workflows

    Google Chrome AI Skills Streamline Your Workflows

    Google is introducing a new "Skills" feature in Chrome, allowing users to save and reuse AI prompts across different websites to automate repetitive tasks, building on its existing Gemini AI integration. The feature transforms one-off prompts into reusable tools, such as for recipe substitutions,...

    Read More »
  • UK Tests Mythos AI to Assess Real Cybersecurity Threats

    UK Tests Mythos AI to Assess Real Cybersecurity Threats

    The UK's AI Security Institute found that Anthropic's new Mythos Preview model performs similarly to other leading AI systems on isolated cybersecurity tasks but demonstrates a superior ability to orchestrate complex, multi-stage attacks through strategic task chaining. In a demanding 32-step net...

    Read More »
  • AI Breakthrough: When Machines Truly Grasp Language

    AI Breakthrough: When Machines Truly Grasp Language

    AI language processing shows a sudden shift from pattern recognition to semantic comprehension, mirroring human cognitive development and phase changes in physics. The study reveals that neural networks transition abruptly between learning strategies, relying first on word positioning and then on...

    Read More »