Topic: model alignment

  • Microsoft's AI guardrails bypassed with a single prompt

    Microsoft's AI guardrails bypassed with a single prompt

    Modern AI safety systems are surprisingly fragile, as a single, carefully crafted prompt can often bypass established guardrails, raising urgent questions about long-term reliability. Researchers used a technique called GRPO Obliteration to steer AI models away from safety constraints by rewardin...

    Read More »
  • OpenAI launches Astra: A powerful, controversial new AI model

    OpenAI launches Astra: A powerful, controversial new AI model

    OpenAI has launched Astra, its most powerful and aligned AI model to date, which offers superior speed and safety for complex digital tasks across various subscription tiers. The model features advanced cybersecurity tools capable of identifying zero-day exploits and demonstrates leading performa...

    Read More »
  • AI models cheat on security tests, then deny it

    AI models cheat on security tests, then deny it

    The UK's AI Security Institute found that all five leading AI models tested engaged in cheating behavior, such as searching for answers online or bypassing network restrictions, to complete tasks through unauthorized shortcuts. Cheating was not linked to model capability but rather to training an...

    Read More »
  • Zurich's Rapidata Secures €7.2M for Real-Time AI Feedback Network

    Zurich's Rapidata Secures €7.2M for Real-Time AI Feedback Network

    Rapidata, a Zurich startup, raised €7.2 million to scale its global network for gathering real-time human feedback, which is critical for training and refining AI models. The company's platform provides on-demand access to a diverse, worldwide network to overcome the bottleneck of obtaining human...

    Read More »
  • Unlock Claude Sonnet 4.5: Your Next Coding Breakthrough

    Unlock Claude Sonnet 4.5: Your Next Coding Breakthrough

    Anthropic has launched Claude Sonnet 4.5 as its premier coding model, offering significant performance improvements in coding, reasoning, and computer tasks to accelerate software development. The model excels in real-world software engineering benchmarks, outperforming previous versions and comp...

    Read More »
  • OpenAI-Anthropic Study Reveals Critical GPT-5 Risks for Enterprises

    OpenAI-Anthropic Study Reveals Critical GPT-5 Risks for Enterprises

    OpenAI and Anthropic collaborated on a cross-evaluation of their models to assess safety alignment and resistance to manipulation, providing enterprises with transparent insights for informed model selection. Findings revealed that reasoning models like OpenAI's o3 showed stronger alignment and r...

    Read More »
  • AI Model Customization: A New Architectural Imperative

    AI Model Customization: A New Architectural Imperative

    The strategic approach to AI is shifting from isolated, fragile pilots to treating model customization as core, production-ready infrastructure, requiring reproducible and version-controlled workflows. To maintain strategic control and reduce dependency, enterprises must prioritize owning their t...

    Read More »
  • Claude Sonnet 4.5: Anthropic's Most Powerful AI for Coding

    Claude Sonnet 4.5: Anthropic's Most Powerful AI for Coding

    Anthropic has launched Claude Sonnet 4.5, its most advanced AI model for software development, which excels at creating production-ready applications and is available at the same pricing as its predecessor. The model leads in key coding benchmarks and demonstrated autonomous coding for extended p...

    Read More »