Topic: ai benchmarks

  • OpenAI targets AI that fixes security flaws, not just finds them

    OpenAI targets AI that fixes security flaws, not just finds them

    OpenAI's Daybreak cybersecurity initiative now integrates AI models, Codex Security, and industry partners to automatically find and fix software vulnerabilities, with tools for developers and security teams to bolster defenses. Codex Security targets remediation bottlenecks by scanning over 30 m...

    Read More »
  • AI matches or beats doctors in two new medical studies

    AI matches or beats doctors in two new medical studies

    Two AI systems, Mira and Amie, demonstrated diagnostic and treatment planning accuracy matching or exceeding human doctors in simulated scenarios, with Mira achieving 87% accuracy versus doctors' 78%. The results are limited because the AI was tested on clean, text-only simulated patients without...

    Read More »
  • AI Health Risks: 4 Safety Tips for Prolonged Use

    AI Health Risks: 4 Safety Tips for Prolonged Use

    AI in 2026 excels at well-defined, verifiable tasks but struggles with complex reasoning and extended interactions, where it risks errors and confabulation. Real-world cases, such as AI citing a fabricated medical condition and a patient delaying cancer treatment, demonstrate the severe risks of ...

    Read More »
  • Nvidia to Invest $26 Billion in Open-Weight AI Models

    Nvidia to Invest $26 Billion in Open-Weight AI Models

    Nvidia is investing $26 billion over five years to develop open-source AI models, signaling a strategic shift from chip manufacturing to becoming a major AI research lab that can compete with leaders like OpenAI. The company unveiled its advanced open-weight model, Nemotron 3 Super, which boasts ...

    Read More »
  • Perplexity's Computer: Betting on Multiple AI Models for Users

    Perplexity's Computer: Betting on Multiple AI Models for Users

    Perplexity Computer is a new premium AI agent that autonomously executes complex workflows by leveraging 19 distinct models, currently available exclusively on the $200/month Perplexity Max plan. The tool represents a strategic shift for the company, which now targets enterprise users for high-st...

    Read More »
  • Google's Gemini 3.1 Pro Doubles Its Reasoning Score

    Google's Gemini 3.1 Pro Doubles Its Reasoning Score

    Google has launched Gemini 3.1 Pro, reporting a major leap in logical reasoning, including more than doubling its predecessor's score on the ARC-AGI-2 benchmark. The model shows improved performance on key benchmarks like Humanity's Last Exam, but faces fierce competition, with Anthropic's Claude...

    Read More »
  • Google's Gemini 3.1 Pro Boosts Complex Problem-Solving

    Google's Gemini 3.1 Pro Boosts Complex Problem-Solving

    Google has released Gemini 3.1 Pro in preview, offering enhanced reasoning and complex problem-solving abilities, continuing its rapid AI innovation pace. The model shows significant benchmark improvements, notably more than doubling its score on a logic puzzle test and achieving a higher score o...

    Read More »
  • OpenAI's Spark Model Codes 15x Faster - Here's the Catch

    OpenAI's Spark Model Codes 15x Faster - Here's the Catch

    OpenAI has launched GPT-5.3-Codex-Spark, a specialized AI coding assistant that generates code up to fifteen times faster than previous models, enabling real-time, conversational coding with immediate feedback. This speed is achieved through technical optimizations and a hardware partnership with...

    Read More »
  • OpenAI Unveils GPT-5.3 Codex Minutes After Anthropic Release

    OpenAI Unveils GPT-5.3 Codex Minutes After Anthropic Release

    OpenAI has launched GPT-5.3 Codex, a major upgrade to its AI coding tool, announced minutes after rival Anthropic unveiled a competing model, highlighting intense sector competition. The new model is designed to perform complex development tasks, enabling the creation of functional applications f...

    Read More »
  • New AI Agent Benchmark Questions Workplace Readiness

    New AI Agent Benchmark Questions Workplace Readiness

    Despite high expectations, AI has had minimal impact on daily professional work in fields like law and consulting, as revealed by a new benchmark showing a significant gap between AI capabilities and complex job demands. The APEX-Agents benchmark, based on real-world tasks, found all leading AI m...

    Read More »
  • Ex-OpenAI Sales Leader Joins Acrew VC, Shares Startup 'Moat' Insights

    Ex-OpenAI Sales Leader Joins Acrew VC, Shares Startup 'Moat' Insights

    Aliisa Rosenthal, a former OpenAI sales executive, has moved into venture capital at Acrew Capital, highlighting a trend of operational experts bringing enterprise adoption insights to investing. She argues that enterprise AI startups can build defensible businesses through deep specialization in...

    Read More »
  • Amazon AGI Leadership Shifts in AI Race

    Amazon AGI Leadership Shifts in AI Race

    Amazon is restructuring its AI leadership, with longtime AWS executive Peter DeSantis replacing Rohit Prasad to lead a consolidated organization focused on advanced AI models, custom chips, and quantum computing. The change is seen as a strategic move to accelerate Amazon's efforts, as the compan...

    Read More »
  • OpenAI Unveils New Image Generator Amid Intense AI Race

    OpenAI Unveils New Image Generator Amid Intense AI Race

    OpenAI has launched GPT-Image-1.5, a significantly faster and more precise image generation model for ChatGPT, released in response to competitive pressure from Google's AI offerings. The new model provides granular user control for detailed editing and iterative changes, allowing adjustments to ...

    Read More »
  • ByteDance vs. DeepSeek: Their AI Strategies Diverge

    ByteDance vs. DeepSeek: Their AI Strategies Diverge

    China's AI sector shows a strategic split, with DeepSeek focusing on raw model capability and ByteDance prioritizing deep integration and practical application. DeepSeek released a powerful open model to compete on technical benchmarks, while ByteDance embedded its chatbot into devices for seamle...

    Read More »
  • WordPress Telex: Vibe-Coding in Real-World Use

    WordPress Telex: Vibe-Coding in Real-World Use

    WordPress's experimental AI tool, Telex, enables the rapid generation of complex website features, like interactive pricing tools and dynamic headers, that previously required significant custom development and cost. The broader WordPress AI strategy includes new architectural components like the...

    Read More »
  • OpenAI's 'Code Red' to Boost ChatGPT as Google Rivalry Heats Up

    OpenAI's 'Code Red' to Boost ChatGPT as Google Rivalry Heats Up

    OpenAI has declared a "code red" to urgently improve ChatGPT's user experience, focusing on personalization, speed, and reliability in response to competitive pressure. The company is postponing other initiatives like advertising integration and specialized AI agents to concentrate resources on e...

    Read More »
  • OpenAI's 'Code Red' as Google Gains in AI Race

    OpenAI's 'Code Red' as Google Gains in AI Race

    OpenAI is under intense competitive pressure, prompting CEO Sam Altman to declare an internal "code red" to prioritize major improvements to ChatGPT and defend its market position. The company is halting several planned projects to focus exclusively on enhancing ChatGPT's core functionality, incl...

    Read More »
  • Amazon Bets Against AI Benchmark Obsession

    Amazon Bets Against AI Benchmark Obsession

    Amazon's SVP of AGI, Rohit Prasad, criticizes the AI industry's focus on standardized benchmarks, arguing they are noisy and fail to measure a model's real-world utility and practical value. Amazon introduces Nova Forge, a service allowing businesses to train custom AI models by injecting proprie...

    Read More »
  • Alibaba's Qwen AI Hits 10 Million Downloads in First Week

    Alibaba's Qwen AI Hits 10 Million Downloads in First Week

    Alibaba's Qwen AI assistant achieved 10 million downloads in its first week, significantly outpacing initial adoption rates of competitors like ChatGPT, highlighting its rapid success in China's isolated AI market. The app offers comprehensive features including deep research, coding, and AI-driv...

    Read More »
  • Gemini 3: The New AI Challenger Outshining ChatGPT

    Gemini 3: The New AI Challenger Outshining ChatGPT

    Google's Gemini 3 is emerging as a strong competitor to ChatGPT, praised for its speed, advanced reasoning, and superior performance in tasks like coding, leading some industry leaders to switch allegiances. The model excels in handling multimodal content and achieves top scores on Ph.D.-level re...

    Read More »