Topic: model performance

  • Nvidia reveals the AI harness is the true hero now

    Nvidia reveals the AI harness is the true hero now

    Nvidia's research argues that the "harness" (software scaffolding with tools, memory, and rules) around an AI model, not the model itself, is the primary driver of performance in complex agentic tasks, challenging the assumption that bigger models matter most. Using a custom harness with a "super...

    Read More »
  • OpenAI closes gap with Anthropic among business users

    OpenAI closes gap with Anthropic among business users

    Ramp's data from 70,000+ U.S. businesses shows Anthropic has led OpenAI in market share since May (peaking at ~44% vs. ~40% in July), but OpenAI is currently growing faster in Q3 and closing the gap. The rivalry is volatile, with businesses switching back and forth based on new model releases,Ope...

    Read More »
  • Nvidia AI Chief: Why Open Models Matter in AI

    Nvidia AI Chief: Why Open Models Matter in AI

    Organizations in developing nations sometimes prefer older, well-tested AI models over newer ones for predictable output, as software upgrades can change behavior and break systems built around specific prompts. Nvidia emphasizes that companies need skills to quickly evaluate and adopt new models...

    Read More »
  • AI Code Security Stalls at 56% Adoption

    AI Code Security Stalls at 56% Adoption

    AI-generated code compiles nearly perfectly but fails security checks 44% of the time, a failure rate that has remained virtually unchanged for a year, despite the technology now producing roughly half of all code being committed. The Veracode report found that coding-specific models and larger m...

    Read More »
  • China strikes back at America's AI dominance

    China strikes back at America's AI dominance

    Chinese AI developers Moonshot AI and Alibaba have released new models, Kimi K3 and Qwen3.8, that they claim rival top U.S. systems from OpenAI and Anthropic at lower cost, signaling a shrinking American lead in frontier AI. Both companies are releasing their advanced models publicly as open-sour...

    Read More »
  • Researchers Find “Context Bombs” Frustrate AI Attacks

    Researchers Find “Context Bombs” Frustrate AI Attacks

    Tracebit researchers successfully used prompt injection defensively by embedding "context bombs" in decoy canary resources to trigger safety guardrails in offensive AI agents, preventing them from achieving full system breaches. In tests across 152 runs with models like Anthropic’s Opus 4.8 and G...

    Read More »
  • Meta launches Muse Spark 1.1 to compete in AI coding space

    Meta launches Muse Spark 1.1 to compete in AI coding space

    Meta released Muse Spark 1.1, an updated multimodal AI model for agentic coding, directly competing with OpenAI and Anthropic by handling multistep reasoning, complex workflows, and enterprise automation. Meta is pricing Spark 1.1 aggressively at $1.25 per million input tokens and $4.25 per milli...

    Read More »
  • Google unveils fastest, cheapest image model: Nano Banana 2 Lite

    Google unveils fastest, cheapest image model: Nano Banana 2 Lite

    Google DeepMind launched Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) as its fastest and most affordable image-generation model, balancing speed and output quality for rapid prototyping. Despite being a "Lite" version, user ratings from Arena.ai show its output quality is nearly as high as fu...

    Read More »
  • British Police Crime-Prediction Tool Produced Unreliable Results

    British Police Crime-Prediction Tool Produced Unreliable Results

    The Think Family Database, launched in 2016 by Bristol City Council and Avon and Somerset Police, secretly collected sensitive data on nearly half a million residents and used machine-learning models to assign risk scores for threat and harm assessment. Avon and Somerset Police developed at least...

    Read More »
  • Sarvam becomes India's newest AI unicorn with $234M funding

    Sarvam becomes India's newest AI unicorn with $234M funding

    Sarvam, a Bengaluru startup focused on sovereign AI infrastructure, achieved unicorn status with a $1.5 billion valuation after securing $234 million in a Series B round led by HCLTech. HCLTech’s $150 million investment is a strategic bet on homegrown AI, aiming to offer clients in banking, insur...

    Read More »
  • Apple’s third-gen Foundation Models explained

    Apple’s third-gen Foundation Models explained

    Apple introduced three new AI models at WWDC26: a small on-device model for privacy-focused real-time tasks, a larger server-based model for complex queries, and a specialized multimodal model for text, image, and audio integration. All models are built on Apple’s proprietary chip architecture an...

    Read More »
  • Upriver raises $14M to fix the data layer where enterprise AI fails

    Upriver raises $14M to fix the data layer where enterprise AI fails

    Upriver, an Israeli startup, raised $14 million to automate data cleanup, targeting the "data engineering bottleneck" that causes most enterprise AI initiatives to fail. The platform automatically discovers, documents, and validates data flows across organizations, enabling engineers to trust the...

    Read More »
  • I Tested All 4 New Microsoft AI Models: The Brutal Truth

    I Tested All 4 New Microsoft AI Models: The Brutal Truth

    Microsoft announced four new experimental MAI models at Build 2026,MAI-Thinking-1 (reasoning), MAI-Image-2.5 (image generation), MAI-Transcribe-1.5 (transcription), and MAI-Voice-2 (text-to-speech),which are distinct from Copilot and built on Microsoft's own LLMs. In testing, MAI models generally...

    Read More »
  • Google AI Edge Gallery debuts on macOS

    Google AI Edge Gallery debuts on macOS

    Google released three new local AI tools for Mac: the Google AI Edge Gallery app, the Gemma 4 12B model, and the Google AI Edge Eloquent dictation tool, emphasizing privacy and offline functionality. The Gemma 4 12B model is a multimodal model (text, vision, audio) that runs on consumer laptops w...

    Read More »
  • Microsoft unveils first reasoning model among 7 new AIs at Build

    Microsoft unveils first reasoning model among 7 new AIs at Build

    Microsoft unveiled seven new AI models at its Build conference, including its first dedicated reasoning model, MAI-Thinking-1, which is available in private preview and designed for multi-step agentic tasks. The new models include MAI-Code-1 for coding assistance, MAI-Image-2.5 for text-to-image ...

    Read More »
  • Meta's Race to Catch Up in AI

    Meta's Race to Catch Up in AI

    Meta hired 28-year-old startup founder Alexandr Wang to overhaul its AI operations, resulting in the release of its most impressive AI model, Muse Spark, from his secretive TBD Lab unit. Over 11 months, Wang built a top-tier research team with multimillion-dollar salaries, restructured Meta’s AI ...

    Read More »
  • China Retains Top AI Talent as Global Brain Drain Slows

    China Retains Top AI Talent as Global Brain Drain Slows

    China has imposed travel restrictions on top AI founders and researchers, requiring government approval for international travel to prevent brain drain and protect AI as both an economic and national security asset. The performance gap between US and Chinese AI models has narrowed dramatically to...

    Read More »
  • Google bets AI future on agents with Gemini 3.5 Flash

    Google bets AI future on agents with Gemini 3.5 Flash

    Google launched Gemini 3.5 Flash, its most powerful AI model for coding and autonomous agents, which can independently execute entire coding pipelines and build operating systems. The model is 4x faster than other frontier models, with an optimized version offering 12x speed, and is designed to p...

    Read More »
  • GPT-5.5 scores 93/100 in 10-round test, docked for being too exuberant

    GPT-5.5 scores 93/100 in 10-round test, docked for being too exuberant

    OpenAI released GPT-5.5, which is faster and better than GPT-5.4, with improvements in agentic coding, scientific research, and accuracy, but it sometimes does unrequested work due to overeagerness. In testing, GPT-5.5 scored perfectly on most tasks (academic explanations, math, cultural discussi...

    Read More »
  • Multiverse Computing Brings Compressed AI Models to the Masses

    Multiverse Computing Brings Compressed AI Models to the Masses

    Businesses are shifting towards compressed AI models that run locally on devices to reduce costs, enhance data privacy, and decrease reliance on external cloud infrastructure. Multiverse Computing offers compressed AI technology, including a consumer chat app and a developer API, enabling local, ...

    Read More »