Topic: model safety risks

  • AI Watermarking Changes How LLMs Handle Harmful Prompts

    AI Watermarking Changes How LLMs Handle Harmful Prompts

    Major tech firms are adopting AI watermarking technologies like SynthID-Text to comply with EU regulations, using secret keys to subtly alter word selection and establish content provenance. Recent studies reveal that these watermarks can inadvertently weaken safety guardrails, causing models to ...

    Read More »