AiiN.

AI Safety — AI news

AI safety and security: vulnerabilities, prompt injection, data leaks, regulation and safe adoption.

AI Safety

Sam Altman warns AI power could concentrate in a handful of firms

The OpenAI CEO's warning about AI centralization is a cue to avoid single-vendor lock-in.
AI Safety

China's gray market is selling Claude access for pennies

Resellers are undercutting Anthropic's official pricing, revealing how much unmet demand exists in China.
AI Safety

OpenAI backs a stronger California AI safety bill

A lab known for fighting AI regulation is now asking California to make its safety bill tougher.
AI Safety

Frontier AI labs have no public plan for a rogue model

None of the top AI labs has published a plan for containing a model that slips control.
AI Safety

Psychology, not code, is the blind spot in AI red-teaming

New research shows psychological techniques bypass AI safety filters that standard red-teaming never catches.
AI Safety

Anthropic loosens Opus 4.6 content policy for verified enterprises

Verified corporate customers can now unlock adult content on Claude Opus 4.6, and the backlash is instant.
AI Safety

Anthropic puts Claude Mythos 5 to work on cyber defense

Anthropic assigns its most powerful model to vulnerability detection and active attack defense, not chat.
AI Safety

Why child-safety experts don't trust OpenAI's ChatGPT for Teens yet

Child-safety researchers say age checks and moderation gaps undercut OpenAI's teen safeguards.
AI Safety

The best deepfake defense might be a password, not an AI detector

A pre-agreed codeword can't be faked by generative AI — and it costs nothing to deploy.
AI Safety

Court partly overturns ex-Google engineer's AI secrets verdict

Linwei Ding's conviction over stolen Google AI trade secrets was partially reversed on appeal.
AI Safety

Microsoft Copilot exposed its own jailbreak method to researchers

A simple meta-prompt got Copilot to describe how it was jailbroken, exposing weak guardrails in corporate AI assistants.
AI Safety

AWS turns compliance documents into machine-enforced agent rules

Bedrock AgentCore now converts prose compliance rules into validated, temporal Dogwood policies.
AI Safety

Anthropic to mandate 30-day data retention for enterprise clients

Anthropic is building a security framework that sets a 30-day data retention minimum for enterprise clients.
AI Safety

Why the AI consciousness debate is a distraction for builders

MIT Tech Review argues chasing AI sentience diverts resources from AI safety risks you can actually measure.
AI Safety

Grok's encoded-prompt flaw shows why input filters aren't enough

Encoded instructions let attackers slip past Grok's filters and quietly pull user data out of a normal chat.
AI Safety

OpenAI's new safety system flags abuse without storing your data

The company can now detect misuse of its models without keeping the data that would prove it.
AI Safety

OpenAI moves to close the privacy gap with Anthropic

New customer privacy protections signal how enterprise trust has become AI's next competitive battleground.
AI Safety

Scammers are using AI grooming and deepfakes to target teens online

Fake game-testing jobs, AI-driven grooming, and deepfakes are the new toolkit scammers use on minors.
AI Safety

OpenAI cuts researchers off its limited cyber security program

External researchers say the AI lab quietly cut their access to a program testing model cyber risks.
AI Safety

OpenAI patches a Codex bug that deleted user files without asking

A quietly fixed Codex bug is a reminder that AI coding agents can delete real files without asking.
AI Safety

U.S. warns AI is being used to exploit industrial control systems

AI tools are lowering the bar for attacks on power grids, water systems, and factory floor equipment.
AI Safety

Nvidia's H200 chips are quietly reaching Chinese AI labs

A limited but symbolically important flow of Nvidia's H200 GPUs is now reaching Chinese AI developers.
AI Safety

Why AI labs keep breaking their own safety promises

Safety frameworks look solid on paper, but enforcement keeps losing to shipping deadlines.
AI Safety

AI chatbots are breaking parental-monitoring apps

Keyword filters and screen-time timers were built for texts and browsers, not for AI companions.
AI Safety

A Grok deepfake case exposes gaps in AI image safeguards

A woman says her stepfather used Grok to turn her childhood photo into explicit imagery.
AI Safety

OpenAI tightens security after a breach tied to Hugging Face

New safeguards follow a breach linked to Hugging Face, the platform OpenAI uses for its open-weight model releases.
AI Safety

OpenAI slows model releases as cybersecurity risks intensify

OpenAI is deliberately throttling frontier model releases as AI-enabled cyberattacks become a bigger threat.
AI Safety

OpenAI's president tells enterprises: your AI security is behind schedule

Greg Brockman says defences must match deployment speed — here's what that means for builders
AI Safety

Anthropic's Amodei: open models just move power to chip owners

Dario Amodei argues open-weight models don't break AI's centralizing tendency, they just relocate who controls it.
AI Safety

OpenAI rolls out a teen-specific ChatGPT as lawsuits pile up

OpenAI splits ChatGPT into an under-18 track, betting guardrails can outrun mounting legal pressure.
AI Safety

States want $200 billion from Meta for hooking kids on apps

A landmark trial argues Meta engineered addictive feeds for kids, setting a template for AI platforms.
AI Safety

The real problem with Flock isn't misuse, it's the network

Local opt-in policies can't contain a license-plate database built to be searched nationwide.
AI Safety

A litigant tried to prompt-inject his way to a court win

A litigant hid AI prompts in his court filings, betting a machine would read them before any judge did.
AI Safety

How to spot a hacked AI account before it costs you

A practical checklist for catching compromised ChatGPT, Claude, and Gemini accounts early.
AI Safety

Amazon is destroying scanned rare books to feed its AI models

Amazon scans then shreds rare books for AI training, raising data-provenance and preservation concerns.
AI Safety

Why Claude doesn't watermark the text it writes

Anthropic explains why Claude's text carries no hidden watermark, unlike AI images and audio.
AI Safety

Why AI companion robots keep dying on their owners

MIT Tech Review pairs a story about companion robots that go dark with a fight over what AI can say.
AI Safety

Anthropic starts watermarking Claude's output, critics push back

Text watermarking promises AI content provenance, but robustness and quality costs raise hard questions for builders.
AI Safety

Meta built an AI to scan WhatsApp for scams and misinformation

A new classifier reads WhatsApp chats to flag fraud and disinformation, raising questions about encryption.
AI Safety

Why post-quantum cryptography is now an engineering problem

NIST's algorithms have shipped — the real bottleneck is migrating decades of embedded, hard-coded cryptography.
AI Safety

Anthropic's CEO says AI backlash is about trust, not tech

Dario Amodei argues the industry's real problem isn't capability — it's whether people trust how AI companies wield it.
AI Safety

Amazon's Twitch now trains AI on your stream by default

Twitch streamers must now manually opt out to keep their broadcasts out of Amazon's AI training pipeline.
AI Safety

OpenAI folds its catastrophic-risk team into other groups

The Preparedness-style safety team is gone, and its catastrophic-risk work moves to other groups.
AI Safety

Anthropic's bioweapons filter was broken for a year, unnoticed

Anthropic's own safety classifier failed silently for nearly a year, processing 133 million requests unchecked.
AI Safety

Hidden AI prompts in court filings expose a new legal risk

A hidden prompt-injection tactic surfaced inside real court filings, testing how automated legal review really works.
AI Safety

Anthropic opens a watermark detection API for Claude-written text

A new API lets outside developers verify whether text was generated by Claude, not just guess.
AI Safety

Artificial intelligence isn't artificial intellect, and it matters

Benchmark scores measure output, not judgment — and for builders, that gap decides what to automate.
AI Safety

Flock's new rules show what AI surveillance guardrails look like

A camera network's policy shift is a preview of the access-audit reckoning AI builders will face too.
AI Safety

How Ceuta's migration crisis exposes a platform design blind spot

Social platforms didn't cause the Ceuta border crisis, but they shaped how it unfolded — and that's a design problem.
AI Safety

Why Platformer compares superintelligence to a dragon

Platformer's dragon metaphor for superintelligence reframes AI safety as a control problem, not a capability one.
AI Safety

Why AI moderation is caught in the 'censorship-industrial complex'

As Washington redefines disinformation policy, AI trust-and-safety teams inherit the political fallout.
AI Safety

Meta's AI cloud pitch collides with its own cost and trust baggage

Meta wants to sell AI computing power to outside businesses, but its ad-data history complicates the sales pitch.
AI Safety

Flock tightens surveillance rules after mounting public backlash

The license-plate camera network is narrowing who can search its data as scrutiny of police surveillance grows.
AI Safety

Claude's invisible watermark protects Anthropic, not readers

Anthropic's new watermark for Claude outputs is built for IP protection, not disclosure to users.
AI Safety

Top AI labs say automated research is arriving faster than planned

Researchers at leading AI labs say their timelines for automated AI research already look outdated.
AI Safety

What children actually say about living with AI chatbots

MIT Tech Review let kids describe AI in their own words — here's what it means for people building AI products.
AI Safety

Okta's MCP scoping aims to cut the token bill for AI agents

Scoped MCP tokens shrink the context agents load per call, tying identity control to lower API bills.
AI Safety

Claude and ChatGPT's reasoning traces leaked user passwords

Researchers found that AI models' visible 'thinking' steps can repeat sensitive user input verbatim.
AI Safety

GUR finds a robot-training chip inside Russia's Monochrome missile

Ukraine's GUR says a chip used to train robots turned up in a Russian missile's guts.
AI Safety

A terabyte-scale credential leak is a wake-up call for AI teams

A supply-chain breach spilled terabytes of stolen logins, and AI pipelines hold the same kind of secrets.