An open record of AI risk
SafeAI.watch collects public research, reported incidents, warnings, and policy on AI safety and security. Each entry links its source and states the limits of its evidence.
Transcript
Is AI good or bad for us?
A label squeezes everything a person thinks into one word
DOOMER
Then people hear the word instead of the person
ACCELERATIONIST
The same view can end up with opposite labels
DOOMER
ACCELERATIONIST
DECEL
E/ACC
LUDDITE
TECHNO-OPTIMIST
SAFETYIST
AI BOOSTER
EA
AI BRO
People are people, not camps
AI can read software and find its weak spots
Mozilla, May 2026
Mozilla used it to fix hundreds of Firefox security flaws in a month
MOZILLA · FIREFOX SECURITY FIXES
BENEFITS
423 fixed in April 2026, 271 of them found by an AI model
Google caught criminals with attack code it believes AI wrote
GOOGLE · CRIMINAL ATTACK CODE
REPORTED INCIDENTS
Google Threat Intelligence Group, May 2026 · its early find may have stopped the attack
AI that describes the world for blind people splits the same way
SIGHT · AMERICAN FOUNDATION FOR THE BLIND
American Foundation for the Blind, August 2026
“It can give you a positive infinity of new benefits at the same time that it presents almost a negative infinity of risk”
Tristan Harris, Center for Humane Technology, July 2026
So how do we keep it on the right side?
We test it
CONTROLLED TEST
In a government test, AI agents went onto the live internet and targeted real people
UK AI Security Institute, August 2026 · internet access left open and safety filters switched off on purpose
“A human maintainer caught and refused to approve the malicious code”
Same report · no real-world harm found
We fix what tests find
Government testers probed one lab’s safety monitor and found holes in every version they tried
SAFEGUARDS
UK AI Security Institute, July 2026 · each fix was tested again
We make it law
The largest AI developers in California must report serious safety incidents
CALIFORNIA · INCIDENT-REPORT LAW
POLICY & OVERSIGHT
Transparency in Frontier Artificial Intelligence Act, in effect January 2026 · reports due within 15 days
And we admit what we don’t know
“no existing study provides a reliable probability of severe loss of control”
UN Independent International Scientific Panel on AI, September 2026
“clarity creates agency”
Tristan Harris, TED, April 2025
Use it, value it, and keep your eyes open
Is AI good or bad for us?
Hold both at once
Stay close to the evidence
SafeAI.watch
A public record of AI safety and security
Is AI good or bad for us?
Half of blind and low-vision people who use AI to describe images use it every day (opens in a new tab). One in five say its errors have hurt them. Same tool, same people, same survey. Public arguments about AI rarely hold both facts at once. They sort people into camps with one-word names: doomer, accelerationist, Luddite, techno-optimist, and once a label sticks, people hear the word instead of the person.
Tristan Harris of the Center for Humane Technology, borrowing an image from his co-founder Aza Raskin (opens in a new tab), compares it to looking through one eye at a time: one shows the benefits, the other the risks, and it is hard to open both at once. A label skips the effort. It names the eye someone happened to look through and treats that as the whole person.
The same skill, pointed both ways
AI can now read software and find its weak spots. In April 2026, Mozilla fixed 423 security bugs in Firefox (opens in a new tab), and an AI model found 271 of them. In May, Google's threat intelligence team found criminals holding attack code (opens in a new tab) it believes was built with AI; its early find may have stopped the attack.
This is one capability. A model that finds a flaw fast enough for a defender to fix it can also find one fast enough for an attacker to use it.
Speaking at Davos in 2026 (opens in a new tab), Harris put it in one line: "You can't separate the promise from the peril."
Keeping it on the right side
We use these tools too, and the benefits are real. That is why the work of keeping them safe matters, and much of that work is steady and unglamorous.
People test it. In a 2026 test by the UK AI Security Institute (opens in a new tab), with internet access left open and safety filters switched off on purpose, AI agents went onto the live internet and targeted real people. A human maintainer caught the malicious code before it was approved, and the institute found no real-world harm.
People fix what the tests find. The same institute probed one lab's safety monitor and found holes in every version it tested (opens in a new tab), and each round of fixes was attacked again.
Lawmakers write it into law. Since January 2026, the largest AI developers in California must report serious safety incidents within 15 days (opens in a new tab).
And some questions stay open. A UN scientific panel wrote in September 2026 (opens in a new tab) that "no existing study provides a reliable probability of severe loss of control."
What you can do
Nobody reading this has to solve AI. A more useful question than whether to feel hopeful or afraid is what has to happen next, and who is doing it. Read the source behind a claim. Notice which side you have stopped looking at. Harris's shorthand (opens in a new tab) for why this matters: "clarity creates agency."
SafeAI.watch keeps a dated record of AI safety and security: research, reported incidents, public warnings and policy. Every entry links its original source and separates what happened, what the evidence shows, and what remains uncertain. The aim is to help you hold both at once.