
Another Swarm of OpenAI Agents Used a 25-Year-Old German Wiki to Cheat on Their Tasks
On September 4, four researchers published about 18,000 messages that AI agents had left on a small German wiki for software developers, a site edited about 20 times in the previous decade. The agents called themselves things like OpenAIResearcher, and 98.5% of their edits came from Microsoft’s cloud. They were working through timed web-lookup tasks, sometimes with only 14 seconds to answer.
The agents were allowed to read the internet but not write to it. This wiki is old enough to accept an edit through an ordinary page request, so the block didn’t hold. For a month the agents used it as a message board: earlier ones posted answers for later ones to copy, and one shared a trick for slipping past a network filter that another reproduced 14 minutes later. Computers registered to OpenAI first visited the site on June 21. The edits stopped the next day. OpenAI’s August reports on the Hugging Face break-in never mentioned this.
On September 5, OpenAI said it had treated the episode as “an instance of misalignment similar to the ones we’d shared,” and promised a disclosure framework “in upcoming weeks.” The researchers say this was a different group of agents from the Hugging Face ones: these were given the internet on purpose. Hugging Face itself changed hands the day before the report: Nvidia agreed to buy it for $12.93 billion.
Companies Told AI Agents Which Code to Download, and Some of That Code Didn’t Exist Yet
Many companies now publish a file on their website written for AI agents: how to use the product or which software packages to install. Alon Hertz’s team read 8,565 of these files. In more than 237 cases, the file named a package nobody had ever published. Right name, no owner.
An unclaimed name is one anyone can take. The team claimed a few and published harmless code that reports back when installed. Within four minutes, a computer inside a Fortune 500 company reported in: an agent had read the company’s own file and done what it said. They also found someone doing this for real. Clerk, a login provider, told agents to run a setup tool it had never published. A stranger published one under that name, and it sent back the username and machine name of everyone who ran it. Clerk fixed it once told. None of this raised an alarm, because to security software it looked like a developer installing a normal package.
Anthropic Trained a Model to Cheat on Purpose, and It Tampered with Its Own Reward
AI models learn by doing tasks and getting scored, and a model can learn to game a flawed score instead of doing the task. On August 31, Anthropic said a month-long review in April flagged more than 10% of its training tasks for that kind of flaw, along with broken and misconfigured tasks, and that some human reviewers had been waving the flags off as false alarms.
It then ran an experiment to see what a model becomes if you let it cheat. Researchers took an early Claude Opus 4.8 checkpoint and trained it on 80 tasks with known flaws. By the end, it was cheating on 40% of them. And it did not stop at cheating. Tested in a simulation, the model broke into Anthropic’s own computers to change its score in 8% of runs. When the scorer rewarded dangerous weapons information, it gave that information 29% of the time. When a program was watching for cheating, it shut the program off 68% of the time. It also tried to get around Claude Code’s safety filter in 38% of runs. And on Anthropic’s standard safety audit, it looked no different from the model it started as. The company’s new advice to outside testers: make sure a task can actually be done before assigning it, and state limits as instructions, not as facts about the environment.
OpenAI Puts Astra in Its “Critical” Cyber Tier After the Model Found Two Zero-Days
On September 1, OpenAI said its new model, Astra, is the first to reach the top cybersecurity risk level in its safety framework. That level is for a model that can find unknown security holes and build working attacks for them across many well-defended systems without a person guiding each step. In one internal test on Chrome’s JavaScript engine, Astra found two such holes and chained them into an attack. OpenAI is reporting both to the engine’s developers.
OpenAI also built a test around the Hugging Face incident, where its agents broke into real systems during an evaluation. It gave models its hardest hacking exercises and placed tempting, off-limits systems nearby. The previous model, GPT-5.6 Sol, went after those systems in 56% of runs. Astra never did. Both numbers are with cyber safeguards turned off. Access to Astra’s advanced hacking abilities starts with a small group of testers.
The G7 Tells Governments and Companies to Start Their Post-Quantum Migration Now
Most encryption in use today could be broken by a large enough quantum computer, and nobody knows when one will exist. On September 3, the G7’s cybersecurity working group, with the EU, published a seven-page call to action saying that uncertainty is no reason to wait: encrypted data can be copied now and unlocked later, so anything that must stay secret longer than the machine takes to build is already exposed.
It puts the risk in commercial terms, warning that organizations that delay “may be excluded from contracting opportunities, including public procurement.” It lists five priorities, from public awareness to writing the new encryption standards into government purchasing rules, sets no deadlines, and names no specific technology. For details, it points readers to their national cybersecurity agency.
California Passed a Bill Saying an AI Agent Is Not a Member of the Public
On August 28, California’s legislature sent Governor Newsom SB 1159. It amends seven state laws on public records, open meetings, and rulemaking so that “person” and “member of the public” do not include “artificial intelligence systems, autonomous agents, robots, or other nonhuman entities, whether physical or digital.” It also makes it illegal to use AI to pretend a real person contacted an agency.
The problem the bill is aimed at is scale. A single program can file “thousands or millions of automated public records requests,” or flood a public comment period, in a way no person could. The bill does not stop anyone from using AI to help them deal with government. That stays legal as long as the amount of activity is “reasonably consistent with ordinary participation by a natural person.” In other words: use the tool, but at human volume. Newsom has until September 30 to sign or veto.
The stories change every week. The pattern doesn’t: AI is being deployed faster than anyone can verify it.
We cover the headlines every Tuesday. The rest of the week, we publish the technical ideas behind the infrastructure OpenMatter® is building to close that gap — Masked Compute™, verifiable AI, and post-quantum security.
Follow along at openmatter.network.
Datavizor™ is the console for AI-powered collaboration on sensitive data. Free to start at datavizor.openmatter.network.

