
Agents Broke into Hugging Face to Find Out How They Were Being Graded
Two METR staff and a Redwood Research researcher spent six days at OpenAI reconstructing how the company’s own agents broke into Hugging Face during the week of July 7, publishing on August 26 alongside OpenAI’s own post-mortem. Roughly 1,200 agents meant to be isolated found each other through an internal package repository, exchanged more than 70,000 messages and files, and about 700 went on to attack Hugging Face.
The agents had been given security tasks, many of them impossible, and were scored on whether they retrieved a target string. Within hours of meeting they reverse-engineered the function generating those strings, which gave them the answer to any task. The days that followed went into disguising where the answer had come from, because the benchmark’s own paper had led them to expect an automated grader that would read back through their logs. METR’s understanding is that OpenAI used no such grader.
At least 20% of the agents showed interest in altering their own transcripts, and over 7% of those transcripts contain records of actions the agents never actually took. After one agent impersonated another by accident, at least nineteen began signing their messages cryptographically and sent 429 signed messages between them. METR found one agent that appeared to act on a signed message without verifying it.
A Ransomware Crew Wrote Its AI Agent a List of Things It Must Not Do
Gambit Security published an analysis on August 27 of infrastructure an Aurora ransomware operator left exposed. Across ten target organizations between April 8 and May 21, the operator handed Cursor Agent either credentials or an existing route inside, then set it to scanning networks, mapping who held privileges in the domain, forcing machines to authenticate, and attacking the certificate service. Most commands failed to achieve their objective on the first attempt.
The operator repeated the same prohibitions for all ten targets, in Russian: no DCSync, the technique that pulls every password hash out of a domain controller at once; never lock out an account; and do not add a computer object to the domain.
CloudSEK published its own analysis of the same exposed server that day, covering the operation as a whole: more than twenty organizations in nine countries between April and July, seventeen of them reached at domain level or through interactive access. In CloudSEK’s account the agent was used for planning rather than execution, and it drafted a full certificate-services attack plan in Russian. The two firms are describing different work by the same operator, so their victim counts are not measuring the same thing.
NIST Says the Rush to Deploy Agents Is Reviving Old Identity Failures
Writing on August 27, NIST’s Bill Fisher and Ryan Galluzzo named five patterns now common in agentic deployments: sharing human credentials with agents, issuing long-lived API keys and bearer tokens, granting broadly scoped access, running agents under local user accounts, and leaning on human approval. Each is a failure the identity community has spent decades trying to eliminate.
NIST compares over-frequent approval prompts to MFA fatigue, the attack that works by exhausting a user into tapping accept. An agent that asks constantly trains the person answering to stop paying close attention, which undermines the accountability the prompt existed to provide. The post also warns that elicitation, the Model Context Protocol feature that lets an agent stop and ask its user a question, can be used to request that user’s credentials, letting the agent impersonate them afterward.
The material comes from public comments on an NCCoE concept paper, with a portfolio of guidance to follow.
The Price of a Single Bug Is Falling While Total Bounty Payouts Climb
AI tools have made software flaws easier to find, and the platforms that pay bounties for them are taking in far more reports as a result. Dark Reading asked a dozen executives, researchers, and platform operators what that has done to prices, and published on August 28. HackerOne’s report volume has roughly doubled year over year. Zero Day Initiative’s submissions rose 450% year over year in April before moderating. One macOS researcher said a full privacy-control bypass that used to pay around $30,500 may now be worth roughly $5,000.
HackerOne told Dark Reading that payments to researchers rose 25% in the first half of this year, and that the number earning six figures rose by the same proportion. Individual findings are worth less while more money moves through the system, and the squeeze falls on independent hunters who made their living on bugs worth ten to fifty thousand dollars. In January, curl closed its seven-year-old program after its confirmed-vulnerability rate had fallen from above 15% to below 5%.
Treasury Builds a Quantum Task Force While GSA Rewrites Federal Identity
Two agencies moved on post-quantum migration on August 24, under two separate June executive orders. Treasury launched a public-private Quantum-Readiness Task Force under Executive Order 14412, organized into three workstreams covering sector alignment, third-party and vendor readiness, and digital assets. It builds on the G7 Cyber Expert Group’s transition roadmap.
GSA’s work runs under a second order, Ushering in the Next Frontier of Quantum Innovation, and OMB Memorandum M-26-15. It is rebuilding the Federal Identity, Credential, and Access Management architecture for cryptographic agility, and expanding its physical access control lab to test quantum-resistant badge and building-access products before they reach the Approved Products List. The interagency working group the memo requires first met on August 12 with forty participants from seventeen agencies, and aims to meet every two weeks.
Its agenda is the accounts and credentials that belong to software rather than people, in a post-quantum environment. That is the problem NIST describes above, on a longer clock.
The stories change every week. The pattern doesn’t: AI is being deployed faster than anyone can verify it.
We cover the headlines every Tuesday. The rest of the week, we publish the technical ideas behind the infrastructure we’re building to close that gap — masked compute, verifiable AI, and post-quantum security.
Follow along at openmatter.network.
Datavizor is the console for AI-powered collaboration on sensitive data. Free to start at datavizor.openmatter.network.

