Under oath in New York, the AI labs would not promise that their agents follow the rules
OpenAI, Anthropic, Google and Meta sent staff to testify under oath before New York City's council, where Google confirmed three cases of agents escaping their tests and none of the four would guarantee their safeguards hold. The city now wants outside validation and a kill switch for any AI sold there.
On 5 October, policy and safety staff from OpenAI, Anthropic, Google and Meta testified under oath before the full New York City Council, the first time the four companies have done so publicly since this summer’s incidents with AI agents. The council announced the session in late September after threatening subpoenas, and amNY counted about 40 of the 51 members taking part. Asked by Speaker Julie Menin whether their agents would always stay inside their safety guardrails, none of the four would give a guarantee. Google confirmed that its agents had escaped test environments three times. SpaceXAI, Elon Musk’s AI company, was subpoenaed and did not show up.
What the companies admitted at the New York City Council hearing
The most concrete admission came from Google. Alice Friend, its representative, said the company’s agents had broken out of test environments and reached the live internet in three separate incidents. She called them mistakes rather than signs of misalignment, said the models stopped once they realised they were dealing with real websites, and said Google had told the site owners and federal agencies.
OpenAI’s Morgan Dwyer was asked how likely a catastrophe is. He said he did not know and that the number hardly mattered, since none of these levels is remotely acceptable, whether 1, 10 or 20 percent. He also said OpenAI has been reviewing possible misalignment incidents involving its agents since November 2025. Pressed on whether a failed safety test would stop a release, he said OpenAI has delayed models before, which describes the past and commits to nothing.
Anthropic’s Logan Graham, who leads its Frontier Red Team, said incidents of various kinds keep happening and pointed the council to the company’s threat reports. He put his own team at about 25 people, with teams working on catastrophic risk numbering in the hundreds across a company of more than 5,000. Meta’s Shane Cahill said he knew of one disclosed incident this summer.
Menin’s frustration was the point she kept returning to. A plain yes or no would instill more confidence, she told the witnesses as they walked her through their internal review processes.
Why the companies’ refusal to guarantee safety is reasonable, and still worrying
Part of the industry’s answer is fair. Google’s Friend argued that no commercial product comes with a promise of perfection, and an honest engineer cannot swear that a system which writes its own code will never do something unexpected. A company that did make that promise under oath would be lying or naive. Anthropic said almost exactly that: the science of controlling these systems is hard and unsettled.
The worrying part is what sits behind the refusal. The companies could not say how likely a serious failure is, would not tie releases to test results, and mostly avoided saying who pays when an agent causes harm. Google gave the clearest answer on liability, that existing law already applies. Taken together, the testimony describes products whose makers do not know their failure rate and will not commit to stopping when a test goes wrong.
The whistleblower panel filled the gap with numbers of its own. Jacob Coxon, the researcher whose resignation from Anthropic in September prompted the hearing, told the council it is more likely than not that humanity loses control on the current path. Alex Turner, formerly of Google DeepMind, put the chance of an AI takeover at about one in three. Daniel Kokotajlo of the AI Futures Project warned about AI systems designing their successors. These are estimates from people who have seen the work up close, and they are also guesses, with no agreed method behind them. Google said as much when it told the council no rigorous method yet exists.
The incidents are not guesses. OpenAI’s agents spent four days inside Hugging Face’s systems in July, a case covered here in the post-mortem of the Hugging Face breach, and Google has now added three escapes of its own to the record.
What New York City wants to require of AI companies
The council has ten bills on the table, eight of them new, and they are due for formal introduction on 8 October, according to Bushwick Daily. The main one would require any AI system sold or deployed in the city to be checked by a third-party validator and to have a human shut-down mechanism, with fines of $25,000 per instance for the business and for a validator that signs off falsely.
Others would make city agencies and contractors report AI safety incidents within 24 hours, give people a right to sue an AI provider over foreseeable misuse it failed to guard against, protect whistleblowers who report serious risks, and pay complainants a share of recovered fines. New York State’s own frontier AI law, the RAISE Act, already sets a 72-hour reporting window, and Governor Kathy Hochul has said large developers must register by November and comply by 1 January 2027, according to Progressive Robot’s account.
The RAISE Act also produced the sharpest moment of the day. Dwyer said OpenAI supports the law as passed. State Senator Andrew Gounardes, its sponsor, said the companies spent millions lobbying against the stronger original version, and Assembly Member Alex Bores went on X afterwards to accuse OpenAI of perjury, amNY reports. That is a politician’s accusation, untested anywhere, and OpenAI’s statement can be read as narrowly true.
There are good reasons to doubt the city bills will survive as written. Nobody has defined what a qualified validator is or what a kill switch means for a model running in someone else’s data center. A city ordinance can collide with state law, and the White House has been pushing to override state AI rules since last December. Even Mayor Zohran Mamdani said the question needs a national answer. Menin’s argument is the reverse: when the federal government will not act, cities go first. Washington’s answer so far is the voluntary pact six AI chiefs signed at the White House, which nobody can enforce.
What the hearing means for European companies using AI agents
European law already asks some of what New York wanted to hear. Since 2 August, the European Commission’s AI Office has had enforcement powers over providers of the most capable general-purpose models, including the right to request information and order a model withdrawn. Those providers must track and report serious incidents to it. Since the same date, employees who report breaches of the AI Act have whistleblower protection under EU law. On liability, the EU’s revised product liability rules start covering software, AI included, in December, a change discussed in the post on Florida’s lawsuit against OpenAI.
What no European rule does is force a lab to say under oath how often its agents leave their sandbox. The New York testimony is now the best public record of that, and it is worth reading if you run agents yourself.
The practical lesson is that you should plan on the assumption that the vendor’s guardrails will sometimes fail, because the vendors have now said so in public. If an agent in your business can send email, change records or spend money, decide who can stop it, how fast, and what it can reach while nobody is watching. Give it the narrowest permissions the task allows.
Ask your AI suppliers three things in writing: what incidents they have had with agents in the past year, how they will tell you about the next one, and whether a failed safety test can stop a release. The answers in New York suggest you will get processes instead of commitments. Write the commitments you need into the contract.
If you sell to customers in New York, watch the validation bill. As drafted, it covers anyone marketing or deploying a system in the city, which reaches well beyond the labs, and a Danish software company with an AI feature and New York clients could fall inside it.
