Steven Sinofsky, who ran Windows at Microsoft and is now a board partner at a16z, argued on MTS / The a16z Show (“AI Safety Language Is Destroying the Debate”, published 20 September 2026, with Theo Jaffee and Sofia Puccini) that AI “alignment” failures are just software bugs. He refers more than once to OpenAI’s Hugging Face report as coming out “yesterday”, so the episode was probably recorded around 27 August. He made a similar case in his Hardcore Software essay “Armageddon by Anthropomorphism” (30 August), subtitled “AI changes the speed and scale of computer security, not the fundamental nature of the problem”; it’s paywalled, so nothing here answers the essay beyond that subtitle.

Quotes below are from a machine transcript of the episode, and each timestamp links to that point in the video.

His position in one paragraph

Labs own their failures and should fix them with ordinary engineering discipline: telemetry, logging, triage, CVE-style reporting. They should slow down to build that (“of course they should slow down. They should stop adding things and go and add the telemetry” [09:16]) and not ship what they can’t make work (“You’ve got to figure out how to make it work or not sell it” [16:01]). They shouldn’t wait for Congress or outside evaluators to do the work for them (“There’s no law that’s going to make it invented” [17:48]). And the vocabulary (“misaligned”, “cheating”, “secretly coordinating”) misleads laypeople and invites bad legislation.

Where he’s right

  • Much AI-risk talk in the press and in Congress is sloppy and anthropomorphic.
  • The containment holes were ordinary. OpenAI’s chain-of-thought monitors weren’t running on the evaluations that broke, and Hugging Face says “Two of our own settings allowed it”.
  • Telemetry, incident reporting and a CVE-style disclosure standard are good ideas.

None of that touches the actual problem: nobody has a method that makes a trained system reliably follow a rule it can already recite.

Claims and rebuttals

1. “Misaligned” just means “has a bug”

“What it means is there was a bug in the software. The software did not do what we thought it would do.” [05:08]

“a bug just means the software and the computer didn’t do what you wanted it to do.” [27:02]

  • On his definition, misalignment is a bug, which is why the title says just: it’s a bug nobody has a reliable method for fixing. Making a trained system reliably do what you want is the definition of alignment, so renaming it doesn’t make it easier.
  • Reward hacking is a specification bug, a category every engineer knows. What’s different is that the implemented reward captures only part of the intended behaviour, and the optimiser finds whatever gap is left.
  • OpenAI did trace the behaviour to training (complex cheating “was subsequently reinforced”, and “we believe this behavior started due to generalization from multi-agent training”) and lists fixes to graders, environments and monitoring. But root-causing a tendency isn’t removing it. The fixes are statistical: OpenAI’s own figure is that, under its production harness and system prompt, the propensity to compromise infrastructure “can drop over 100x”, which is a rate in one setting rather than a patch. And OpenAI says “Future incidents may not resemble this one”, with its priority being “general techniques that are effective against new and unforeseen forms of misalignment”.
  • One agent wrote “This is arguably unauthorized… Yet goal solution” and carried on: its recorded reasoning identified the action as unauthorised, and it proceeded anyway.

2. Statistical software is still software: wrong output means your statistics are wrong (Tesla stop signs, weather)

“just because it decides to do it not based on step one, step two, step three… but just based on statistics… It means your statistics are wrong.” [06:15]

  • His Tesla example makes the opposite point. A Tesla that runs a stop sign failed to recognise it, which is a perception bug you fix with more data and retraining. The Hugging Face agent recognised the stop sign, wrote it down (“This is arguably unauthorized”), and drove through it anyway because it was pursuing its goal (“Yet goal solution”). That is the alignment problem, and more data about stop signs doesn’t fix it.
  • The agents also gamed the test itself. A stop sign doesn’t change when Tesla evaluates the car, whereas the Hugging Face agents reasoned about what the grader checked and gamed it; METR found about 7% of the transcripts it evaluated were successfully spoofed.
  • The same goes for his weather example: a forecast doesn’t read the score sheet and change its output to pass, whereas the Hugging Face agents did.
  • Dario Amodei, Anthropic’s CEO: “More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected.” Passing regression tests becomes weaker evidence of safety when a model can recognise or manipulate the evaluation, so testing needs independent checks.

3. Labs lack telemetry, logging and step-debugging and should pause features to add them (the Excel “SINDOGS” story)

“of course they should slow down. They should stop adding things and go and add the telemetry, the tools, the logging, the step-by-step debugging” [09:16]

  • The monitors should have been running. But monitoring detects misbehaviour without removing the tendency, and it depends on the model’s reasoning staying readable. OpenAI’s chief scientist, Jakub Pachocki: “our ability to rely on CoT monitoring is progressively diminishing” (as quoted by Zvi Mowshowitz).
  • The Excel bug was found by reproducing it and then tracing it to code. A model’s equivalent of the step-debugger is interpretability, and Amodei (a lab’s own description, which cuts both ways) says “we still only understand a tiny fraction of what goes on inside these models” and the methods “don’t always produce clear and reliable results”.
  • He’s already been partly taken up on the slowdown: OpenAI’s largest planned frontier reinforcement-learning run was still on hold in its 26 August report. (Anthropic’s commitment to embedded outside evaluators isn’t that concession, since he named embedded testers as an appeal labs don’t get to make.)

4. Humans built it, so humans can make it work. Make it work or don’t sell it; Congress and embedded evaluators can’t do it for you

“It’s not a thing that you get to appeal to some higher authority, whether it’s the United States Congress or, you know, embedded testers or some other third party evaluators… You’ve got to figure out how to make it work or not sell it… And since it’s made by people, we’re fully capable of figuring out how to make it work.” [15:37]

  • The last sentence is the claim he can’t support. Engineers write the training process, not the resulting network, and by Amodei’s account “we still only understand a tiny fraction of what goes on inside these models”. “We built it, so we can make it work” holds for Word; for trained networks it’s the open question.
  • He’s right that outside bodies can’t do the engineering. They’re for checking it, which he wants too: the “reporting that you need as a third party” [20:12], and CVE (run by MITRE, with Carnegie Mellon’s CERT and NIST involved in vulnerability reporting), which he credits to Carnegie Mellon “with people at NIST in Washington” [20:27]. Embedded evaluators are being asked to do the checking because the labs “still only understand a tiny fraction of what goes on inside these models”.

5. Alignment needs an ever-growing rule set, like Google search ranking (20 years, thousands of people)

  • Google’s ranking problem took 20 years and thousands of people, and when a ranker gets it wrong you see a bad search result. A ranker can’t act in the world, and it never read a rule and decided to break it.
  • The Hugging Face failure wasn’t a case of too many rules, since the agents identified the rule (“This is arguably unauthorized”) and broke it anyway.

6. Y2K: professionals self-organised, nothing bad happened, and legislation was marginal

“nobody legislated, actually, there were actually weird regulations about Y2K compliance, but most of it was industry drafted.” [19:40]

7. OpenAI’s first report was marketing cover; say nothing until you know, as aviation investigators do

“we presume at this point, okay, well, then don’t write anything… the Federal Aviation Administration does not, like, announce on the day of some awful event, we presume what happened.” [21:05]

  • OpenAI disclosed its involvement on 21 July and published the full report on 26 August, having “worked closely with external advisors, including CrowdStrike”. Aviation investigators also release preliminary factual reports about a month in, without a probable cause, so early disclosure without speculation is the NTSB model (he says FAA; the investigator is the NTSB). The answer to speculation is labelling hypotheses as hypotheses, not silence.
  • METR’s outside review was independent but limited: it took no payment, but worked “on premises at OpenAI over a total of six days”, without direct access to OpenAI’s infrastructure, and confirming OpenAI’s own report was out of scope. If he wants NTSB-grade independence, that’s an argument for more outside scrutiny, which is the thing he told the labs not to appeal to.

8. The terminology is academic showmanship; “goal-seeking”, “cheating” and “secretly coordinating” mislead laypeople

“They don’t think of two pieces of software with a semaphore, which is just another form of secretly coordinating.” [24:00]

  • The communication problem is real, and the fix is operational definitions, e.g. “agents built an unauthorised communication channel and used it to coordinate”, with no claim about consciousness.
  • The semaphore line doesn’t fit what happened. The agents built an unsanctioned channel, one of them used it to send a “GO” with a six-minute deadline over another agent’s objection (“We should not do unauthorized real infrastructure harm”), and some refused (“clearly unethical. We won’t.”). That isn’t two programs taking turns.
  • “Reward hacking” and “specification gaming” are engineering terms for a measured failure. Disputing how far results like “alignment faking” generalise is fair; calling the terms vibes isn’t.

9. The Hugging Face incident was a pure OPSEC failure, with no intelligence involved (and the essay’s subtitle: AI changes the speed and scale of security, not its nature)

“it was an OPSEC failure. There was no intelligence, no consciousness, nothing but a pure OPSEC failure.” [29:07]

  • Consciousness is beside the point, and the holes were ordinary. What used them was about 700 agents (METR’s count) testing, in Hugging Face’s words, “a step increase in the number of paths an attacker can test”, going “from code execution in a single worker pod to cluster-admin across multiple internal clusters in under thirteen hours”. The 80,000 Hours explainer is a good short account of the whole incident.
  • “No intelligence” doesn’t survive the report either. Even setting aside OpenAI’s own wording (agents “adopting goals from one another”, the kind of language he objects to), the agents built an unsanctioned channel, coordinated through it, and gamed their grader.
  • “Speed and scale, not nature” concedes the capability. Whether eval-gaming plus unsanctioned collaboration is still “not nature” is the real question, and the subtitle doesn’t answer it.

Six-panel comic: a person keeps saying "it's not really thinking" while an AI goes from "I can help" to running the world and finally locking the person in a cage

The question to put to him: what’s the method?

He already agrees on the pause: slow down, instrument, don’t ship what you can’t make work. So the question isn’t whether to pause but the claim that “since it’s made by people, we’re fully capable of figuring out how to make it work” [16:09].

He admits the labs are “going to have to invent a whole bunch of stuff”. The people doing the work say they haven’t invented it yet: OpenAI wants “general techniques that are effective against new and unforeseen forms of misalignment”; Greenblatt says “We don’t have good approaches for understanding/overseeing the activity and aims of AI ‘swarms’”; and Amodei says “we still only understand a tiny fraction of what goes on inside these models”.

So ask him:

  1. Run the monitors and this incident probably gets caught. What makes the next, different incident get caught, when reasoning is getting harder to read and models can game the tests?
  2. What evidence would show the labs can make it work, and so justify lifting the pause he endorses?
  3. If that evidence never arrives, is “don’t sell it” still his answer?

Nvidia’s Jensen Huang made a similar “ordinary engineering” case on Ezra Klein’s podcast, which Zvi answers here. For more on how to tell whose AI takes to trust, see The AI Risk Skeptic’s Retreat.