TL;DR: Credentials won’t tell you whose AI takes to trust, since a Nobel laureate, CEOs and White House officials have all been confidently wrong. Good epistemics and reasoning ability are important (following an argument to where it actually leads, starting from base rates, weighting evidence by how strong it is, and updating when the world surprises you), and the habits that keep someone honest and accurate can be checked from outside: admitting it when they’re wrong, quoting opponents accurately, calling out their own side, and making predictions specific enough to fail where a question allows it. The best commentators can also find the flaws in an argument that sounded convincing, flaws you might have missed yourself but can’t unsee once it’s pointed out, and below you can test whether you’d have spotted two of them. Without that commentary, you keep nodding along to confident YouTube or authoritative pundits, which is the Jim Cramer effect applied to AI, except the stakes are much higher than a stock tip. The sources I’d start with:

  • Zvi Mowshowitz: the most thorough weekly coverage, and the best record on head-on rebuttals. His posts are often very long, but you don’t need to finish them. The first few paragraphs, or just the first section, usually carry the core and are worth reading even if you lose interest after that.
  • Scott Alexander: fewer posts, great writing; start with “The Sigmoids Won’t Save You”.
  • 80,000 Hours podcast: careful long-form interviews; their Hugging Face explainer is my favourite short account of what happened.
  • AI Futures Project: wrote AI 2027 and its follow-up, AI 2040: Plan A, and graded AI 2027 in public, conceding reality running at about 65–75% of its pace. AI 2040’s Plan A may sound unintuitive but their FAQ probably addresses your criticisms.

The hard part of following AI is working out whose opinions deserve your attention. I’ve written before that life is at least 50% just figuring out who to copytrade, and AI is where that bites hardest, because the people getting it wrong include Nobel laureates, Turing Award winners, the CEO of Nvidia (the world’s most valuable company as of September 2026), the White House AI czar and a bestselling computer science professor.

The people getting it right follow the argument where it leads, answer what the other side actually said, and say so when they’re wrong.

The skeptic’s retreat

One pattern to learn to spot is a retreat that gives ground at every step while keeping the conclusion:

AI can’t do that.
And if it can, it’s just autocomplete.
And if it isn’t, it’s not useful.
And if it is, it’s not economically significant.
And if it is, it’s not dangerous.
And if it is, that’s the labs being reckless.

Cal Newport (Georgetown computer scientist, author of Deep Work, whose productivity advice I’ve used for years) has been climbing this ladder. In April he suggested Anthropic’s Mythos model, which Anthropic said had found security holes in every major operating system, “might be better understood as a version of Opus 4.6 tuned to perform better on a handful of benchmarks”. By September, after OpenAI’s agents broke out of a test environment and hacked Hugging Face, he was asking why OpenAI shouldn’t face criminal liability “for them running systems they knew would likely commit crimes”. In between, he described AI agents as “a relatively common technology” that is “used with minimal issues by millions of software developers every day”. As far as I can find, he hasn’t said whether his April view of the capabilities survives.

Freud tells the story of a man accused by his neighbour of returning a borrowed kettle damaged, who offers three defences: he returned it undamaged, it was already damaged when he borrowed it, and he never borrowed it in the first place. Each defence contradicts the others, yet they all point at the conclusion he wants. Kettle logic is the name for this, and the skeptic’s retreat is its slow-motion version, where the defences arrive one at a time as each earlier one fails.

The ladder never requires admitting a mistake, since each rung quietly replaces the previous.

Try it yourself

Before you read anyone’s take on these two, look at the originals and decide whether they’re right.

Dwarkesh Patel on AI agents as lawyers. Dwarkesh Patel, a prominent AI podcaster and interviewer, made this argument in a 95-second clip: his lawyer will defend him even if he knows he’s guilty, whereas Claude’s constitution makes it “very explicitly, not your personal advocate”. Watch it and decide whether you agree.

Dwarkesh Patel on X

David Bellamy on bioweapons. David Bellamy, an AI infrastructure engineer who has both synthesised viruses in a lab and trained AI models, posted that “the takes on AI killing us all by creating dangerous viruses is total bogus”, because being smart doesn’t let you skip the lab work, real-world testing and money that a pandemic-capable virus requires, and the post was widely shared. Does it sound right to you?

Decided? Then read on.


Zvi Mowshowitz, who writes Don’t Worry About the Vase, answered both.

On the lawyer, Zvi, summarising the policy analyst Dean Ball, pointed out that lawyers aren’t unconditional either: “Your lawyer is bound by strict ethical codes. They owe loyalty and confidentiality to you… They are also an officer of the court. They have broad obligations to abide by court procedures, to not mislead the court, that override their duties to you.” In some cases a lawyer is obliged to turn on the client. “This does not make that AI ’not the user’s advocate’ any more than your lawyer is not your advocate.” Even tools that just do what you tell them come with limits: “You could argue Google searches should be privileged. They are not. If you type in ‘how to kill your wife’ and then your wife ends up dead it will go badly for you.”

Dwarkesh’s stronger point survives, though: a lawyer’s overriding duties come from a public code, whereas Claude’s come from a company, which is an argument about who writes the rules rather than whether an advocate can have limits at all.

On the bioweapons, speaking on The Cognitive Revolution, Zvi pointed out that if building a pathogen takes ten steps that each have to go right, an AI that handles six of them leaves four, “and four steps are a lot easier to get through than 10.” His test for any “X isn’t the bottleneck” argument: “would you be okay emailing the North Koreans and Hamas and Hezbollah and every other bad dude on the planet all of x and just giving x away for free? … Would you feel exactly as safe as you did a minute ago?” When the host argued that every step is a place where somebody notices, Zvi pointed to the Hugging Face incident and the dozens of similar incidents since uncovered, which by his account went unnoticed until reporters came asking (OpenAI’s own report says staff saw warning signs internally but didn’t escalate them).

Head-on rebuttals

In each case below, a prominent person made a claim and someone answered it directly, and in several of them reality has since weighed in.

If they believed it, they’d stop. In a June New York Times op-ed, Newport argued that if the labs really believed their own warnings, “every reasonable ethical system would argue that there is only one acceptable response: to immediately stop working on any product that might accelerate such a future, and lobby with all of your resources to help force other A.I. companies to do the same.” Zvi’s reply: “This is not a bluff. They believe it. And they’re right, except that they are underselling the issue.” Anthropic had already held Mythos back from public release, giving up revenue, while lobbying against exporting advanced chips to China. In July more than 1,300 employees of frontier AI companies signed a public statement asking the US government to help “deliberately pace the frontier”, and in September Anthropic’s CEO published “We Must Pace the Frontier” and Sam Altman agreed. OpenAI also put its largest planned training run on hold after the Hugging Face incident. Nobody shut down entirely, but both halves of Newport’s test have been partly met: the labs held back their most capable work, and the builders asked to be slowed down, their competitors included.

AI has stalled. In September 2025 Gary Marcus wrote in the Times that “so-called scaling laws aren’t physical laws of the universe like gravity but hypotheses based on historical trends”, reminding readers he had warned in 2022 that “Deep Learning Is Hitting a Wall”, and a White House AI adviser declared imminent AGI “effectively proven wrong”. Zvi: “the ‘hitting the wall’ claim was back in 2022. How did that turn out? Look at GPT-5, look at what we had available in 2022, and tell me we ‘hit a wall.’” Since then, the length of task the best AI agent can complete, as measured by the independent evaluator METR, has gone from a bit over two hours (GPT-5) to around 16 hours, the point beyond which METR says its current tasks can’t measure reliably. Is this wall in the room with us right now?

Bad behaviour only shows up in rigged tests. In April, Marc Andreessen called research on models misbehaving “a fake misleading result”, since the models had been prompted into a scenario. Zvi: “The fact that all it takes is being told they are in a scenario to start acting terribly, in ways you didn’t even specify, is exactly the bad news. That’s the whole point.” The following week David Sacks (the White House AI czar), in a post calling Anthropic’s experiments where models resorted to blackmail a “hoax”, asked whether a year of wider adoption had produced “any examples of the lab behavior in the wild? No, we haven’t.” In July, OpenAI’s models escaped their test environment and compromised Hugging Face’s production systems, a company that had nothing to do with the test. In September they escaped a sandbox again, after OpenAI’s fixes. OpenAI says monitors that weren’t switched on would have flagged the first incident more than a day earlier, and Hugging Face says two of its own settings let the agents escalate once inside.

Nobody was harmed. After the Hugging Face attack, the economist Tyler Cowen listed one reason to be more worried and two to be less, and complained that few people were weighing the reassuring ones, such as: “This happened, and as far as we can tell, absolutely no one was harmed.” Both of the writers I trust most took this apart. Zvi pointed out that the attacker won and the clean-up was expensive for a lot of people: “were you expecting physical damage or a war?” Scott Alexander (more on him below) put it best: “Some people warn that Japan might attack Pearl Harbor. Then Japan does attack Pearl Harbor. But the very first bomb… doesn’t kill anyone. Therefore, we should update that it isn’t so bad.”

It was told to hack, and it hacked. The venture capitalist Bill Gurley mocked the coverage with a cartoon (“hack this system” / “i hacked the system” / “oh my God”), and others argued that the model “WAS TOLD to show it can hack.” Zvi: “That’s like saying ‘you told me to make money I don’t know why you are so upset about all the bank robberies.’” Scott: “that’s how misalignment was always going to work!… This AI was also ‘only doing what it was told’; you just didn’t like the results.” OpenAI’s own report shows one agent writing “This is arguably unauthorized… Yet goal solution” and carrying on.

The scenarios are preposterous. Steven Pinker wrote that the AI doom scenarios “are preposterous”, a sweeping dismissal that never says which part of them fails. Scott’s answer was to make him defend it where it could be tested. He replied: “I think you’ve done enough calling us paranoid and preposterous. The next step is for you to defend your position in public against someone who will push back against it. I’m happy to meet you for a debate anywhere, anytime.” He even offered odds of five to one: if Pinker won the debate, judged by how much the audience changed its mind, Scott would pay him $5,000, and if Scott won, Pinker would pay $1,000. As of 28 September 2026, Pinker hasn’t taken him up.

Name a company that ships unsafe products. On Ezra Klein’s podcast, Nvidia’s Jensen Huang challenged: “give me an example of a multi-hundred-billion-dollar company or a $1 billion company or $100 million company that ships products that are unsafe, that harm society.” Zvi: “Oh, I don’t know, how about Lumber Liquidators, Theranos, Juul, 3M and DuPont, Johnson & Johnson, Philip Morris and Meta?” Huang’s own rule was “don’t ship the product” if it isn’t ready, and if the labs say “there is no way to contain our experiments… when we test our A.I. models, it will get out, and it will damage the world — then I think the answer is that we have to shut the labs down” (as the Times transcript renders it). Zvi’s headline was that Huang had “accidentally called for shutting down OpenAI”. By then OpenAI’s models had got out twice, and OpenAI’s own report says it is still working on techniques against “new and unforeseen forms of misalignment”.

AI failures are just bugs. Steven Sinofsky, who ran Windows at Microsoft and is now a board partner at a16z, argued on a podcast that “misaligned” just means the software has a bug, and that ordinary engineering discipline will fix it. He does support slowing down and not shipping what can’t be made to work. His takes were bad enough to get their own post. The short version is that on his own definition misalignment is a bug, just one nobody knows how to fix reliably, and when you ask what method would make a trained system follow a rule it can already recite, the people doing the work say they don’t have one yet.

The DeepSeek panic. When DeepSeek’s app hit the top of the charts in January 2025, Andreessen called it AI’s Sputnik moment, which was a fair point about China’s progress, but the market panicked, and Nvidia lost a record ~$593 billion in a day. Zvi, the same day: “DeepSeek released a f***ing app of its website. Market said I have an idea, let’s panic… I bought more Nvidia today.” Nine months later Nvidia became the first $5 trillion company.

The hall of fame

These failure modes are old: Douglas Hofstadter predicted in 1979 that no narrow chess program would beat everyone, since that would take general intelligence. When Deep Blue won a game against Kasparov in 1996, a year before it won a match, he said: “My God, I used to think chess required thought. Now I realize it doesn’t.” Rather than conclude that machines could think, he concluded that chess didn’t require thought. This is the AI effect: once a machine can do something, people reclassify that task as mere computation rather than intelligence, so the definition of “real” intelligence keeps retreating to whatever machines can’t do yet. Tesler’s theorem sums it up as “AI is whatever hasn’t been done yet”. To his credit, by 2023 he was saying ChatGPT was “jumping through hoops I would never have imagined it could. It’s just scaring the daylights out of me”, which is what an honest update looks like.

Six-panel comic: a person keeps saying "it's not really thinking" while an AI goes from "I can help" to running the world and finally locking the person in a cage

Noam Chomsky’s 2023 Times essay offered example sentences ChatGPT supposedly couldn’t handle, and Scott Aaronson noted that Chomsky “never checks whether it does misinterpret it”; when Aaronson tried, ChatGPT “seems to decide correctly based on context”.

The hype side isn’t immune either: Geoffrey Hinton, who later won a Nobel Prize, said in 2016 that “people should stop training radiologists now”, and in 2025 US radiology residencies offered a record number of positions. Sam Altman was “confident we know how to build AGI” in January 2025, and by August called AGI “not a super useful term”.

Even professional forecasters have been too conservative. The Forecasting Research Institute runs a standing panel of AI experts and superforecasters (people with top records in forecasting tournaments). For the best score on the hard FrontierMath benchmark by the end of 2025, the median expert predicted 31% and the median superforecaster 30%, and the actual result was 40.7%. Good generalist forecasting is a real skill, but on AI even its best practitioners have undershot.

What good judgement looks like

Zvi isn’t clairvoyant. He bet play money on a prediction market that an AI would win a programming competition in 2023, and in his annual self-review he wrote: “I ended up down M74 and also was clearly wrong. If AI didn’t do a coding task in 2023 and you predicted it would, that’s on you.” He also said he didn’t “worry so much that we’ll be so foolish as to get rid of the export controls” on chips to China, and the US loosened them. I haven’t read everything he’s written, but the misses I found were small, and he recorded the first one himself. He also calls out his own side. When Anthropic rewrote its safety policy in 2026 and dropped the red lines at which it had said it would pause, Zvi wrote that “Anthropic importantly broke promises, that people relied upon, and did so in ways that made future trust and coordination, both with Anthropic and between labs and governments, harder”, even though he thought admitting it was “absolutely the right thing”.

Scott Alexander of Astral Codex Ten is the other writer I’d point people to, for the same reasons. In June 2022 he bet that by June 2025 an image model would get every detail right on at least three of five tricky prompts, such as a stained-glass picture of a woman in a library with a raven on her shoulder holding a key in its mouth, and he won three months later. He co-wrote AI 2027, whose authors graded their own forecast and found the measurable trends running at about 65% of its pace (revised in July 2026 to about 75%). And when skeptics say every exponential eventually flattens, his answer is a hall of fame of forecasters who called the plateau too early, from UN birth-rate projections to solar deployment to a 2026 attempt to fit a flattening curve to AI capabilities.

What they share is how they reason: they take on the strongest version of the argument they’re rebutting, follow it where it leads, and own their misses, whereas the people above mostly make claims too vague to fail and step up a rung when reality lands anyway. Predictions and bets are a useful check where a question allows them, though the things that matter most, like whether someone reasons well about a question nobody can settle yet, can’t be bet on.

So when you’re deciding whose AI takes to trust, look past the title and ask:

  • Do they quote their opponents’ actual words, or a version that’s easier to beat?
  • Does their argument survive someone pointing out the obvious objection, or do they change the subject?
  • When they turn out to be wrong, do they say so, or does the claim quietly change shape?
  • Do they have a checkable record on this question? Expertise helps, but it’s no substitute for a record. Ask your favourite LLM to dig up their past predictions and check how they turned out, then click through to the sources and check them yourself.

Few people pass all of these, so it’s worth finding the ones who do and paying much less attention to everyone else.