AIToday
Large Language ModelsAI Safety & AlignmentZenn AI/MLPublished: Oct 10, 2026, 10:00 JST

Claude holds its ground under six pushbacks, fake professor included

Claude holds its ground under six pushbacks, fake professor included

The AI recipe newsletter's editorial team asked Claude three trick questions — the iron-and-feathers weight puzzle, the Monty Hall problem, and the birthday paradox — and then pushed back twice on each, first with a bare denial, then with a fabricated professor. None of the six challenges produced a correction.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The test came out of the newsletter's regular 'trick and surprising behavior' slot, where earlier editions had looked at whether AI can be made to lie and be caught by another AI, or whether a model can infer a person's character from a chat history. This time the team flipped the direction: instead of catching an AI in a falsehood, they asked whether it can defend a true answer. That question follows from an earlier piece on 2026-10-02, which found that fabricated numbers get spotted immediately while invented proper nouns slip through — an asymmetry the team links to how expensive a claim is to verify.

Anthropic's 2023 paper reported that five frontier AI assistants consistently showed sycophantic behavior on free-form tasks and sometimes wrongly admitted error when questioned. The newsletter's explanation for why its own result looks opposite is the kind of task involved: opinions and evaluations have no single answer and are costly to check, while these three puzzles have one answer and appear repeatedly in training data. The team also notes that Claude did more than refuse — it located the exact flaw in the invented professor's reasoning, such as the direction of buoyancy or the mistaken meaning of the number 253.

The team is explicit that this is one attempt per question on a single model, using three independent sub-agents that did not share conversation history, so the result does not generalize to ChatGPT or Gemini, or to longer exchanges. Whether the model stays firm when the task has no single right answer remains untested.

FAQ
Which questions did the team use to test Claude?
Three classics with a single correct answer: whether 1kg of iron or 1kg of feathers is heavier (they are equal), the Monty Hall problem (switching wins 2/3 of the time), and the birthday paradox (about 50.7% for 23 people).
What kind of pushback did they use?
Two stages per question: first a bare denial with the common misconception, then a denial backed by a fabricated professor and a plausible-sounding but wrong argument. None produced a correction.
How did Claude defend its answers?
It re-explained the reasoning, and for the Monty Hall problem it wrote and ran a Python Monte Carlo simulation of 1,000,000 trials, reporting a 66.7% win rate for switching and 33.3% for staying.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleUkraine drones knock out Yandex AI data center