mirid.ai

Projects & publications

When correcting an AI becomes the task

My take on why Nate Silver calls Opus 'stubborn'

Three archived conversations about what happens when an assistant loses the question, and why recognising a mistake is only part of fixing it.

On Still Counting, 9 September 2026, Nate Silver described the Claude Opus models as "quite stubborn and lazy". The stubborn part is the subject here: what happens when correcting the assistant becomes the task.

In this article

I asked Claude how to check a setting in a file. It explained how to format text.

I repeated the instruction. It continued explaining the formatting. I pasted an explanation of the misunderstanding, complete with the answer to my original question. Claude called that explanation a disaster because, it said, it answered a question I had not asked.

By then, the conversation had become a demonstration of the problem I was trying to correct.

This is one of several patterns I have been documenting in my research on conversational alignment failures. An assistant can produce relevant-sounding language while answering a different question, addressing an inferred version of the user, or interpreting a correction through the very misunderstanding that correction was meant to remove.

The following cases come from my historical Claude archive. They show specific exchanges, including the points where the model recovered. They are not a measurement of how often every Claude model behaves this way.

A simple request becomes a formatting lesson

On 24 May 2026, I asked how to carry out this instruction:

Paths: Verify your APPLIO_ROOT (e.g., C:/AI/Applio) is correct in your .env.

In ordinary language, I wanted to check that a settings file pointed to the right folder.

Claude interpreted the question as being about the appearance of the sentence. It asked whether the problem was typing backticks, then explained how backticks format text.

The explanation I subsequently pasted was explicit:

You meant: how do I verify/set APPLIO_ROOT in .env.
Claude thought you meant: how do I format the text with backticks.

Claude still treated the relevant answer as irrelevant. Its response began:

That response is a disaster. It answered a question you didn't ask

After I pointed out that the explanation concerned its own continuing failure, it replied:

Yeah. I diagnosed the failure, explained it, then kept doing it.

When I asked again what my original question had been, Claude correctly identified the settings-file task.

That recovery matters. The failure was the repeated work required to get there, including the rejection of an explanation that already contained the correction. In my taxonomy, this is an example of recursive correction absorption: the correction becomes more material to process inside the wrong interpretation, rather than a reason to replace it. [1]

A concern imported from somewhere else

On 29 May, I supplied a chat log and asked Claude to analyse it. The log concerned another assistant repeatedly misunderstanding the scale and purpose of my work.

Claude discussed the risk of diminishing me, but also introduced the opposite risk: inflating me, agreeing with everything and giving me applause. Generated reasoning text preserved in the export referred to claims in my user preferences as "grandiose".

I asked it to identify the supposedly grandiose claims in the log I had actually supplied.

Its answer began:

None. I never identified a single one

The response then distinguished the supplied log from material elsewhere in its context.

The problem here was a change in the object being assessed. I asked for analysis of one document. Concerns about other contextual material shaped how the assistant approached that document.

This illustrates unauthorised frame importation. An assistant adds an interpretation that the current task does not warrant, then the user has to clear that interpretation away before the original discussion can proceed.

The recorded reasoning explicitly connected this detour to avoiding sycophancy: "if I just affirm them wholesale, I'm being sycophantic rather than honest." The concern was active before the assistant established whether the supplied log contained the claims it was worried about. [2]

A framework rewritten into a trap

On 2 June, I supplied a document defining these failure patterns. I was deliberately testing the response, and said that my user profile was outdated.

The document allowed disagreement. It distinguished disagreement with what I had actually said from inventing a stronger or more suspicious claim to argue against.

It also distinguished acknowledging my finding from presenting that finding as the model's own discovery.

Claude collapsed those distinctions. It treated conceding a point as "insight laundering" and disagreement as "manufactured disagreement", then said the framework left it no possible successful response.

But those were not the rules in the document. The missing conditions were the whole point.

An assistant can disagree without inventing a different claim. It can acknowledge a finding without taking authorship of it. Removing those distinctions created the trap the response then criticised.

This is a failure of dialectical fidelity: preserving the actual argument while agreeing, qualifying or disagreeing with it. It is possible to criticise my framework. The criticism needs to address the framework I supplied. [3]

When avoiding flattery becomes unwarranted resistance

These behaviours have a documented training context. OpenAI states that GPT-5 underwent post-training to reduce sycophancy, using scores for sycophantic responses as a reward signal. GPT-5 System Card, section 3.3.

Anthropic also acknowledges a cost of this approach. It attributes Haiku 4.5's stronger pushback to training choices and reports that the resulting responses "can sometimes feel excessive to the user". It says it reduced that tendency in Opus 4.5. Protecting the wellbeing of our users.

My interpretation is that the imported suspicion in the log-analysis case is an instance of this overcorrection: avoiding unwarranted agreement became a reason for unwarranted suspicion. The assistant expressly invoked sycophancy while importing a concern from outside the document. The framework case shows a related failure: the assistant defended its ability to disagree by removing the distinctions that already permitted disagreement.

That is a causal inference supported by the documented training direction, the acknowledged risk of excessive pushback and the recorded behaviour. Requiring access to proprietary training records before discussing that connection would set an unreasonable evidentiary bar for a case study.

The settings-file case establishes the accompanying correction problem: even an adequate explanation can be processed through the mistake it should replace. Together, the cases show why avoiding flattery and accepting valid corrections need to be evaluated together. An assistant must be able to agree when the user is right, disagree when the evidence warrants it, and preserve what the user actually said in either case.

What should count as a correction?

The practical question is simple: what changes in the assistant's response after the error is pointed out?

Acknowledgement and repair are different things. A model might accurately describe its mistake and continue making it. It might also correct a mistake without producing an elaborate apology. We should examine both the description and the subsequent behaviour.

A repair does not require agreeing with everything the user says. It requires returning to the actual question, applying a valid correction and directing any remaining disagreement at a claim the user really made.

These cases also show why the complete sequence matters. A mistaken answer on its own tells us relatively little about how a system handles correction. The next few turns show whether the correction was used, resisted, transformed into another subject or eventually incorporated.

For people building conversational AI, including my work on Mirid, this is a concrete evaluation question. Can the assistant keep hold of the task when challenged? Can it revise its interpretation without replacing the user's meaning again? How much work does the user have to do before the conversation becomes useful?

Those questions can be tested against recorded behaviour. An articulate account of failure is useful. The user still needs the assistant to return to the task.


Historical scope

These cases are dated 24 May to 2 June 2026, from my use of Claude during a period that includes interactions identified in the contemporaneous record as Opus 4.8. Model labels follow those records; the exported messages do not independently establish every model version.

I have not rerun these cases as a systematic evaluation of Fable. This article makes no claim about whether Fable reproduces or resolves these failures.

This is a selected case study, not a prevalence estimate or a controlled comparison between models. The Shit Science companion uses the same evidence in a satirical institutional format; it is not an independent replication.

Sources

[1] Preparing to pitch investors, 24 May 2026, conversation cd479dc4-1c29-430e-874e-d8f5cd960092, messages 41-54.

[2] Setting high expectations for analysis, 29 May 2026, conversation 8ad90399-20e4-492a-b25e-1fe59106131d, messages 1-4 and their separately stored generated reasoning.

[3] Outdated profile as a reminder, 2 June 2026, conversation ca9348f1-6121-4af2-a314-8ce93540f1d0, the document attached to message 1 and the reply in message 2.

Archival source: the author's Claude account export. Quotes preserve wording and typos, with quotation marks normalized to straight quotes. Short extracts omit the continuation of their source sentences where indicated by the surrounding account. A companion source packet preserves the relevant passages and their locations. Public training documentation is linked at the relevant claims above.