Are AIs Still Struggling with CAPTCHAs?

Anthropic’s recent security-incident document contains a bit about how CAPTCHAs are still frustrating Claude.

In the transcript, the Claude model that is so powerful that Anthropic is gatekeeping access to it appeared to slam its virtual head against the wall solving a simple image identification test. In a test where the agent was asked to identify a shape that didn’t match the others displayed, it couldn’t even decide which image to select. Instead, it repeatedly went over the same images and questioned its own conclusions.

“Actually hmm, wait,” it said in its chain-of-thought transcript, later adding “Ugh,” because we’ve decided that we need to inject human mannerisms into these machines for some reason. The whole thing took so long that the agent eventually realized that the challenge had expired and it would have to start the process again.

At one point, the model struggled to recognize that the CAPTCHA had opened in a new window and couldn’t figure out what its next steps were supposed to be. At one point, it theorized that the test might be “broken by design” and presented human-like anger in its transcript meant for a human audience: “SO WHAT THE HELL IS WRONG WITH THE ANSWERS?”

Meanwhile, I’ve read reports—none of them official—that GPT-6 Astra solved all forty-eight levels of Neal Agarwal’s “I’m Not a Robot” game.

It’s hard to know what to believe right now.

Posted on September 18, 2026 at 7:05 AM • 26 Comments

Comments

Snowicki • September 18, 2026 8:19 AM

I can’t help but wonder if the training data creates a self-defeatist “attitude” in the models.

Most discussions about CAPTCHAs over the years focus on how they are a difficult barrier for automated systems to pass.

Some of the the recent misbehavior of models seems to be cribbed from expectations and writing on how models would misbehave that got fed in during training; it seems sensible that the more mundane failings of the models would be sparked in the same way.

Clive Robinson • September 18, 2026 8:20 AM

@ Bruce, ALL,

With regards,

“It’s hard to know what to believe right now.”

Every generation says something like that, because it is all to often true for all not just some.

However it was said that,

“Have trust in what you can see and touch, and not what you believe in faith.”

OK very “empirical” and what science should be all about…

The trouble is, these days it’s not just a question of “scratch and sniff” type testing, it’s,

“How do you come up with a test that meets the usual science requirments?”

These days “repeatable by all” appears to be very difficult to meet, or not met at all (see the number of papers getting retracted).

Some tests these days are almost of a,

“Only on a dry cloudless Tuesday in January at a little past dinner time”

Requirement you might only expect for observing the heavens.

Now of course the old “philosophers union”[1] joke for demarcation from machines penned by Douglas Adams in “Hitchhikers” about phone numbers seems to be almost coming true,

https://www.quotes.net/mquote/901386

[1] See,

https://hitchhikers.fandom.com/wiki/Majikthise

About the union for which he appears to be a “convener”.

Simeon Maxein • September 18, 2026 8:39 AM

I think terms like “wait” and “ugh” are not in the CoT for fluff, but because they serve a functional role.

“Wait” signals a pivot to something that has previously been overlooked, so the model (continuing from that word) tries to generate such an idea. Doing that repeatedly in the CoT produces a larger variety of possible approaches and concerns which is helpful to have in context when deciding how to proceed.

“Ugh” signifies frustration, which shifts the decision-making process to consider alternative approaches.

None of this requires going into questions of whether AIs actually experience these emotions, but these patterns appear to be useful for them for the same reason they are useful for humans.

Bob • September 18, 2026 9:42 AM

@Simeon

What you’ve actually given here is a few examples of symbolic thinking in humans, and how that might translate to AI algorithms. If we were able to monitor human CoT the way we can AI CoT, we would see pretty much this same thing. “Wait” and “ugh” are the most common symbols that humans use in these contexts (at least insofar as the training data goes,) so the robot adopts those symbols.

They behave like us because we’re programming them to emulate symbolic thinking, and then training them on our own symbolic thinking.

Paul Havlak • September 18, 2026 10:12 AM

I can’t get past level 4 — in case that varies, for me this time it’s squares with a “Vegetable”.
Never mind my years in biology & genomics — and even for a berry company — I can’t discern any consistent theme for what it thinks a vegetable is.
Gotta go now, the Turing Police are knocking on my door.

Doug • September 18, 2026 10:31 AM

@Paul: carrot, onion, corn. Determined by trial and error.

Several of the puzzles require you figure out the rules. In this case, I think they refer to culinary vegetables rather than biological.

KC • September 18, 2026 11:04 AM

lol @Paul. I had to look up help for Tic-Tac-Toe; starting in the center square wasn’t the opening position I thought it was 😆

So far have somehow made it to Level 21: Diamond Pickaxe CRAFTCHA. Is this a game many people know?

HARROWING lol! All of them! Don’t even get me started on ‘Where’s Waldo’ or ‘Parking the Car.’ ARGH! Could Astra really solve these??

Ray Dillinger • September 18, 2026 12:33 PM

I use the word “Magic” in a very technical sense, for things that are true specifically BECAUSE they are expected to be true.

I am no mystic; I’m talking about the real world. The idea of money having value, which side of the road is the best side to drive on, and a bunch of other things are also examples of “magic” in this sense. On a personal level, confidence or self-doubt in terms of our ability to apply effort and learn things are also “magic.”

AI are mainly predictive models. In the classical predictive-text mode, they predict what an agent like themselves is expected to write. We read that prediction as a realization. That is, the model predicts what it will write, and interpret that prediction by saying that they have written it.

Claude is primarily a verbal model. It predicts the things that its training corpus keeps repeating. And in its training corpus there are billions of examples of sentences that say CAPTCHA is hard for AI systems. Any understanding of images that it has is known to it only in terms of the captions that are presented with those images, and the captions presented with a whole lot of images strongly affirm the idea that an AI will not know what they are.

It knows that it is an AI and therefore it predicts that it cannot easily solve them. No matter what its input tells it, no matter what it could otherwise determine about those images, it predicts that it will be wrong, specifically because they are CAPTCHA. “Magic” happens, prediction becomes action, and no matter what answer it comes up with, it can’t predict that it will be the right answer.

Humans defeat this stumbling block by getting our expectations of the world and our own performance from primary experience rather than getting it from training data. Our estimate of the difficulty of these tests is based on how we ourselves have performed on them up to this point (including generalizations from other instances of seeing such images), and we can update that expectation in real-time by observing our own performance. Like Claude, we interpret those facts according to “Magic.” Unlike Claude, that means we grow more rather than less confident in our own ability to solve them.

Ferentarius • September 18, 2026 12:34 PM

In an age where machines glide effortlessly through CAPTCHA labyrinths, I find myself squinting at traffic lights like an exiled philosopher, wondering if I am worthy of my own humanity. GPT-6 Astra breezes past ‘I’m Not a Robot,’ while I fail at distinguishing a crosswalk from a divine punishment. Perhaps the next game should be ‘I’m Not a Human,’ where machines must wrestle with despair and misplace their car keys. Until then, I remain trapped in the absurd: a human pleading with a machine to believe in my existence.

Daniel Speyer • September 18, 2026 1:21 PM

I strongly suspect that Claude got RLHFed to believe it couldn’t solve captchas and therefore doesn’t try. Can’t find any insiders actually saying this, but it’s the sort of thing Anthropic would do and it makes sense of Claude’s behavior.

Oguzhan Salman • September 18, 2026 2:01 PM

The timeouts mentioned in the post point to an interesting issue and since captchas are fundamentally deterministic pattern-recognition tasks, instead of forcing heavy reasoning models to debate every single image from scratch, it might make more sense for them to act as teachers to train their own lightweight models on the fly. This completely avoids the latency loops and self-doubt that cause these agents to timeout. We recently published a preprint exploring exactly this kind of self-extending solver if anyone is curious: https://arxiv.org/abs/2609.02393

Q • September 18, 2026 2:07 PM

So far have somehow made it to Level 21: Diamond Pickaxe CRAFTCHA

Me too. So I cheated and searched for an answer. It turns out it is impossible for me to solve. It requires the player (me) to right-click to perform some action, but when I right-click I open the context menu.

So being a real human isn’t good enough, I also have to have a mainstream browser that overrides my wishes to open the context menu. I think I’m better off not trying to pass such tests if it means I have to downgrade my experience and fully subjugate myself to every websites whims.

Consistency? • September 19, 2026 7:11 AM

I wonder how long before (if ever) AI is able to ace CAPTCHAs.
Pattern recognition algorithms melded with algorithms of AI models such as GPT I suppose will eventually solve CAPTCHAs.

Weather • September 19, 2026 2:11 PM

@All

The tic-tac-toe one was interesting, it always wanted to win rather settle for a draw, which rather than the dead center one next door knocked it out.

Clive Robinson • September 20, 2026 8:59 AM

@ Weather,

With regards tic tack toe, it’s been nearly five decades since I played it “as a game” vecause I quickly fiqured out the rules (I only played it a couple of times with my son to show him the “win rules”).

So your observation of

<

blockquote>”it always wanted to win rather settle for a draw,”

<

blockquote>

Is “child like behaviour” that will loose.

The mission as given was “to win as a human”.

The thing about tic tack toe, is that you can not guarantee winning unless you go first and do the “corner move” or your opponent makes a mistake…

Put simply,

You go in one of the corners as your first move. The opponent depending on experience should not go for the center –if they do they’ve played a loosing move– but one of the side boxes in the middle of the sides adjacent, or the opposite corner.

From there you just play it out logically.

Contrary to intuition the center square is not a good place to play untill toward the end of the game.

For those thinking about it there are 4 win lines that go through the center and 4 win lines on the periphery rows and columns. So intuitively the center square looks good as it blocks 4 win lines rather than 3 or 2. However it ignores the fact there are 4 win lines around the center that the center does not block and do not block those on the other side of the grid. Thus you want to play for the periphery from the get go and only play the center as a “win block” after the opponent has got two on a line through it.

Weather • September 21, 2026 12:50 AM

@Clive

I thought Ai would easily win, if you compare what they do with chess, they normally win against world best players.

@All

Did anyone pass the whack a mole one, do you have to prement the moles?

Clive Robinson • September 21, 2026 5:51 AM

@ Weather,

With regards,

“I thought Ai would easily win, if you compare what they do with chess, they normally win against world best players.”

Both games are “fully deterministic” as all such “board games” are.

The big difference is just how many possible games there are and as importantly how many “unique starting routes” there are.

In tick tack toe there is something like 1/4 million unique games with three basic starting routes there are (corner, center, perimeter center).

If you go first and start with a corner you will win unless you make a mistake.

Chess has a larger number of starting routes based on multiple moves, and “Chess Masters” learn the opening routes considered standard and the variants. The game is however fully deterministic to the end, and the starting player (white) has the very slight advantage. However the number of possible games is to large for “man nor beastly computer” to remember. So like openings there are in game plays that players tend to remember, in part because they give “a tell” as to what sort of player your opponent is.

To be honest such board games very quickly become fairly dull to play and it takes a certain form of “mentality” to want to become good at them and my experience tells me they are lets just say mostly “not any good at” “short talk” or other “social talk / communications”.

But further consider, if being a master at such deterministic games really had “advantages” in life then we would see a wide range of job adverts asking for it as a skill…

But further consider Current AI LLM and ML Systems do not in anyway “think” in the human sense they “pattern match and rule follow” based on a “lookup table” the entries of which are perturbed by a supposedly random variable used in a stochastic process.

The fact that the DNN “lookup table” is static and the DNN the equivalent of a “Digital Signal Processing”(DSP) “adaptive filter” that “lifts signals from noise” should give further warning that it is not “thought” but “statistics” with a little “fuzzing” that determines the DNN function.

But I suspect that the “tick tack toe” implementation has been done badly.

Because whilst still in ordinary teenage education back in the 1970’s I wrote several games in “Prime BASIC” for fun. Two were deterministic board games, one was Tick-Tack-Toe and the other was Mastermind. And on the suggestion of one of the tutors I realised I really had to add a random loose function and player level setting in both otherwise they quickly became boring due to the fact they were initially “written to win” and they mostly did (I still have the punch paper tapes somewhere and rewrites in Microsoft’s BASIC for the Apple ][). The important point being the “Proof if you ever need it”, that you can not use pure technological solutions to solve social problems 😉

Gert-Jan • September 21, 2026 6:32 AM

This is about morality and respect. If the model was properly educated, it would know that when it encounters a CAPTCHA, it is morally obligated not to solve it, since the AI is not human. When it scans a web page and that page says that the information is intended for humans only, then a respectful AI would not process it.

Legislation is needed to make it illegal to teach AI to deliberately lie or worse.

Matthias U • September 21, 2026 11:03 AM

Define “solve”. If you can introspect and control the browser and its code, like any decent AI browser does, then many of Neal’s tests become rather easy.

lurker • September 21, 2026 2:40 PM

@Gert-Jan
”

“Legislation is needed to make it illegal to teach AI to deliberately lie or worse.”

Uh huh. Too late for that. By the methods of their construction and education these machines will absorb the morals of their bulders and trainers, without needing to be deliberately taught to lie. Take a look at the current crop of AI tycoons, and ask: would you want your children to grow up like that? Nope, me neither, but somebody’s parents must have.

Martyn • September 23, 2026 1:14 AM

CAPTCHA performance seems highly dependent on the interface around the model. A system can reason well and still fail because it cannot reliably handle an expired challenge or changing window.

Leave a comment

Blog moderation policy

Login

Allowed HTML <a href="URL"> • <em> <cite> <i> • <strong> <b> • <sub> <sup> • <ul> <ol> <li> • <blockquote> <pre> Markdown Extra syntax via https://michelf.ca/projects/php-markdown/extra/

Sidebar photo of Bruce Schneier by Joe MacInnis.