People expect a new AI model almost every month now. GPT-5.4, then 5.5, then 5.6. Opus 4.6, 4.7, 4.8. Most launches barely get a reaction any more, even from me, and I’m the sort of person who reads the release notes.
But under every AI decision a marketing team makes, there are questions nobody has a straight answer to yet. Do we commit to a tool now, or wait for the next model? Is it expensive? Do we pick one model or keep switching? Is this stuff actually getting smarter, or are there just more launches? Are things really moving that fast?
A month ago I’d have told you progress had flattened. So I went digging in the data to see how fast things are really moving. Is AI improvement speeding up? Have we hit a wall? And what does it mean for marketing teams trying to build with it today?
I sat down with Claude* and went through every model ChatGPT and Claude have put in front of us since ChatGPT launched in November 2022. All 57 of them. Dated, checked against the release notes, and scored against two independent measures: Epoch AI’s Capabilities Index and METR’s task-length results .
*Claude did most of it. A lot. Definitely. My PC crashed halfway through, which felt about right.
Why does it feel like AI has stalled?
In July I had breakfast in the garden while Opus finished the first version of a client tool. Then I asked Fable to check it. It ran for 45 minutes. 76 separate agents. 7 million tokens. It found the bugs, wrote a seven-step plan to fix them, and when I said I had to cycle in, it moved itself to the cloud, carried on, and wished me a nice ride in. I put the Stones on and got on the bike.
That should feel like a revolution. It didn’t, quite. Because a month earlier a big AI-assisted data job had come back with over half its links wrong, and I found myself typing “just thought we were beyond this kind of rookie AI error”. (I did not spell it that well.)
That’s the whole story in two mornings. Every launch is a bit better. None of them is a “wow” on its own. In the last 12 months the two apps launched 15 new top models between them, up from 6 two years before, and each new record added about 2.3 points on average.
(Quick one on the points, because they matter for the rest of this. Epoch’s index squashes more than 50 benchmarks into one number. It allows for how hard each test is, so it doesn’t stop working when the easy ones max out, and it isn’t a percentage, so there’s no 100 to hit. It’s pinned so Claude 3.5 Sonnet, from June 2024, scores 130 and GPT-5, from August 2025, scores 150. GPT-4 is about 126. Today’s best are about 166. So 20 points is roughly Claude 3.5 Sonnet to GPT-5, which felt like a big deal at the time.)
So we get the progress in slices. Nobody notices a slice.

Each dot is the strongest model an app launched on that day. The solid lines are each app’s best so far. GPT-2 and GPT-3 are in grey for context, as they were never chat apps.
What were the real leaps?
Two. You can see both on the chart.
GPT-4, in March 2023. It added 16 points in one go and stayed the best model you could use for about a year, longer than anything since. We all remember it. It felt even bigger because we’d waited so long for it.
Reasoning models, in late 2024. o1-preview, then o1, then o3 added 17 points in seven months. It came in three releases, so it never felt like one moment. But it was the same size as GPT-4, and it did more than jump once. It roughly doubled the pace of everything after it.
In between, not much happened. For 14 months after GPT-4 the best score moved about 3 points.
I’ll own one bad call here. In January 2025 I wrote on LinkedIn that OpenAI’s Operator was “a million miles ahead of the Claude version”. Nineteen months later I basically live in Claude Code . I was wrong about who’d lead. I was right about the line I tacked on the end: “Right now this is the worst the tech will ever be!”

Each block is one model that beat everything before it, sized by how much it added.
So is everything since then just small improvements?
Each step, yes. The total, no.
Since o3 in April 2025 the best score has gained 20 points. That’s more than GPT-4’s leap. It works out at around 14 points a year, five times the pace before reasoning models arrived.
Epoch AI, who build the index, found the same thing another way . The jump from GPT-4 to GPT-5 was about as big as the jump from GPT-3 to GPT-4. It just came spread across a lot of models in between.
Last December I posted a chart on LinkedIn and wrote “LOOK AT THE INCLINE OF THAT GRAPH!” I stand by the capitals.
Where do tools and agents show up?
Not in the capability score. In how long a task AI can finish by itself.
METR, an independent research group, gives models real software and research tasks with tools (a terminal, code, files) and no human help. It measures the length of task, in the time a skilled person would take, that the model gets right at least half the time.
In March 2023, GPT-4 managed about 5 minutes. In February 2026, Claude Opus 4.6 managed about 12 hours. That’s 130 times longer in three years, doubling roughly every four months.

This chart is on a log scale, so each gridline is about four times the one below. A straight line here means steady doubling.
I feel this one every day. In April I tracked a day: about 70 hours of manual work squeezed into 9. In August my frontend terminal started messaging my backend terminal BY ITSELF because it needed something before it could carry on. And in July my wife and I picked a holiday using a comparison page Opus built for us. We used it for 25 minutes. Then we never looked at it again. Nobody will ever maintain that code. That’s what software is now.
What’s the catch?
Useful lags impressive. GenAI is brilliant at the first 50 to 80% of a brief. The last bit is where it gets hard.
- 12 hours is at 50% success. Ask for 8 out of 10 and Opus 4.6 manages about 70 minutes, according to METR’s own data.
- These are tidy software tasks. METR found that about half of AI code fixes that pass the tests wouldn’t be accepted into a real codebase as they are.
- Real paid work is further behind. On the Remote Labor Index , which tests agents on actual freelance jobs, the best agent now completes about 21% to the client’s standard. That’s up from 2.5% a year ago, which is fast. It’s still a fifth.
It isn’t just us. The Content Marketing Institute’s 2026 B2B research found almost every team now uses AI, but only 39% say it’s actually improving performance .
We see it with clients too. One comms team built their own prototype in Perplexity in an afternoon. It got them about 20% of the way. Then it hit data pipelines, logins and governance, and stopped. The afternoon demo was real. So was the other 80%.
Dwarkesh Patel put it better than I can: “Models keep getting more impressive at the rate the short timelines people predict, but more useful at the rate the long timelines people predict.”
Are we waiting for the next breakthrough?
This is where the smartest people in the field disagree.
- Ilya Sutskever, who co-founded OpenAI, says the “age of scaling” is over and we’re back in an age of research. In July he said his company, SSI, now has “research that is worthy of scaling up” . In other words, he thinks he’s found the new idea.
- Demis Hassabis at Google DeepMind thinks we may need “one or two more breakthroughs” , and names continual learning and memory as the gaps.
- Dario Amodei at Anthropic says the “hitting a wall” story comes round every few months, while underneath there’s “a smooth, unyielding increase in AI’s cognitive capabilities” .
My read? Nothing in the data has slowed yet. The missing piece is AI that learns on the job and gets reliable at messy work. Whether that needs a new idea or just more of the same, I genuinely don’t know. Neither, it turns out, does anyone else.
And honestly, even if it stopped improving today, it’s already transformational. I’m dyslexic. I’ve been hiding typos for 35 years. LLMs are the first piece of software in my life that has done what I meant, not what I typed. (You should see my WhatsApps.)
What does this mean for marketing teams?
- Don’t wait for the next model. It’ll be about 2 points better. If a tool, app or build isn’t working for your team today, the next release won’t fix it. How you use it will.
- Plan for bigger jobs, not cleverer answers. The thing that’s doubling is how much work AI can take off your plate in one go. Whole pieces of work, like the research, the first draft of a campaign or a full set of assets, with a person checking the result.
- Keep a human on the last mile. At 8 out of 10 reliability, tasks top out at around an hour. That’s still a lot of hours back. But the checking is the job now.
- Don’t marry one model. New top models ship about every eight weeks, and the lead keeps swapping. It’s one of the reasons we build our own tools, like Compass, rather than renting someone else’s. More control, more customisation, more functionality.
- Don’t let cost be the reason to wait. The model is rarely the expensive bit, and it keeps getting cheaper. Anthropic says Opus 5.5 costs about 40% less to run than Opus 5, which launched two months before it. The real cost is setting it up properly and checking what comes out.
Claude Opus 5.5 came out on 22 September. On paper it’s a couple of points better than what came before. I told the team it feels like talking to a normal person again.
That’s the thing about slices. Each one looks small. Then you look up.
How I worked this out
I listed every model you could pick in the ChatGPT or Claude web chat, dated from the day it arrived there, and checked each date against OpenAI’s and Anthropic’s release notes. For each release day I kept the strongest model. Capability scores come from the Epoch Capabilities Index (Epoch AI, CC BY 4.0). Task lengths come from METR . Where Epoch hadn’t scored a model yet, like Claude Opus 5.5, I estimated it from its raw benchmark results using Epoch’s own method, which lands within about a point of Epoch’s figures when tested on the models it has scored. (Yes, Claude did most of this bit too.)
Chris Wright , Founder, Fifty Five and Five
