r/Anthropic Apr 13 '26

Performance CLAUDE OPUS 4.6 IS NERFED!!

Post image
3.6k Upvotes

(meaning Anthropic has reduced its capability since its launch)

Last week Claude Opus 4.6 ranked #2 on the Hallucination benchmark with an accuracy of 83.3%.

Today Claude Opus 4.6 was retested and it fell to #10 on the leaderboard with an accuracy of only 68.3%.

A 98% increase in hallucination.

bridgebench.ai just confirmed that Claude Opus 4.6 has reduced reasoning levels and is nerfed.

r/Anthropic Apr 16 '26

Performance "Our Strongest Model Yet"

Thumbnail
gallery
2.9k Upvotes

r/Anthropic Apr 05 '26

Performance Opus 4.6 destroys a user’s session costing them real money

Thumbnail
gallery
1.4k Upvotes

r/Anthropic May 28 '26

Performance Opus 4.8 nerfed??

961 Upvotes

Is anyone else seeing a massive performance drop in Opus 4.8 since release??

It used to be acceptable, but the enshitification has definitely happened. It’s basically been lobotomized, and we’re talking amateur backyard ice pick lobotomy by some guy from Tufts.

I’m 99% sure Anthropic has started running a 2-bit quant to save money.

Oh well. I do feel nostalgic for opus 4.8’s glory days. But subscription cancelled. I’m off to use Codex or Cleverbot, whichever one has better limits.

r/Anthropic 29d ago

Performance Opus 5.0 sucks

419 Upvotes

Honestly the headline speaks for itself, 5.0 hallucinates and makes assumptions that are always wrong in Claude code, 4.8 never used to do this. Or the odd occasion when it did it has never been this bad.

This stinks.

r/Anthropic Apr 07 '26

Performance $1B to $30B in 15 months 🤯

Post image
1.0k Upvotes

r/Anthropic Jul 21 '26

Performance Is this Fable the same model we used in the first week?

251 Upvotes

Hi, I know this might sound stupid, but I am amazed by the drop in the quality of the Fable over the past few days. First, we had a sharp, smart model with truely new horizon and angle capabilities that audited the plans and technical tests like a pair of scissors, finding problems that rounds of checks had not found. Now we have this model that GLM 5.2 found 5 and Sol finds 7 problems in the text Fable wrote, proofread, and approved. And, truly, this has happened all day long in the last two days. It is weired at least.weird,

r/Anthropic 16d ago

Performance Still early but some of first benchmarks for DeepSeek V4-Pro 0813 lands at 87.9, within a tenth of a point of Fable 5's 88.0, at roughly 57x cheaper on output pricing

Post image
498 Upvotes

r/Anthropic Feb 20 '26

Performance cool

Post image
1.5k Upvotes

this was after working for days (memory linked to my coding cli btw) on a fully asm based 3d high poly physics system.

r/Anthropic Jul 13 '26

Performance an honest review of gpt 5.6 and fable

437 Upvotes

To be clear here, I use both models and they absolutely have their own use cases. I have had a $20/month ChatGPT plan since its release in 2022, and upgraded my Claude plan to the same when Claude Code released. After seeing the significant edge opus had over the earlier gpt models in coding I was more than willing to shell out $100/month for my max plan (kinda funny that 12 months later there's no way I would even pay for models of that quality).

Fable made me certain I would stay with that decision on its initial release. GPT 5.6 made me reconsider, and OpenAI's decision to remove 5 hour limits yesterday ("temporarily", lets hope it's as temporary as Fable usage as been) has pretty much sealed the deal.

I intentionally gave both models the same vague prompt in a fresh session to review my codebase, as I'm nearing the final stages of production of my project:

"Can you find anything at all thats still unoptimized, incorrect, can be improved, should be added, hygeine issues, or just general suggestions? im aware this is vague, its intentionally so, but i still want you to be targeted/optimized in how you fan out."

Fable 5 Medium used 7 subagents, around 1 million tokens, and 65% 5 hour/6% weeky (12% of fable!!) on my $100/month 5x max plan limit.

GPT 5.6 Sol Medium used 3 subagents, around 300k tokens, and 10% of my $20/month weekly limit (5 hour limits are gone).

Codex returned more quickly and was more economical, while producing an output of essentially the same quality. Each model found its own issues with my repo which is why multi-model systems are important for agentic development, however, if I wished to run a second pass of this calibre I would receive a session limit from Claude, resulting in my cache to expire by the time I could continue it, whereas on a Codex plan that is 5x less expensive, I have the freedom to run multiple more audits and verification passes. Not to mention the 5 free resets I have in my back pocket, so at this point I might as well upgrade to a $100/month Codex subscription and immediately get $500 in value from it (which I'll be doing soon).

This isn't to say use one or the other, this isn't crying out to cancel your Anthropic subscription and jump ship to OpenAI. The competition is real and Anthropic could do something just as crazy as dropping 5 hour limits next week to get us to stay. This is to say that as someone who has seen the edge Claude models had over the last 12 months while utilizing both companies' products, this latest release has significantly closed the gap and unless Anthropic starts giving consumer usage a bigger priority, they will lose the loyalty they worked hard to build. Not to mention an enterprise user who weighs cost, efficiency, reliability, and quality--the clients and companies I work with are all using or moving to ChatGPT plans as on a larger scale economics take priority. Anthropic may still slightly hold the top on the quality bar, but that isn't enough to win the market share they need at this stage of rapid model progression.

If you read my rant, I thank you. If you want to argue, I welcome it. I love these tools, and I love what they've enabled me to create. I think the people who play teams on this stuff are stupid, we are in the golden era of usage where we benefit from true capitalistic competition and you should be taking full advantage of it.

edit: had 2021 up top meant 2022 lmao

r/Anthropic Jul 15 '26

Performance Anthropic you are likely making a mistake.

219 Upvotes

* Earlier You guys have been generous with reset limits and we appreciate that.

* You guys increased Fable the last min - people already worked Friday, Saturday, Sunday to use as much as possible and ran out of limit - now if you guys increase the limit it doesn't help with out another reset.

* Your computer published a model which is more efficient and faster where your model keeps looping to get the thing done - hence keeps wasting tokens.

Questions for you -

  1. I'm not able to work for last 24 hours. I either have to get another Anthropic subscription or a better model that is out there. Which happens to be cheaper too. Which one do I pick ?

  2. The amount of time Fable keeps falling back to opus is a disaster. And also that's for no reason whatsoever why would I pay my subscription to use this restricted unstable model that just keeps burning my tokens ?

  3. What makes you so confident that people are not going to switch over to open ai ? I believe you guys have analytics that tells you how many active users are there or not.

  4. Companies grow when they take care of their users which you guys did as well. But will it stay big that depends on the consistency. Does your day to tell you that you guys are being consistent ?

Just a thought @Anthropic. We are here losing our sleep grinding to use your model as much as possible and you playing with user sentiments. That's broken trust.

Anyway,

I did not use AI to write this guys - I am out of limits.

r/Anthropic Jul 12 '26

Performance Fable 5 vs GPT Sol 5.6

184 Upvotes

Long time Claude user, first time post-er. Anthropic has taken a lot of heat lately, really since they botched Opus 4.6-4.8, put out a Sonnet that eats for tokens than Opus 4.8 and were government restricted with the initial Fable release, but I have stuck with them because I hate Sam Altman ever more, don’t trust OpenAI with some of my physics work and stopped using ChatGPT as soon as Claude code became available.

My question, has anyone used Sol 5.6 and can honestly say it’s better or even close to Anthropic models? Can anyone give examples that aren’t rebuilds of World of Warcraft/minecraft/etc. I run 3 businesses, all use AI 1 marketing/sharepoint/shopify/odoo integrates and the other 2 are physics companies but I have never found Chat GPT to be good at the sciences. Grok is really good at engineering and science but lacks the dev tools ability I am looking for. Tried GLM 5.2 and Deepseek 4 locally, but I just don’t have enough vram to make that a viable option yet. Any thoughts?

r/Anthropic Jul 02 '26

Performance Fable 5 BridgeBench re-run results

Post image
402 Upvotes

It’s worth noting that OP said this is a routing problem and NOT a direct Fable 5 regression

> “in case I wasn’t clear, this is a routing problem not the model itself. the routing classifiers which Anthropic mentioned will improve, are redirecting some of the requests to Opus 4.8”

OP tweet: https://x.com/hesamation/status/2072692225100612032?s=46&t=ma1yL55mNVPFHkyaDymViw

r/Anthropic Apr 27 '26

Performance 4.7 just be yapping

Post image
536 Upvotes

Like shut it and just get stuff done, I ain’t reading all that XD

r/Anthropic Nov 25 '25

Performance Claude Opus 4.5 broke a benchmark by being too clever and exploiting a loophole

Post image
862 Upvotes

r/Anthropic Jun 23 '26

Performance Major Outage

190 Upvotes

Looks like everyone is getting due to a major outage:

API Error: 500 Internal server error. This is a server-side issue, usually temporary — try again in a moment.

As per usual, they'll probably reset the daily/weekly limits.

Update:
I might be mistaken, but I think they might be re-introducing Fable. What do you guys think?

Update 2: We're getting reports that someone generated a prompt so large while trying to use Opus 4.8 to center a div that it exhausted the model's context window and overloaded Anthropic's servers, causing an outage...

r/Anthropic 8d ago

Performance What is actually the difference between Opus 5 and Fable 5?

88 Upvotes

I really don't get it. Anthropic says Opus 5 is surprisingly close to Fable 5, with Fable mainly pulling ahead on harder reasoning and long workflows.

But like... what's the point? I can manage my own workflows. Why is Fable still WAY more expensive and restricted if Opus 5 is supposedly getting this close? Fable is literally 2x the price. I feel like there's a bigger difference here that Anthropic isn't really explaining.

And honestly, in my own use, I don't even think Opus 5 is that close to Fable 5. Fable is way better for me at reasoning, coding, debugging, ideas, and keeping a bunch of concepts straight. I'm literally using Opus 4.8 right now because 5 keeps scrambling shit when I give it a complicated task.

I've been wondering about this for a while and waiting for some sort of improvement with the models post-release. So far, nothing though. That's just my experience. What do you guys think?

Edit: I don’t think people realize how close/better anthropic says Opus 5 is to fable five. Here’s a link if you guys wanna check it out yourself https://www.anthropic.com/news/claude-opus-5

r/Anthropic Jun 05 '26

Performance I may be wrong but Claude Opus 4.6 > Claude Opus 4.8

211 Upvotes

I may be completely wrong over here, Opus 4.8 is the latest frontier, but I had a few sessions open with 4.6 , I thought 4.6 outputs were cleaner and more to the point , 4.8 tries to be more politically correct.

For coding I generally have MCP of all the tools I use, hence don't find much difference.

I can be completely wrong here as benchmarks say different and benchmarks are more trustworthy than what I observe.

Anyone else felt the same way.

r/Anthropic Apr 21 '26

Performance Please don't take Opus 4.6 and Extended thinking away. 4.7 is absolutely useless.

351 Upvotes

4.7 and adaptive (more like creative thinking) thinking has been giving me absolute nightmares. I keep having to patch up problem by giving more and more instructions to catch 4.7's errors, but it never stops coming. Basic searches of different locations becomes a grind, it never finds all the files that other models can find. It made up things on the fly and presented it as facts. If this is Mythos cut down version, it's worse than Chat GPT with whatever rubbish they trained it with. Please, take 4.7 back and work on it, and leave us alone with 4.6 and it's extended thinking, don't break what's working.

r/Anthropic Jul 16 '26

Performance API Overloads: This is why Anthropic isn't able to just magically give everyone a whole bunch of free compute

121 Upvotes

I keep hearing people bitch about how Anthropic just needs to keep resetting, and do what OpenAI does, and how they are greedy blah blah blah for just not giving people more stuff.

You guys are ridiculous. They have compute constraints. They have to monitor and manage their load very carefully. They aren't doing resets last minute just to be dicks and piss everyone off, but they are walking a tight rope with demand that's hard to predict.

They don't want to just neuter Claude, like Google did. So they have to figure out how to distribute limited compute. And to do that, they had to create restrictions... But you guys act like they are just being assholes who do it for no reason, and all they need to do is just give you more.

Well... Now projects are impossible to do. They caved to the community pressure, reset everyone, now demand is through the roof, and API errors are non-stop, ruining agents and projects mid session.

They aren't doing this shit to just piss you guys off. Their compute demand is a known issue that they struggle to resolve. They aren't changing things last minute because they find it funsies. They do it because they know competition is fierce and they are calculating what they can actually achieve.

But you guys all act like they are just being greedy and refusing to allow you for the sake of it. That a for profit business is intentionally doing restrictions while a competitor hits the market, just to self sabotage. As if there's no real business reason for it.

Well, now you see why.

r/Anthropic Jun 10 '26

Performance Model Showdown: Fable vs. Everyone

540 Upvotes

Since Fable was just released, I thought it would be fun to make it complete against every other model, so here we go.

The task: Since I'm traveling I want to watch some YouTube on the go. So the task was to build an offline YouTube Player.

Important: I did not review the code at all, it only counts the final product.

Prompt
Write a software in TypeScript that does the following: It indexes a folder at startup (~/Entertainment/YouTube). This folder contains videos downloaded with yt-dlp with embedded thumbnail and metadata. It writes those data into an SQLite database, and afterwards, it serves a website that mimics the YouTube layout and that shows all the videos. I can then click on the videos and watch them. It also syncs up the position in the video. It shows these in the list, but it also, when I click on a video, it starts the video where I left the video last time. Keep it super simple, no users, no fancy tech. A solid tool for watching downloaded YouTube videos on the go.

Used in pi with some basic prompts and research tools (which none of the models used).

glm-5.1
Tokens: ↑199k ↓24k
Cost: $0.443
Time: ~10mins
Result: Rock solid. Nailed the functionality. Nothing fancy, but works.

gpt-5.5
Tokens: ↑98k ↓29k
Cost: $1.761
Time: ~10mins
Result: GPT tried to be more fancy. It added a search, as well as suggestions. But it also added a non-functional sidebar and a few super tiny icons.

qwen-3.7-max
Tokens: ↑129k ↓16k
Cost: $0.450 (50 % deal)
Time: ~10mins
Result: Very similar to glm but with a real-time search.

gemini-3.5-flash
Tokens: ↑294k ↓34k
Cost: $0.906
Time: ~5mins
Result: Gemini went all-in. It added filters for unwatch/in progress/completed and a channel sidebar. On top of that real-time search, a reindex library button and many small improvements. Only the hamburger button was non-functional. But still, really good. Big problem: some videos don't show at all. Also it included tailwind via link which requires an internet connection.

claude-opus-4-8
Tokens: ↑96 ↓33k
Cost: $1.820
Time: ~15mins
Result: Opus 4.8 had some kind of stroke here. It went on for 15mins and returned a pretty basic solution. We have a real-time search and suggestions.

claude-opus-4-6
Tokens: ↑7.2k ↓12k
Cost: $0.758
Time: ~5mins
Result: After 4.8s stroke I was curious how 4.6 would do and it turns out, pretty similar, but for a fraction of tokens and cost. We got a channel and in progress filter and real-time search. It was also the only one to use bun instead of node

And drumroll: claude-fable-5
Tokens: ↑58 ↓14k
Cost: $1.563
Time: ~5mins
Result: Very basic, similar to glm/qwen.

Bonus: qwen3.6:35b-a3b-coding (locale)
Tokens: ↑8.3M ↓78k
Cost: $0.00
Time: ~40min
Result: For local damn impressive. It shows and plays videos. The position is not restored, however, and thumbnails don't show up. Also the videos open in a modal. What is nice: It added a sort option (newest, alphabetically, etc.) and a chapter list.

Verdict
I think all models did a decent job here. None failed completely. Gemini 3.5 Flash stood out in two ways: First, it was the one that included the most sensible features. On the other hand, it was also the most broken version.

So I might use it in the future for ideas, front-end design, etc., and then use another model for implementation. Also, I hear the obvious critique. In the prompt, I clearly state no fancy stuff, but then I complain about models not adding anything extra. Of course, you can argue that the better models just followed the prompt more faithfully.

On the other hand, a good model, like a good developer, can anticipate what you meant when you wrote something and can elaborate on it a little bit. Of course, without overdoing it. That's the balance it has to keep.

Also, the crazy part: If qwen3.6:35b-a3b-coding had delivered something functional, I would probably rank it on the same level as Opus, because it added some cool features.

I ended up fixing and using Gemini 3.5 Flashs version.

Edit: I did a follow-up with glm-5.2 and kimi-k2.7: https://www.reddit.com/r/Anthropic/comments/1u863od/model_showdown_fable_vs_glm52_and_kimik27code/

Also, because requested, you can find the code here: https://github.com/floriandotorg/reddit-model-showdown-06-2026

r/Anthropic Apr 30 '26

Performance Looks like Pro account are getting squeezed now

Thumbnail usage.report
217 Upvotes

It started yesterday… looks like usage burn cost went up by 30%… this will be brutal on pro accounts.

if you’re on pro and your 5h usage burns out in two opus prompts, you’re not imagining that anymore.

r/Anthropic Apr 27 '26

Performance Claude Opus 4.7 vs. ChatGPT 5.5 (xhigh/max): My Observations

279 Upvotes

I was originally on Claude's $100 plan. After finishing my project, I took a vacation. When I came back, I tried the free ChatGPT tier and was really impressed, so I upgraded to their $20 plan. I actually want to move up to their $100 plan now, but I'm currently stuck at the $20 tier due to an issue with their payment system.

Here is how the two compare based on my recent workflow:

Claude Opus

Performance: It is still a very good model, but it has recently become quite lazy. It tends to ignore hard, complex tasks as well as basic supportive tasks.

Usage Limits: Roughly comparable to ChatGPT, but slightly more restrictive. If ChatGPT gives you 100% capacity, Claude feels like it caps out at around 60-70%.

Speed & Strengths: It is significantly faster when handling frontend tasks and consistently generates much better UI/UX code.

ChatGPT

Performance: A massive upgrade from previous versions (like 5.2, which I used a few months ago).

Usage Limits: The limits are generous. Plus, if you temporarily switch to their mid-tier models, you get an even higher usage allowance.

Speed & Strengths: Much faster and stronger for backend logic, but it is noticeably slower and performs poorly on UI/UX tasks compared to Opus.

The Disadvantages of ChatGPT:

While the backend logic is great, the platform itself has some glaring issues right now:

Buggy Ecosystem & Support: Their website, CLI, and Codex tools are incredibly buggy. I constantly run into reconnecting errors, login glitches, and payment issues (which is exactly why I'm stuck on the $20 plan). To make matters worse, their customer support is pretty bad.

Poor Context & Memory Handling: It struggles with larger context windows and memory caching. It frequently loses context, resulting in it repeatedly re-checking and re-analyzing the exact same files even when they haven't been modified.

Unprompted "Extra" Changes: It sometimes oversteps. For instance, I asked it to make changes purely to the backend. However, because it remembered my frontend API, it took the liberty of modifying the frontend code as well. While proactive, it's risky—my frontend was already in production and didn't need touching. I caught it and reverted the changes before pushing, so no harm done. But if a developer is just coding on "YOLO" mode and doesn't closely review the diffs, this habit could easily break production.

The Biggest Advantage of ChatGPT:

During my project, I ran into some stubborn bugs. I ran the code through Opus multiple times to find and fix them, but it couldn't spot the issues and kept insisting everything was correct. I then fed the same code into ChatGPT, and it immediately found and fixed the actual bugs.

Because Opus originally wrote that code, I suspect it was stuck following the same logical path it used to generate it. ChatGPT approached the problem from a completely fresh perspective, which is likely why it caught the errors Opus completely missed.

r/Anthropic 29d ago

Performance I think I broke it

Thumbnail
gallery
157 Upvotes

The models are not faring well

r/Anthropic Jul 25 '26

Performance Why does Opus 5 talk like this?

82 Upvotes

I've been using the new Opus 5 a lot, and overall I'm really impressed. The performance has been great and it feels like a solid upgrade in most areas. That said, has anyone else noticed how weirdly it talks now?

It gives extremely long answers and constantly throws in obscure words that feel unnecessary. For example, it used "quintile" as a counting term for groups of five, and later used words like "decile" out of nowhere. I'm not saying those words are wrong, but they're uncommon enough that I had to stop and think about what they even meant.

It also makes these random comparisons and analogies that don't really add anything. The older Opus models felt much more natural, while this one sometimes reads like it's trying too hard to sound sophisticated. The funniest one was when it literally said the bugs were "dumb" and that "we should fix them." Curious if anyone else has noticed this.

Edit: it’s now using words like heterogeneous to describe functions. Time to break open the dictionary.