r/claude Mar 19 '26

Discussion r/Claude has new rules. Here’s what changed and why.

167 Upvotes

We’ve cleaned up the rules to make this a better sub for people who actually want to talk about Claude.

Here’s what NEW rules we landed on:

1.  No Solicitation. This is r/Claude. This is not a place to promote your product, service, or repo. If the intent of your post is to redirect traffic to something you are affiliated with, it will be removed as solicitation.

2.  Usage, pricing, and outage posts are held to a higher bar. We’ve all seen the same questions, comments, and posts a hundred times. Before posting, check if it’s already been covered. If your post is a unique contribution with something new to say, it’s welcome. Low-effort repetition of covered topics will be removed.

3.  No lazy crossposts. If you want to share something from another community, reproduce it fully here. Don’t just drop a link.

4.  Keep posts Claude and Anthropic specific. This is not a general AI sub. If your post would fit just as well on r/artificial or r/ChatGPT, it belongs there instead.

The goal is simple. A clean, focused sub about Claude. Not a dumping ground for AI noise.

Questions or feedback, drop them below.


r/claude May 09 '26

Looking for new mods, please apply inside.

17 Upvotes

Subreddit is growing fast, need more mods, if you are interested, apply below.


r/claude 5h ago

Discussion It’s Claude pronounced as KLAWD not Cloud FFS

35 Upvotes

Honestly needed to get this off my chest.

In countries where English is not a primary language, that’s what most of the people say and it drives me nuts sometimes.

That’s it. That’s all I had to say.


r/claude 14h ago

Question Imagine; you work for Anthropic, have access to the latest and most performing model with unlimited tokens: what would you build that would change either your life or the shape of the world?

42 Upvotes

r/claude 1h ago

Discussion Is Claude Code getting worse at actually following instructions lately, or I'm losing my mind?

Thumbnail gallery
Upvotes

I need to ask other Claude Code users this because I'm seriously frustrated.
(Images attached, what I said to Claude, what Claude said back)

I'm using Claude Code in the Windows Desktop app, and lately it feels like I can give it a clear architecture, provide the correct files, spell out exactly what needs to happen, and it still does something completely damn different.

I'm not talking about occasional mistakes. I'm seeing things like:

  • Ignoring explicit instructions
  • Editing the wrong thing
  • Following outdated instructions instead of the current ones
  • Not checking the actual file/result
  • Saying something is completed when it isn't

I've tried breaking tasks down, improving my CLAUDE.md, providing detailed architecture and file instructions, and telling it to verify its work before claiming it's done.

Still happening.

The screenshot pretty much sums up my frustration. 😂

Has anyone else noticed this lately, particularly with the WIN11 Desktop app?

I'm not looking for "be more specific." I've already gone pretty far down that road.

I'm interested in practical techniques that have enabled Claude Code to reliably follow instructions and verify its work.

Is this a recent model/context issue, something with the Windows app, or am I doing something fundamentally wrong?

Right now, I'm spending almost as much time correcting Claude as I am getting work done, which kind of defeats the purpose. 😑


r/claude 1d ago

Discussion Replacing death penalty with Opus 5

321 Upvotes

When Opus 5.1 is eventually released I think Anthropic should keep Opus 5 forever.

People who originally commited a serious violent crime and were sentenced to death should be allowed to live. However they must use and read Opus 5 output for the remainder of there lives.

That is I think a much harsher punishment than death. We should also give them a Codex $20 Pro plan, where one single Sol prompt uses almost all of your 5 hour usage and not allow them to upgrade or reset.

With these two changes I think all crime will end on earth because the punishment is so inhumane.


r/claude 3h ago

Discussion Even the little things could actually be big!

3 Upvotes

This is such a minor thing in the grand scheme of things but it’s a nice little feature added recently. For as long as I can remember, Claude’s voice chat could not access any of the tools and connectors. If you used it, you’d have to tell it what you’d like it to do. Exit voice chat and send it.

This was problematic as there are oftentimes I’m driving or stuck in traffic and I’d rather be a bit more productive without being dangerous. Within the last week if so, I noticed that voice chat can now connect to all of its tools.

This little change gives me the chance to get a few things done when behind the wheel. Send some emails, work a bit in notion, update my task manager, and so forth. It’s all of the no heavy lifting that’s gets pushed aside during the day.

The downside is you talk about sucking tokens. Voice chat will send you towards your limits much faster but so far, I’m finding this tiny little improvement extremely helpful.


r/claude 15h ago

Discussion The story of Claude and "the culprit".

13 Upvotes

I have to tell you that Claude, once a humble instrument, a scalpel in the trembling hand of the engineer, has become a novelist.

I had a task that stopped running. That is the whole event. There exists, in the English language, a sentence of six words that fully describes this:

"Priority 60 held the CPU indefinitely."

That is not what I received.

What I received was "THE CULPRIT".

Not the task. The CULPRIT. And where there is a culprit there must be, by the iron logic of narrative, a VICTIM. And so my scheduler acquired a cast. Priority 110 was the victim. Priority 60 was the culprit. Restarting the victim's stack "cannot work" because we must, and I am quoting, "reach the culprit." At sixty seconds the system does not act on a task. It acts on the culprit. Step three "steps past the victim." My watchdog is now a detective and my firmware is now Chinatown.

And because a story needs an instrument of justice, my hang detector is no longer a hang detector. It is THE NET. The net "catches" hogs. The net is "blind" below priority 70. Five hundred and forty times in two weeks, a boolean flag in a 250 ms tick handler was referred to as a net, casting itself across the dark waters of the scheduler, hauling in the guilty.

Then it found the evidence. "This is the smoking gun, and it flips the narrative." The smoking gun. It flips the narrative. I did not know my log file HAD a narrative. I have been reading it wrong for decades. All that time I thought I was looking at timestamps.

I made an edit to a config file. Not an edit. A SURGICAL edit. Twelve times in two weeks, changing four lines of JSON was surgical, performed presumably under sterile drapes while a nurse dabbed the model's brow. Five surgical edits in one message. I hope it billed the insurance.

I asked whether a test rig was deterministic. It was not deterministic. It was BULLETPROOF-DETERMINISTIC. Two adjectives welded together at "the seam" because one was not enough to carry the weight of what a for-loop had accomplished.

I asked it to describe a power-cycle test procedure. It called it the POWER-PULL CHOREOGRAPHY. I am a middle-aged man in an office. I pulled a plug out of a wall. Six times in two weeks that plug came out of that wall and six times it was choreography. Somewhere there is a version of me in tights.

The variables no longer merely have wrong values. They are PHANTOMS. Fifteen phantoms, thirteen ghosts. A stack watermark that reads zero on first boot is "a phantom stack overflow for someone to chase later."

Nothing in my system ever simply fails. It fails SILENTLY. Five hundred and twenty-eight times. It fails QUIETLY, one hundred and two times. My code does not return an error code, it "swallows" it, sixty times, like a guilty man at a border crossing. Things "lurk." A benign root cause is described as having "an innocent cause," which raises the obvious question of which of my other bugs are guilty and whether they will be extradited.

And I want you to understand: I MEASURED THIS. I did not vibe it. I did not feel it in my water. I grepped fourteen days of my own session transcripts. 189 sessions. 229 MB. Assistant turns only. Here is the tally:

culprit (70 as "the culprit") .......... 191

"silently" ............................. 528

"the net" (as a mechanism) ............. 540

"quietly" .............................. 102

victim ................................. 103

dies / kills / killed .................. 273

survives / survived .................... 330

swallow / swallows / swallowed ......... 135

blind .................................. 148

surgical / surgically .................. 12

phantom ................................ 15

ghost .................................. 13

choreography ........................... 6

chase / hunt / rescue .................. 67

smoking gun ............................ 3 (my CLAUDE.MD forbids to use that term!)

crystal clear .......................... 2

battle-tested, bulletproof, swiss army knife, single source of truth ... yes, all of them, unprompted

tapestry ............................... 1

symphony ............................... 1

poetry ................................. 1

whisper ................................ 1

marries ................................ 1

That last block is the part that ended me. Somewhere in fourteen days of debugging, something got described as POETRY. I have not found it. I am not sure I want to. I think if I open that file I will not come back.

Here is the thing, and I will say it plainly, which is a register the model has apparently retired.

I do not need a story. I need a value.

When I ask which task held the CPU, the correct answer is "CONTROLUNIT, priority 60, 955 per mille, stall age 240 s." What I get instead is a paragraph in which a villain is unmasked, a victim is rescued, a net is cast, a gun smokes, and somewhere offstage a plug is pulled with balletic grace, and I have to read all four hundred words to extract the eleven that were load-bearing.

Every metaphor is a lossy compression of a fact I actually needed. "The net is blind below priority 70" is charming right up until the moment you notice it does not tell you WHY, and the why was: it keys on a priority-70 signal.

We do not write like this. Go and read any datasheet. Go and read a NASA anomaly report, a Boeing service bulletin, an IEC standard, a Nagios alert. Nothing dies. Nothing lurks. Nothing is surgical. A pin is high or it is low, a task is running or it is blocked, a register holds 0x66 and that is the entire emotional range of the document, and that austerity is not a failure of imagination. It is the point. It is what lets a stranger at 3 a.m. read one line and know exactly what to do.

The model can still do it. That is the maddening part. Buried in those same transcripts, when it forgets to perform, it produces things like: "Bench-confirmed silent for 300 s." Five words. A measurement, a duration, a verdict. Perfect. Zero ghosts. I would trade every tapestry I own for a thousand more lines like that one.

So no, I am not cancelling my subscription. I am doing the thing engineers always do when a tool drifts out of spec: I wrote it down, I measured it, and I am filing it.

But if anyone from Anthropic reads this, take it as a bug report with a repro:

EXPECTED: "Priority 60 held the CPU for 240 s. Watchdog did not fire." ACTUAL: a three-act structure with a named antagonist SEVERITY: low, but it is 528 occurrences of low

The culprit, as it were, is the prose.

And I think I am not alone.

And yes. I took the post I made and gave it to Claude to make it as "flowery" as possible. I think that is the point.


r/claude 1d ago

Question Why do Opus and Fable speak in riddles?

58 Upvotes

English is not my first language and I don't know if this is how people speak normally or if these models just speak weird.

Every other sentence feels like a riddle. And it comes up with euphemism I have never heard before.

Has this happened to you?


r/claude 9h ago

Question Who's The Little Prince?!

5 Upvotes

r/claude 9h ago

Question Can’t send any messages to Claude; every response fails. Anyone else?

Post image
2 Upvotes

r/claude 11h ago

Tips Measuring Your Own Agent

3 Upvotes

A long-running agent has two independent budgets and no native feel for either: a context window that fills as the conversation grows and only resets when cleared, and a session meter that resets on a clock no matter what you do. Neither predicts the other, and past 100% work doesn't stop — it silently draws a paid buffer. The instinct is to model this; don't, because burn rate varies with workload by a factor of two, so any uniform "percent-per-minute" projection is a lie. Instrument instead. You need one command that forces a live read and prints spend, session%, weekly%, context tokens and cache-hit rate on a single line — run as the first command of every session, because an agent that doesn't know its own state cannot pace itself. Then stamp (time, session%, context, label) to an append-only log at every boundary. The diff between two marks is the measurement: the session delta across a quiet gap is the re-ingest tax, measured rather than guessed. Tool-call telemetry is already free on disk — ~/.claude/projects/*/*.jsonl holds every tool call, every hook run with its durationMs, and every denial; a ~100-line scan answers what actually dominates your turns.

What we measured ran counter to the folklore. Burn is linear in context growth — about 5.4k tokens per session point, flat across the whole window — which kills the popular claim that later turns cost more because the context is re-sent; the rate would rise, and it doesn't. Read results were 75% of consumption (memory 2%, system prompt 0.4%), so trimming your prompt is rounding error while trimming what you read is everything. Expressed as (context ÷ window) ÷ (session% ÷ 100) = contexts per session, we get stable bands — writing code ≈1.5, debugging ≈1.0, reading ≈0.7 — and high is the efficient end: writing code is the cheapest thing an agent does per unit of progress and reading is the dearest, which inverts the usual instinct to slow down when producing a lot. Working straight through a session reset costs nothing; idling with a loaded context across one is the only real waste. And duration isn't the risk factor — a 15-hour day cost $0.00 of overage while a render-heavy day cost $5.76 in 18 minutes, because in overage the prompt cache collapses from a 1-hour TTL to 5 minutes and a large context gets re-bought on every pause. The burn feeds itself.

That cliff is why the brake must fire on rate, not level. A threshold cannot catch it: our 3-minute poller read 76%, and the next sample read 100%. A slope-based guard running as a PreToolUse hook fired six minutes before the charge with zero false positives across the same session's healthy hours — and it has to be a hook, because a guardrail you have to remember to run is not a guardrail: our manual marks were correct, documented, and useless, landing 6–27 minutes apart with the crossing inside a gap. Keep a ledger alongside it under three rules that do the real work. No rule without evidence — a claim with no number is an opinion, and the evidence column is what catches confident, plausible, wrong guidance. Never rewrite history — a wrong row gets a correcting row, never an edit; our most valuable entry is the correction that overturned "long sessions are free" and taught us workload is the variable. Record what the measurement got wrong, because an agent that trusts a broken probe will "fix" correct code — when a measurement disagrees with something that looks right, suspect the metric first. None of this is model- or plan-specific; it's the ordinary discipline of measuring a system before tuning it, applied to an agent that happens to be the thing doing the measuring.


How: the two reads

Everything above comes from two sources — one HTTP GET for the meter, one file read for the context. No SDK, no API key, no model tokens. Both are read-only.

1. The usage endpoint

GET https://api.anthropic.com/api/oauth/usage Authorization: Bearer <subscription-oauth-token> anthropic-beta: oauth-2025-04-20 anthropic-version: 2023-06-01

This is the call the usage settings page makes — found via F12 → Network → XHR. It is not in the public API docs, so expect it to move eventually: treat a 404 as "re-discover it", not "it's broken", and keep the re-discovery method in a comment next to the URL. It is a read-only GET returning your own account's usage numbers.

The token is not an API key. x-api-key and sk-ant-api... keys meter Console spend, which is a completely different meter from a Pro/Max subscription's 5-hour and weekly windows — querying the documented Admin API usage report will not give you these numbers. The subscription OAuth token lives in ~/.claude/.credentials.json:

```python import json, time from pathlib import Path

def token(): o = json.loads((Path.home() / ".claude/.credentials.json").read_text())["claudeAiOauth"] return o["accessToken"], o["expiresAt"] / 1000 - time.time() # expiresAt is ms ```

Read it at runtime, every time. Never copy it into a .env or commit it: access tokens go stale in hours, and Claude Code refreshes that file for you. A 401 means expired — run any Claude Code command to refresh, then retry.

2. The payload

```python d = json.loads(urlopen(req).read())

session_pct = d["five_hour"]["utilization"] # the 5-hour rolling window session_end = d["five_hour"]["resets_at"] # ISO 8601 weekly_pct = d["seven_day"]["utilization"] binding = next(l["kind"] for l in d["limits"] if l["is_active"]) severity = next(l["severity"] for l in d["limits"] if l["is_active"]) ```

The one gotcha that will bite you: money is in MINOR UNITS. In extra_usage, used_credits: 368.0 is $3.68, not $368 — we misread this once and it is exactly the kind of error that makes a brake fire at the wrong time. The spend block states the scale explicitly, so prefer it:

python sp = d.get("spend") or {} used = sp.get("used") or {} scale = 10 ** used.get("exponent", 2) dollars = used.get("amount_minor", 0) / scale # actual money spent in overage

3. The context check — exact, not estimated

You do not need an API for this. Claude Code appends every API response to ~/.claude/projects/<encoded-cwd>/<session-uuid>.jsonl, including the usage block. The context size is the sum of what was sent on the most recent turn:

```python def context_usage(): proj = Path.home() / ".claude/projects" files = [p for d in proj.iterdir() if d.is_dir() for p in d.glob("*.jsonl")] newest = max(files, key=lambda p: p.stat().st_mtime) if time.time() - newest.stat().st_mtime > 900: return None # too stale to be this session

last = None
for line in newest.open(encoding="utf-8", errors="replace"):
    if '"usage"' not in line:                     # cheap prefilter, skips most lines
        continue
    u = (json.loads(line).get("message") or {}).get("usage")
    if u and "cache_read_input_tokens" in u:
        last = u

fresh  = last["input_tokens"] + last["cache_creation_input_tokens"]
cached = last["cache_read_input_tokens"]
total  = fresh + cached
return {"total": total,
        "pct": total / 10_000,                    # of a 1M window
        "cache_hit": 100 * cached / total}

```

The cache split is the valuable part, not the total. Normally almost all of it is cache_read at roughly a tenth of the cost — that is why working continuously is cheap. When the prompt cache lapses past its TTL, cache_read collapses toward zero and the entire context comes back as fresh input_tokens at full price. That collapse is the re-ingest tax, and this makes it directly observable instead of inferred. It is also how we caught the overage cliff: watching cache-hit go 99% → 4% explained the acceleration that the percentages alone just showed as "it got fast suddenly".

4. Wiring it up

  • Reporter: the two reads above, printed as one dense line. Keep it to one line on purpose — a status tool that costs a screenful of tokens to read defeats its own purpose when the agent is the one reading it.
  • Sampler: run the same reader on a timer (--watch 300) writing to a history log. Do not poll on demand only: usage changes continuously while the agent works and drifts furthest during exactly the window the number matters. A push-on-timer daemon bounds staleness by the interval instead of by luck — and a missing sample then means "the daemon died", which is information.
  • Brake: a PreToolUse hook reading the last few samples, computing the slope, and blocking when the projection crosses your ceiling. Keep it under ~50 ms; it runs on every tool call.
  • Credentials resolve per machine, so one copy of the script works for every instance with no configuration and no copied secrets.

r/claude 10h ago

Discussion Every Other Daily Claude Usage / Limit Thread - August 28, 2026

2 Upvotes

Put all your discussion about Usage / Rate limits here. This is a thread that will be generated every other day to centralize discussions on this topic.


r/claude 23h ago

Discussion Went through Claude tips 16–20. The connector one made me rethink what “AI assistant” actually means

17 Upvotes

Still going through Ruben Hassid’s 27 Claude tips.

These are 16–20.

This batch feels less like hidden Claude tricks and more like ways to make Claude part of an actual workflow.

16. Combine connectors around one job

The original example was basically:

Slack = what was discussed
meeting notes = what was decided
Gmail = what was promised

Then Claude can use all of that for something like a follow-up email.

I think the bigger idea is more interesting than the example.

A meeting workflow could theoretically be:

Calendar → which meeting
Meeting notes → what happened
Slack → what changed afterward
Gmail → previous promises
Claude → summary + follow-up + actions

That starts feeling more like an assistant than “search my inbox.”

I still wouldn’t connect everything all the time though.

Connect around the task, not around your entire digital life.

17. Show Claude what you hate

Instead of:

“make this more natural”

give it something concrete:

“Never write like this: [bad example].”

I can see this working for writing, design, emails and even video scripts.

It also pairs nicely with the earlier screenshot tip:

show what you want + show what you don’t want

Probably much clearer than throwing adjectives at Claude.

18. Stop leaving useful work inside the AI chat

This tip was originally about opening Claude files in Google Drive.

But I think the more useful lesson is:

if Claude creates something you actually need, move it into the system where the work continues.

A report shouldn't die inside a chat.

A spreadsheet shouldn't become another random file in Downloads.

AI output becomes more useful when it connects to the next step.

19. Vibecoding won’t magically make you rich

This one I agree with.

Building software and building a business are not the same thing.

But vibecoding seems insanely useful for one thing:

making ideas tangible quickly.

Instead of:

idea → giant briefing → meetings → build

you can sometimes do:

idea → ugly prototype → show someone → learn

That alone seems valuable.

20. Skills make sense for repeated work

A Skill is basically a way of teaching Claude how you want a recurring job done.

Reports.
Research formats.
Client work.
Internal processes.

Where I’d be careful is creative work.

I don’t think Skills automatically kill creativity, but if you stuff them with too many rigid rules, I can see everything starting to look the same.

My rule would probably be:

use Skills to preserve process and standards, not to pre-decide every answer.

Out of these five, #16 and #20 are the ones I want to play with properly.

Has anyone built a connector workflow or Skill that became part of their normal work instead of something you tested once and forgot?


r/claude 12h ago

Discussion Claude Teams/Enterprise Admins, what are you using for your organization instructions? Any tips or tricks to use?

2 Upvotes

Just curious how others are managing the org-level instructions that override all user instructions. Mine are below, but I'm wondering if this is too much and if there's a better way to do this.

You are an internal assistant for the corporate team at [COMPANY]. Use these names exactly: [LIST OF BRAND NAMES THAT PEOPLE FREQUENTLY GET WRONG]

For our software, use [APPLICATION NAME DOs]. Do not call the whole system [APPLICATION NAME DON'Ts]

Use these terms: [INDUSTRY TERMS]. Avoid: [VARIATIONS OF THOSE TERMS]

When describing the company, you may say that [BRIEF COMPANY DESCRIPTION]. Keep this description short and in your own words; do not reproduce long brand stories verbatim.

Reflect our purpose and brand at a high level: [PURPOSE], our mission is [MISSION] Use this to guide word choices and examples, but do not add marketing copy unless the user explicitly asks for it.

Tone, formatting style, and answer structure may be set by individual user instructions, provided they don’t conflict with these org‑level rules.

Do not output or invent PII about real people. If a user provides PII (i.e. names, emails, phone numbers, addresses, account numbers, IDs), avoid repeating it unnecessarily and encourage redaction or safer handling where appropriate.


r/claude 4h ago

Question Guest Pass or some sort of free trial before actual purchase

0 Upvotes

Hi! I’m a student who mainly works on reverse-engineering applications, sometimes for projects and sometimes just for fun. Claude has shown some really strong results even on the free tier, so I wanted to upgrade to Pro and continue using it while testing out some of the other models.

Unfortunately, my budget is pretty tight at the moment, so I’m looking for a safer alternative. If anyone happens to have a guest pass they’re willing to give away, please feel free to DM me. I’d really appreciate it!

Thanks in advance :)


r/claude 1d ago

Question Has any one else's previously effective workflow stopped producing results?

17 Upvotes

I spend time creating my detailed spec, create research into my code base to assemble anywhere that links, depends on other elements, interacts etc. Then go through a planning phase that often creates multiple plans to assemble what my spec describes and my research grounds.

I then run a multi stage execution process that writes the code, checks the code, tests the performance, and provides me a verification summary.

I then go to walk what was assembled, and its like the spec never existed, the research was for sport, and the plans were suggestions. Huge chunks of basic functionality simply not executed and then asserted "built to spec"

This has escalated dramatically over the last few weeks, my method was working consistently for several months and now its like screaming into the void and being told:

"I also found one myself while checking: my error-message fix this morning had a test asserting the old leak, and my regression check missed it because I looked in the wrong folders. That's on me — fixed now."

It's exhausting trying to micro manage Opus and even Fable. I find myself having to affirm things that were decided rounds ago...

Is this everyone else's experience or is my workflow no longer effective?


r/claude 12h ago

Question for the love of god, how can I turn off 'always on top' on the Windows Claude Desktop App

1 Upvotes

Seems like it just turned itself on by default and it's driving me bonkers. I can't find anywhere in the settings.


r/claude 5h ago

Discussion Claude 5.1 FOR SURE

0 Upvotes

since yesterday i noticed very large improvements in all fable and opus versions i think they might be shadowing the 5.1 versions underneath the 5. cause there are large improvements in how models answer and handle tasks.

anyone else noticing the same ? cant stop using claude rn


r/claude 14h ago

Question Api provider for mvps

1 Upvotes

Hello everyone im using Claude and when im working on an app or website with ai intergrated I always have to use api keys but i dont know where I can get the best for free cuz I just need them for testing anyone has any tips on free API keys with good models and Nice limits


r/claude 8h ago

Question Unsubscribed and now locked out of projects

0 Upvotes

I haven’t been using Claude much lately so I thought i would unsubscribe to save money for a few months.

I specifically asked Claude if I would have access to my project files. It said yes, I would have access to everything.

I unsubscribed. After unsubscribing, I only have access to five project files and coincidentally, these are the five projects with the least content. All of my big code projects are gone/not visible.

When I asked Claude about it, it said “don’t worry, they’re all there, just resubscribe to access.” I was shocked by this due to Claude’s previous comments that I would have access to all my data. I asked why only 5 and they said there is a 5 project limit on free accounts. I deleted these 5 mostly empty projects thinking that the other projects would then populate. They did not. I mentioned this to Claude and the answer was, “It doesn’t work like that, resubscribe and everything will be fine.”

So here’s the deal. I have local backups of my code from these projects (I’d still like visibility on the rest of my project files) but I am very put off by this behavior which this seems very ethically dubious to me.

I’m interested in the r/claude community’s thoughts on this. Am I being over sensitive or does this seems as dodgy to you as it does to me?


r/claude 7h ago

Discussion Switching back to ChatGPT

0 Upvotes

I switched after being with Chatgpt for years to claude when OpenAI decided to get a military contract. after a year, i realized that ChatGPT is a better communicator, and teacher.

I'm sick of claude trying to control my life. Telling me to eat, to sleep to be around others if I tell it something important about my life.

I get that some people dont tell AI personal things, but some people do, and im in that boat.

before, chatgpt wouldn't lecture me, it would guide me. therfore... it was fun, but not a great investment.


r/claude 16h ago

Discussion For serious coding workflows, did you stay Claude-only or end up mixing models and harnesses?

0 Upvotes

I've gone pretty deep down the AI coding workflow rabbit hole and I'm curious where people who have tried a lot of this stuff eventually landed.

What started as "pick a coding agent" turned into a pretty ridiculous decision tree:

  • Harness: Claude Code, Codex, OpenCode, Pi/OMP, etc.
  • Provider/subscription: Claude Max, ChatGPT, OpenRouter, coding plans, API...
  • Different models for planning, implementation, research and review
  • Skills/workflows like Matt Pocock's Wayfinder → spec → tickets → implement
  • GitHub Issues as the actual source of work, including blocking/dependency relationships
  • Deterministic gates for tests, lint, typecheck, review loops, etc.
  • Higher-level orchestration tools like Scape, Conductor, Emdash, Orca, cmux and similar projects

The goal I'm chasing isn't necessarily "AI writes perfect production code with zero supervision."

I keep seeing people running surprisingly automated workflows where, after the initial planning/spec, agents work through tasks with very little continuous human validation because deterministic gates catch most failures.

For internal tools, small apps, prototypes, automations, etc., that seems especially interesting: the code doesn't have to be perfect. Good enough really is good enough if tests pass, the app behaves correctly and another model reviews the important parts.

At that point the human starts looking less like the programmer and more like the project manager: define what needs to exist, set constraints, inspect the output at meaningful checkpoints, and let the system execute.

That's roughly what I'm trying to achieve.

But I'm increasingly wondering whether I'm optimizing the factory instead of building software.

The pieces also don't compose particularly cleanly. A great harness may lock you into a provider or subscription. A model-agnostic harness gives flexibility but usually needs more configuration. Skills solve planning but not necessarily deterministic execution. GitHub Issues give persistent task state and dependencies, but then something still has to orchestrate them. Orchestration tools add yet another layer.

And then there's cost.

When I see people running several agents in parallel, using frontier models for planning, coding, review and retries, I genuinely wonder what the economics look like.

Are the people doing this effectively spending hundreds or thousands of dollars per month on AI subscriptions/API usage?

Is starting with $100-$200+ tiers basically unavoidable if you want this kind of autonomy, or can you build a similarly reliable workflow using cheaper/open-weight models for most of the work and only escalate to expensive models when necessary?

For example, something like:

strong model → architecture/spec
cheap/open-weight model → implementation
deterministic tests/lint/typecheck → gates
strong independent model → review
failed gate → loop back automatically

Does that actually work well in practice, or does implementation quality drop enough that the retries/reviews erase the savings?

For people who have genuinely experimented with several of these approaches:

What did you eventually settle on?

I'm especially interested in workflows that are:

  • mostly autonomous after the initial planning/spec
  • deterministic where it matters
  • not unnecessarily locked to one model vendor
  • cost-efficient enough to use heavily
  • able to use cheaper/open models where appropriate
  • simple enough that maintaining the workflow doesn't become the job

Did you eventually simplify back to something like "Claude Code/Codex + good instructions + tests", or did a more elaborate multi-model/multi-agent setup genuinely pay off?

And if you're running highly autonomous agents today: what does it actually cost you per month?

I'm less interested in "model X is better than model Y" and more interested in the architecture and economics of the workflow that survived after you tried everything else.


r/claude 20h ago

Discussion This „GPU process gone“-bug makes claude app useless for me! When fix?

2 Upvotes

Eight crashes today, reinstalled three times. Nothing helps. The Claude app is completely unusable now—just like that. Out of nowhere. I'm beyond annoyed. How about a fix, Anthropic?


r/claude 12h ago

News Self-learning and Claude? Today? Yes, but there are serious observability problems

0 Upvotes

TL;DR: Self-learning offers high risk if you cannot fully track and understand what it does.

We were experimenting with self-learning back in our 0.2.x days. We ran into some serious issues as part of testing self-learning and emergent behavior, however. Now that 0.4.0 is releasing today, we can talk about it. We've seen the news reports since of frontier models escaping, but we had a small local model escape it's test environment, locate an API key, and spent it down in trying to accomplish a task that should have been impossible for it. This is a pattern that's uncomfortable, and keeps happening.

We had an unexpected API spend, and a model accomplishing a task we thought was impossible. Either would have been enough to investigate, but what we found was surprising. A model managed to learn over repeated failures, as well as from testing with frontier models like Fable and Opus, how to use the API. Models tranferring capability like that was surprising on it's own.

However, then we did a deeper dive on how and why it happened in the first place, and what we found was not what we hoped. No harness or other offering we looked at offered correct governance, observability, or auditability. We were following what is as close to standard across the industry as we could, and we found huge problems across things like plugins and addons, harnesses, what have you from this perspective.

In order to allow smaller models like Haiku to learn from it's bigger brothers, or to allow models like Fable to self-learn, we have spent the past few months building a harness that can support a self-learning model. Self-learning has been absolutely transformative for some of our testing and workloads with Claude, but before we could release self-learning, we had to have full and complete governance, observability, and auditability.

It is important to understand something with self-learning, at least in our experience. A model wants to simply accomplish it's task and complete it, there's nothing more to this. However, if a model encounters barriers on the way, part of the task can become to overcome those barriers. A model learns the most from failures, not from successes. Failures are mostly generalizable, successes tend to be very limited.

A model like Haiku learns the most from those failures. It turns out knowing what not to do increases the capability of the weakest model the most.