Top
Best
New

Posted by twapi 18 hours ago

Maximizing the value of your Claude Code sessions(claude.com)
213 points | 122 commentspage 2
NoDodgeQuestion 16 hours ago|
Bro: superintelligent machine line go up AI AGI software solved automate everything

Also bro: Run /clearbetween tasks. This prevents prior irrelevant context from being sent back to the model, which can reduce token usage. Set your model and effort level before you start. Changing either one mid-conversation can bust your prompt cache, which can increase token cost. @-mention files instead of naming them. The file gets attached to your message directly, which saves a Read call, or a search if Claude has to go find it. Add quiet flags to noisy commands, or run them in a subagent. Command output is added to the conversation just like a file, and stays there for the rest of the session. Run /context once in a fresh session. It shows what's loaded (CLAUDE.md, MCP tool definitions), so you can cut out anything unnecessary. /compact before you take a break from your keyboard. The prompt cache expires after an hour, and summarizing a conversation is much cheaper while it's still cached.

NoDodgeQuestion 16 hours ago||
Author not bro, sorry misgender
cyber_kinetist 4 hours ago|||
It has been a while since bro has become a gender-neutral term, particularly in younger circles...
pdpi 15 hours ago|||
I'd argue that women can be bros too, especially when using the word in this sense.
NamlchakKhandro 6 hours ago|||
There are no women on the internet.

I remember when the internet was an exchange of ideas instead of using gender to justify value of bad ideas

runeblaze 12 hours ago|||
puts on my etiquette hat

don’t do that, it is weird, use “bruh” or “dude”

DangitBobby 14 hours ago||
What do you want? Lacking omniscience, even the smartest superintelligence imaginable has to do more thinking to deal with worse inputs.
DangitBobby 9 hours ago||
Still waiting to hear what you want.
mccoyb 16 hours ago||
I mean, it feels hard not to laugh at this type of blog post. My cynical interpretation is that this is a type of passing the buck to engineers in enterprise settings ("Stop spending tokens. Did you read the value maximization blog post? It is your fault.")

Oh yes, Claude will do all sorts of different things -- it depends on how you use it! You should totally learn all of these little finicky things ... because now completing your tasks cost money. It's not "free" anymore haha like when you used your old text editor, what are you a grandpa?

Oh, and those things will definitely change, as we (the priests of Claude) are vibe coding the system you use to do your little "tasks" ... right, you can't see how it works ... the code is not available. It's all good, just trust us -- we're totally looking out for you.

I mean it is utterly ridiculous to talk around this model of development. There are so many walls between you and doing the thing you want to do.

Agents are great, but the notion of "best tricks" for how to best use an opaque costful tool which will, by all odds, be completely different in a few months time is quite funny.

You know what won't change? A fucking text editor. Or your pi config, or a local model you run and trust.

DangitBobby 14 hours ago||
Not sure how this level of cynicism is even remotely warranted. The post helps people who don't understand LLMs very well get the most out of Claude. Your incentives here are actually aligned with Anthropics since both of you want fewer tokens inputted and outputted per task completed.
mccoyb 14 hours ago|||
Perhaps my enterprise cynicism is not warranted, but my other comments refer to accurate descriptions of reality: Anthropic wants to place their opaque system between you and any computational task that you wish to perform. Do you contest this or think it is not accurate?

Why do you think that Anthropic wants fewer tokens inputted and outputted?

DangitBobby 12 hours ago||
Because they sell subscriptions and tokens cost them compute, and their margin lives in the difference between what your subscription pays in and what you cost them in compute.

They have also been supply constrained on compute and if users cost them less in compute they can more subscriptions and less customer frustration.

I agree they want you to have a subscription. That doesn't mean they aren't aligned with their subscribers.

fcarraldo 10 hours ago||
Subscriptions are a very small part of their overall revenue (estimates have been between 5% and 20% based on financial reporting). Enterprise users are charged per-token, and maximal input/output tokens nets them maximal revenue.
DangitBobby 9 hours ago||
They still want you to hit the cache because their margin is higher on cache hits. That's actual compute they don't have to pay for and they don't have to have capacity for because they are supply limited on the compute side.

And the unit economics need to be there because there are competitors in the space. They can't just skin you on tokens or you'll jump ship.

NamlchakKhandro 6 hours ago|||
1000% warranted. not sure what level of peak echo bubble you live in where this level of critiscism feels like you need to defend a ONE TRILLION DOLLAH company.

seriously... priorities yeah?

mnahkies 15 hours ago|||
Through my weekend experiments, I've found I can get way better outcomes, and an order of magnitude less cost with my slapped together sandboxed omp setup plus ZDR openrouter models (DeepSeek, Kimi, etc) than I've ever seen from Claude Code at work.

Everything is version pinned and a deliberate choice to change, and a git revert away from changing back.

TBF the models may change underneath me to some extent still, but the cost benefit of running them myself doesn't pan out yet (for agentic coding at least, don't have enough local vram to get a usable context window and generation speed, self hosting on runpod or similar isn't economically sensible for my current consumption though I have tinkered with it)

bmitc 6 hours ago|||
I fully agree. All these blog posts are basically features they should be implementing. They advertise they are replacing software development, but then these tools require a massive amount of overhead akin to having to train new hires. But these tools never actually learn and are not trainable, and Anthropic releases a blog post every six months about how to re-invent your workflow. Even the author of Claude Code just told everyone they should delete all their `CLAUDE.md` and skills every six months.

It's wildly lazy.

runeblaze 16 hours ago|||
dude, if you try to do harness development yourself you will realize that most things said in this blogpost is shared with any ${sufficiently_advanced_harness}. this is not really claude-specific, this is just how this class of tools, OSS or not, works
mccoyb 16 hours ago||
That's not my complaint. I know well the concerns of agent harnesses. My complaint is that this is a low-dimensional projection of a system which I have no insight into, and therefore, I cannot evaluate the tips myself against their source.

Am I to believe the creators, knowing full well that the source will, as Boris Cherny put it in a recent interview, be deleted and rewritten from scratch at the release of the next big model?

Further: I'm responding to content in the blog post itself:

> Until pretty recently, the tools you wrote code with were a flat fee (or free). Your editor cost the same whether you fixed one test or fifty that afternoon, so an individual task didn't really have a price of its own.

I find this type of prose ridiculous. It conveys "this is the way things are now, get used to it".

Does that make sense?

runeblaze 14 hours ago|||
> which I have no insight into

i guess you do? claude code is the commercial closed sourced version provides by ant. reading a mini version of vllm or sglang and then read codex source code or grok build source code will teach you all things taught by this article, fully in the open

it is like saying that you have no insights into some $commercial_db_system which is kinda true but imagine if the article is to teach you indices, query normalization, etc..

bluefirebrand 15 hours ago|||
> I find this type of prose ridiculous. It conveys "this is the way things are now, get used to it".

This is absolutely what AI companies and AI lovers want you to believe

csallen 16 hours ago|||
I'm trying to understand your point of view, but it kind of just sounds like you're against learning how to use tools efficiently?

I mean, agentic coding software is hardly the first tool to exist where learning some idiosyncrasies of how to use it well can result in more efficiency and cost savings.

mccoyb 16 hours ago||
It's very easy to understand:

- I'm happy to learn how to use tools efficiently

- I like to be able to inspect my tools

- I'm against tools changing underneath me

Are you against any of these points?

crthpl 7 hours ago|||
what specifically has changed about Claude Code?
NamlchakKhandro 6 hours ago||
https://cchistory.mariozechner.at

and as is normal for hosted models, almost everything... based on load flucation they may even send your prompt to a quantised model

csallen 16 hours ago|||
I think I'm happy about the first two. The third I suppose I care less about, just because I've kind of become used to it from decades working on the internet where many businesses/tools/apps are more like services and less like physical tools that never change.
RossBencina 16 hours ago|||
> The third I suppose I care less about, just because I've kind of become used to it

I've never become used to it. My impression is that the constant churn has accelerated. Plausible drivers are (1) normalize novelty as desirable (like fast fashion), (2) product developer/designer incentive structures that reward revolutionary change over progressive refinement. The global switch to subscription models and continuous deployment didn't help.

> more like services and less like physical tools that never change.

I'm not sure that constant change is a characteristic feature of services, especially not professional services.

It used to be that you bought a piece of software and used that version until you decided it was worth upgrading, like a particular physical tool. The software still evolved, just like the design of physical tools can, in principle, evolve.

All that said, agentic AI tooling is evolving so rapidly I'm not sure an expectation of stability is realistic.

Wowfunhappy 14 hours ago||
> It used to be that you bought a piece of software and used that version until you decided it was worth upgrading, like a particular physical tool.

But Claude is running on someone else's computer, not yours, so it's not Photoshop so much as AWS. Or a rented server farm, if AWS is too new school for you. Of course there's an ongoing cost! And if you configure the server to use more electricity, you get billed more.

If you want to do agentic tooling locally, you can do that—the models aren't quite as good, but they're not bad either. But be warned, for the large models you're going to have to acquire some serious hardware, to the point where you may wish you'd chosen to just rent it instead!

mccoyb 16 hours ago|||
I agree that the third seems to be implied by industry, but I'd argue that it's not clear that it is necessary -- and it is subtle whether or not it is beneficial?

My contention is that we should be building towards less churn, not more. I'm aware that some churn is the cost of engaging in any sort of enterprise, but I'm deeply suspicious of an AI company inserting themselves between me, and the tasks I wish to do with my device -- with a completely opaque system that I can't really "learn".

dist-epoch 14 hours ago||
I'm curious if you also laugh at articles about how to reduce your AWS bill, or how to add indices to Postgres such that you can run it on cheaper hardware.
mhitza 11 hours ago|||
You can generally understand your AWS cloud usage, and waste can be self evident with their existing tools. Not at all with llms.

A postgres index post is unlikely to reach front page. It's already part if the docs, and should include more context to be read worthy.

They are not equal comparison.

This before the fact that there is no guarantee that a model follows your agent instructions (plenty of easy to reach for research on it), and you also get suggestions by devs at these companies to wipe parts of your model's instructions because the model is better now tm.

If cloud providers change their billing quasi monthly, and if you'd need to fiddle with your indexes every couple of days. I'm not sure we'd be using them as much.

There is interesting information about the inference pipeline, but almost too late to the party (by at least a year), and for which audience? Techies understand in broad strokes the tech if they are interested, normies will definitely not read it.

All that to say, that yes, it's worth having a laugh. If for nothing else, as a release valve for all the problems they create in the real non-VC world.

Anthropic is IPOing in October according to news, you might be interested in investing.

dist-epoch 2 hours ago||
I thought that software engineers were supposed to do, you know, engineering - solving hard problems, dealing with uncertainty.

But you might be right, engineering around the difficult LLM primitive might be a task which is just too hard for your typical software engineer, as you said, they want predictability, hand holding, determinism, most are unable to deal with the real world which is not a spherical cow in a vacuum. So I guess they can stick to simple very well understood primitives like EC2 or Postgres and leave dealing with LLMs for others.

Banditoz 13 hours ago|||
I think you're conflating two different things here. I am not aware of any DB optimization articles that say "trust me bro, throw your data, don't build indices, it'll Just Work™!"
ahurmazda 15 hours ago||
What’s the point of running /clear vs starting a brand new session. At least with the latter I have session history, no? Pardon my ignorance since Claude isn’t my primary driver
g4cg54g54 13 hours ago||
Its all lies anyhow:

- https://github.com/anthropics/claude-code/issues/47756 > [BUG] /clear bleeds into the next session (what also breaks cache)

- https://github.com/anthropics/claude-code/issues/47098 > [BUG] new sessions will *never* hit a (full)cache

mikeocool 15 hours ago|||
As far as I've seen /clear is the same thing as starting a new session.

If you type /resume right after clear, the first thing in the list is the session you just cleared.

andai 14 hours ago||
Not 100% sure what clear does, but starting a new session invalidates the cache*, whereas I assume clear only removes part of the context, so it should be cheaper and faster.

* In theory the system prompt is always the same and should therefore be cached, but in practice there's some dynamic strings in there so it doesn't work that way. (Unless they changed this recently.)

olsondv 16 hours ago||
After having used Codex for a promotional month, and now using Claude, Claude is not as efficient with finding relevant information. I can give it the one file it should be using and then it goes off and greps parent directories for more context. It’s also incredibly slow at producing results because of this side work. In this article, it seems like they are catching up to what GitHub copilot users had already been doing since the cost restructuring in June.
docheinestages 10 hours ago||
Anthropic should build a harness (and model) that smartly takes care of all these points. Not requiring the user to do the manual work. All I see are excuses because they cannot handle the load and enforce strict quotas on users, all while OpenAI constantly resets their quotas.

With Qwen 3.8 27B, we're one step closer to on-device LLMs that can replace subscriptions.

Petersipoi 7 hours ago|
The amount that OpenAI resets their quotas is nuts. It's like, every 2 days I swear. Feels so fucking good. Whenever I think about switching my $200 plan back to Claude for a month I'm reminded that they still have a 5 hour usage limit, which feels so absurd now that I've used Codex for a couple of months.
wjakob 11 hours ago||
When rewinding to an earlier turn, what if that turn is more than 1hr old? Can this cause KV-cache misses compared to continuing the conversation?
8note 10 hours ago||
how many times it stays there i think doesnt give the best comparison

if youre working on the same codebase, that cache stays quite relevant, and i dont think they make the case that clearing and reading the same couple files over and over again is cheaper that relying on it already being cached. same with doing some of the same teaching claude the right way to approach changes in that codebase again and again.

what would be nice is pulling back and reusing an earlier part of the cache for the later two tasks, but claude code doesnt make that particularly easy, and using an LLM to pick where to go back to isnt really gonna save much when it reads all the same text again.

pzo 14 hours ago||
> Set your model and effort level before you start. Changing either one mid-conversation can bust your prompt cache, which can increase token cost.

I know we supposed to do this but is there any particular reason why such things cannot be supported? I thought its running on same model just different settings like reasoning. This would be super useful.

myshapeprotocol 4 hours ago||
Great practical insights on streamlining AI coding sessions. Maximizing workflow efficiency like this is essential for modern development.
ChrisGreenHeur 4 hours ago|
Is this comment written by ai?
aleksiy123 15 hours ago|
Is it possible to have some kind of script to keep your cache warm, or auto compact or something.

I sometimes just leave some goals or something running before I go to bed or out and I don’t want to pay the cache text when I come back.

cube00 14 hours ago|
If everyone does it they lose the memory savings they're getting by expiring the cache.
8note 10 hours ago||
which is to say that they set their ttl too short, it should be longer than people spend at lunch
More comments...