pi. I wasn't even going that hard. I checked the logs for Sep 18 and I did just shy of 3b input with approx 98.5% cache and 5.8m output, which cost around $35. Most of the was a Rust code review exercise with 1 driving agent and a varying number of subagents (up to 6 some times). I do the same with with Sol med/high driving and Luna x-high reviewing and get at least as much done if not more in a day, but I'd use up two x20 weekly allowances for the week. Worth noting that token cost isn't super meaningfull on it's own, because DS is super token heavy (but also great at caching) compared to Sol. (my stats show DS uses 3x the tokens as Sol)
The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...
Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.
You might be overusing subagents. Especially with a chatty model like DS, you’ll be wasting millions of tokens on re-discovering the project and facts instead of actual reasoning.
I agree, and I started sending well defined review packages to Luna x-high instead, which is how I have this setup when using Codex (Sol drives, Luna async reviews, and Sol keeps moving). Same process with DS really dropped my DS usage by a lot and I think that might actually be the secret sauce. Especially if you use Luna through a subscription (probably the entry pro level would be enough). I'm not sure what a Luna equivalent open model is though. 5.6 Luna max was catching a lot of issues while my implementer keeps rolling. Last couple of days I had 3-4 Sol Mediums running with Luna x-high reviewers (async reviews) and one Sol medium orchestrator and Astra X-high to plan everything out.
Gotta be honest though, I don't love fiddling with this all the time. I would rather be working on my projects than evaluating my usage. Having DS in my back pocket should i need it is a relief. The providers get fiddly though too.
Such a high cache may be a sign that your agents are using tools inefficiently, thus taking too many turns. Or, like someone else said - you may use tools inefficiently many agents that rediscover stuff etc.
Pi out if the box tries to optimise system prompt size, which is not necessarily good and will cause exactly this effect for all but the most simple tasks.
What you want is to give enough context to the agent to minimize the amount of searching within the codebase etc.
If you want to track cache, what you should do, imho, is to check if you have cache expirations mid sessions (ideally you should not), and if you don’t then lower cache use is actually better - it means that your model doesn’t reread what it just wrote.
The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...
Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.