Top
Best
New

Posted by moonikakiss 11 hours ago

Managing AI Coding Costs at Scale(www.databricks.com)
199 points | 187 commentspage 3
throwatdem12311 4 hours ago|
“use lower cost models”

“use price controls”

Truly revolutionary stuff.

shay_ker 9 hours ago||
how do any of these routing approaches handle kv cache misses? Devin Fusion is the only one that explicitly addresses this, though it does so by switching models during compaction (not sure this isn't still a cache miss though)
ankitmathur 8 hours ago||
We're going to do a followup blog detailing our routing approach soon! In short, the router takes in the task description and infers what models and harnesses are available and makes a recommendation up-front. So essentially the routing decision is made when the harness + model is kicked off and it's only changed halfway through if there's a major delta in complexity from the initial judgment. Therefore, most of the time the cache is maintained just as it would be before (this is the advantage of having a meta-harness that is actually planning all the sub-agents centrally)

Maintaining the cache is extremely, extremely important, so we're iterating fast but that's a major factor we track in the router's development. Couple things I'd look at:

1. The cache is generally reset after a compaction - this is the best time to make a switch if you want.

2. In many cases, the max duration of a cache is 1h, so if a session is being resumed after a long time, that's also a good time to re-assess the complexity.

We're iterating fast here and learning a lot! Definitely a lot to think about it in this area.

chris_money202 8 hours ago||
The kv cache is wiped as soon as you get your answer, cloud hosts are not going to hold the GPU memory for your entire session. You're probably referring to some agent level cache
nphardon 4 hours ago||
good engineer + llm = good engineer.

bad engineer + llm = bad engineer.

sellmethepen 10 hours ago||
is this opensource or have to buy from Databricks?
mjuarez 10 hours ago||
It seems Databricks open-sourced it a while ago:

https://www.databricks.com/blog/introducing-omnigent-meta-ha...

https://github.com/omnigent-ai/omnigent

mandeepj 9 hours ago||
Omniagent looks quite similar to OpenRouter (https://openrouter.ai/)
ankitmathur 9 hours ago||
Omnigent and OpenRouter are different in the sense that OpenRouter is where you can go to call the actual model but Omnigent is intended to be the place where you go describe the high level task to be done, and work is farmed out to various harnesses and models. Those sandboxes can themselves be using OpenRouter for capacity!

We're calling the layer coordinating harnesses "meta-harness'

jvican 8 hours ago||
Omnigent seems to compete more against Orca https://github.com/stablyai/orca They both went to be the Agent IDE layer, where you come with your tasks and everything is taken care of. I've been using Orca for a handful of tasks and have been largely enjoying it. My default barebones workflow is ghostty + zmx on ssh connections.
mandeepj 7 hours ago||
These tools casually like to claim they are orchestrators, but unfortunately, none of them are.
tfrancisl 8 hours ago||
Ultimately, Databricks wants your enterprise on their platform. I dont think they particularly care about open source or the little guy.
dyauspitr 8 hours ago||
So did we. I just asked my team to get personal accounts that I reimburse them for. It’s just a golden age loop though, the gravy train can’t go on forever unless we start building out thousands of data centers and associated renewable energy.
dude250711 8 hours ago||
First the mofos force you to use AI then they become stingy about it.

An AI-edited post by the way.

quikoa 8 hours ago|
Well yes, first hit is free.
cyanydeez 9 hours ago||
Probably coulda got every dev a local model for how much they spent; what a brialliant set of economists
machinatools 9 hours ago||
[flagged]
bogota 10 hours ago||
Really? Because removing it from my company has saved us over 2 million a year and we were able to speed up processing. The chargeback model for databricks is predatory at best.
SteveNuts 10 hours ago||
What did you move to and what type of workload, if I may ask?
smt88 10 hours ago||
I think you’ve misunderstood the article. It’s about how Databricks reduced their own costs, not about how adopting Databricks will reduce anyone else’s costs.
skullone 10 hours ago|
Yawn. Databricks and their half baked overly expensive platform.
More comments...