What secret sauce do they have?
Quant company usually squeezing every penny.
pair it with codewhale, 50 agents, 200 MB of ram.
But I find it having a pretty significant problem with tool calling - no idea why, but tool calling with it is SLOW. As long as the model is reasoning, all good. But give it a bunch of tools and it becomes extremely slow.
Am I the only one experiencing this?
China has zero energy concerns in terms of energy production - not literally zero, but they’d be able to prioritize other dimensions and not necessarily worry about efficiency
Here they are though releasing models that sip resources
That's irrelevant when you use $/task as the metric, which the OP does use.
That's how I handle the Qwen27B and 35B
What do you mean by "redirect it to useful output"? Could you give an example? This sounds interesting.
Imagine if they had GPU resources of western labs.
SV companies get way too comfortable when they have enough in the bank to stay running more than three months.