Top
Best
New

Posted by anerli 5 hours ago

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents(github.com)
Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp.

We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case.

Inference engines today all make a performance tradeoff. They are either:

- Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang) - Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama) - Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4)

Plus none of them are designed for running agents locally. Sessions are long, several often run at once, and you still want to use your computer for other things.

Magnitude is built for maximum performance on your hardware and running local agents:

- On-device compilation and tuning: Kernels are written with flexible parameters that are tuned on your actual device before the model runs. This gives you broad hardware compatibility with the same performance ceiling as hardware-specific kernels.

- Focus on best architectures: We write our tunable, highly efficient kernels for the most popular open-weights families. This allows us to achieve and surpass the performance of hardware or model specialized engines, without forcing ourselves to over-generalize at the cost of performance.

- Dynamic memory allocation: Magnitude reserves only enough memory up front to hold model weights. As your agent sessions grow, the memory heap dynamically increases, and frees itself when agents stop. Your hardware can still be used for other stuff while agents run.

- Hybrid paged attention: We borrow the best ideas from engines like SGLang to allow concurrent sessions to share prefix caches, but optimize placement for memory-adjacency so single-session performance doesn't suffer.

Magnitude is fully open source (Apache 2.0). We built it in Rust, including a custom GPU kernel runtime and autotuner. We take inspiration from the best innovations in inference from academics (e.g. FlashAttention, FlashInfer, TurboQuant) as well as other engines (e.g. SGLang radix attention) to reach the performance ceiling.

Benchmarked against llama.cpp with Qwen 3.6 35B A3B (4 bit), 64k context, no speculative decoding:

Metal (Mac M4 Pro 48 GB) - 92% faster decode (30 tok/s → 57 tok/s) - 9% faster prefill (466 tok/s → 507 tok/s) - 28% less per-agent memory usage

CUDA (DGX Spark) - 19% faster decode (49 tok/s → 58 tok/s) - 23% faster prefill (2,033 tok/s → 2,507 tok/s) - 27% less per-agent memory usage

Magnitude ships as a desktop app that you can easily connect with whatever agents you already use (Pi, OpenCode, Hermes, Codex, and more). It automatically runs models on demand when these agents actually need them, and shuts them down after inactivity. Here's what it looks like: https://www.youtube.com/watch?v=0qE8BWEZu7o

We're excited to push Magnitude further to let you run bigger models on the same hardware while continuing to improve performance. Our plans include:

- Expert streaming: store experts on RAM or disk and load them just-in-time. This lets you run models bigger than what otherwise would fit on your GPU.

- Kernel compiler: our current kernels tune a few parameters to fit your hardware. We can take this further with a fully custom compiler that automatically chooses how to fuse kernels and which implementations to use, to make it fit to your hardware even better.

- Multi-device utilization: Make the best possible use of all hardware on a system (CPU, GPUs, RAM, disk) by detecting these and automatically solving for the best model layout.

We'd love for more people to try it out and give us feedback. Feel free to comment here, we'll be around all day!

103 points | 43 commentspage 2
hypercube33 4 hours ago|
From your description looks like this isn't for AMD or Strix Halo at all? Also one of the things I'm not sure of but definitely plays a huge factor is the variant of the model you download - how does this help select the fastest version for your specific hardware / context size?
anerli 4 hours ago|
We support Vulkan as well, we just didn't mention it in the benchmark. When AMD or Strix Halo is detected the engine will use Vulkan.

Regarding model variants - our catalog includes different quantizations, and automatically assesses these against your hardware to determine which ones will fit in your memory and how fast they will run. This lets you pick a model to download based on your desired speed/intelligence tradeoff.

skohan 4 hours ago||
Do you have any plans to support ROCm?
anerli 3 hours ago||
We are actively benchmarking our Vulkan kernels to ROCm implementations in other engines to ensure that we can reach the performance ceiling with them. Vulkan is much more portable and also works on non-AMD hardware even though it can be more awkward to write kernels for. If we find that Vulkan is not sufficient for reaching the same performance as ROCm, we'll consider adding it as a backend
kenzic 5 hours ago||
How long does tuning take (on an M3 MacBook Pro for example)?
anerli 4 hours ago|
Tuning is a one-time process that takes around ~1 minute whenever you download a new model. This is generally enough time to tune all the kernels' parameters to the point where tuning any longer asymptotes. Time can vary a little based on the hardware though.
kenzic 4 hours ago||
Wow, that's impressive.
amirhesham 5 hours ago||
Oh this is so cool. Curious about the business model, too.
anerli 3 hours ago|
As mentioned here https://news.ycombinator.com/item?id=49912327 we'll eventually build an inference cloud for hybrid workloads. For now though, we're focused on making the inference engine great!
paulgerhardt 3 hours ago||
Trying to run this but keep hitting bugs. Can you open up issues reporting on your repo?
anerli 3 hours ago|
Issues should already be open for anyone to report!

https://github.com/magnitudedev/magnitude/issues

Let me know if you keep running into problems for some reason

sgtwompwomp 5 hours ago||
This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too
anerli 3 hours ago|
I would say the overall idea of trying to achieve performant inference for agent workloads is the strongest commonality with Wafer.

It's not a coding agent running on your device optimizing the kernels, we have a system for writing kernels that can be tuned on the target device automatically. So we write the efficient high level kernel structure with tunable parameters, then it fits to whatever hardware it's actually running on.

p-e-w 5 hours ago||
What is the business model?
anerli 4 hours ago|
We envision a future where workloads are hybrid. Average consumer hardware will be able to handle a lot with local models, but you’ll still want to use cloud models for harder tasks. Magnitude will make it seamless to switch between the two, even for the same tasks (without breaking your prefix cache). We’ll charge per token for our inference cloud, using the same efficiencies we unlock for local inference to pass the savings on to you.
yolandac 4 hours ago||
does it allow us to run larger models that weren't possible before?
anerli 4 hours ago|
Right now, since we use less memory for KV, you have more room for model weights when you're running longer sessions.

However we also have expert streaming on the roadmap. This will let you run mixture-of-experts models with unused experts offloaded to RAM or disk, and load them only when needed. This means you'll be able to run models that wouldn't otherwise fit in your GPU memory.

ls-a 1 hour ago|
[dead]