Top
Best
New

Posted by fibo 4 hours ago

From the creator of Redis; run LLM locally with ds4(dwarfstar.sh)
98 points | 25 comments
liuliu 2 minutes ago|
If you are interested in high-end models with high-end Apple Silicon, also try out Local Code: https://releases.drawthings.ai/p/public-beta-of-local-code-b... It is currently in TestFlight (and will open-source next week), supporting vision with DeepSeek 4.1 Flash, Qwen 3.8 27B and DeepSeek 4 Flash 0731 without vision. Custom quants & SSD streaming to make these big models work with 64GiB and above devices (and of course, Qwen works with devices with 16GiB and above).
ttoinou 11 minutes ago||
Ive been using this since it was initially released with deepseek v4 flash, and it is absolutely the best launcher ever on my m5 max 128gb

Now Ive been running qwen 3.8 flash next for more than a week and it’s doing great, really fast and super long context windows. Sometimes the model is behaving stupidly by not remembering something I said earlier but it could be also a problem from the agentic AI harness. Im using oh my pi but Im wondering what people are using ds4 with here ?

gchamonlive 26 minutes ago||

  small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA)
This is local targeting high end consumer hardware like DGX Spark or AMD Ryzen AI Halo.

For our mere mortals that were kids not long ago and can't really believe we've got our hands on a x090 series targeting Qwen3.8 27b, https://github.com/noonghunna/club-3090 is the way to go.

I'm maintaining a web frontend for this, trying to at least. You can follow it here: https://github.com/gchamon/club-3090-server

twoodfin 1 hour ago||
https://github.com/antirez/ds4

The project GitHub page is a much better introduction for the hn crowd.

neomantra 40 minutes ago||
I maintain a fork of ds4 as shared libraries and thus can be used with other languages via FFI, along with public builds/binaries [1]. I made ds4go [2] against ds4 using techniques inspired by yzma.

In addition to the library bindings, we have a small library of tools (workspace for view/edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them.

Even if you don't use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew.

Here's some TUI toy screenshots, sorry I still haven't released that code; it's of different quality than the others. [3]

EDIT: add ds4go TUI screenshot gist [4]

[1] https://github.com/NimbleMarkets/ds4/releases/tag/v0.8.20260...

[2] https://github.com/nimblemarkets/ds4go#install

[3] https://gist.github.com/neomantra/ae47422c8daf7a458212c93992...

[4] https://gist.github.com/neomantra/40180ade13df93290250ce8c6d...

simoiacos 2 hours ago||
Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.

I'm also looking into expanding the protocol and the engine to support various steering techniques.

https://github.com/simoneiacomino/xenolith

ilaksh 36 minutes ago|
I wish someone would add Intel support to ds4. And also improve AMD support.

Maybe Intel and AMD should help them with that.

simoiacos 15 minutes ago||
Yeah I see the value but I built Xenolith to target smaller models.

I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?

ilaksh 4 minutes ago||
The recent Intel GPU/AI cards. Really the same type of models as ds4
vlowther 1 hour ago||
It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
ttoinou 10 minutes ago|
I’m already able to use 1M context windows with the same machine than you and same model. Strange
HoldOnAMinute 48 minutes ago||
How is this different from other LLM runners?
ilaksh 1 minute ago||
Emphasis on performance and usable coding/agentic ability for consumer AI hardware. Does not attempt to handle all models or hardware at once but rather focuses on optimizing the best options for that category of hardware.
pydry 17 minutes ago|||
My instinctive reaction from the readme is that it isnt. It's apparently a vibe coded knock off of llama.CPP.
csmlab_notes 22 minutes ago||
[flagged]
pulkitsh1234 51 minutes ago||
curious, why did antirez go with C instead of something like Rust ?
ilaksh 39 minutes ago||
Antirez has been writing C for a million years so is much more familiar with it than Rust.

Also the goal of the project is to squeeze the absolute maximum performance and capability possible out of limited hardware resources (compared to clusters of B200s or something).

Does Rust even give you good access to low-level code on different platforms? And if so, how much extra work do you need to do to make it acceptable to the compiler? And is that work worthwhile if you are not going to get the security guarantees of normal Rust code? Is it a worthwhile tradeoff when the goal is performance?

Those are real questions by the way, not rhetorical. If Rust could work well for this type of project then I would like to know.

Aeolos 46 seconds ago||
Yes, Rust gives you great access to low-level code on different platforms, including SIMD. It is also alias-free by default, and gives you excellent primitives to write multi-threaded code with compile-time correctness guarantees, which is how projects such as zlib-rs end up significantly faster than their C counterparts.[1]

It's about as good as it can get for this kind of code.

[1] https://www.reddit.com/r/rust/comments/1ixt1ei/zlibrs_is_fas...

GTP 48 minutes ago||
Personal preference of the author, he made at least one video on YouTube on why he dislikes Rust. I think he finds it too cumbersome and not worth it when the software isn't security-critical (not that I agree, just reporting what IIRC his stance is).
doctorpangloss 3 hours ago|
the problem is the dsv4 checkpoint so quantized isn't very good
ilaksh 46 minutes ago||
Which ds4 checkpoint for which model exactly did you test? Don't they have multiple different versions and quantization levels?
c0rruptbytes 52 minutes ago||
not my experience

the ds4 quants were very good beating the unsloth quants https://github.com/michaelasper/benchmarks/blob/main/deepsee...

dotancohen 44 minutes ago||
That's quite the statement - unsloth quants are amazing.
More comments...