From the creator of Redis; run LLM locally with ds4
91 points by fibo 5 hours ago | 17 comments

gchamonlive 7 minutes ago

  small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA)
This is local targeting high end consumer hardware like DGX Spark or AMD Ryzen AI Halo.

For our mere mortals that were kids not long ago and can't really believe we've got our hands on a x090 series targeting Qwen3.8 27b, https://github.com/noonghunna/club-3090 is the way to go.

I'm maintaining a web frontend for this, trying to at least. You can follow it here: https://github.com/gchamon/club-3090-server

reply
neomantra 22 minutes ago
I maintain a fork of ds4 as shared libraries and thus can be used with other languages via FFI, along with public builds/binaries [1]. I made ds4go [2] against ds4 using techniques inspired by yzma.

In addition to the library bindings, we have a small library of tools (workspace for view/edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them.

Even if you don't use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew.

Here's some TUI toy screenshots, sorry I still haven't released that code; it's of different quality than the others. [3]

EDIT: add ds4go TUI screenshot gist [4]

[1] https://github.com/NimbleMarkets/ds4/releases/tag/v0.8.20260...

[2] https://github.com/nimblemarkets/ds4go#install

[3] https://gist.github.com/neomantra/ae47422c8daf7a458212c93992...

[4] https://gist.github.com/neomantra/40180ade13df93290250ce8c6d...

reply
twoodfin 2 hours ago
https://github.com/antirez/ds4

The project GitHub page is a much better introduction for the hn crowd.

reply
simoiacos 2 hours ago
Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.

I'm also looking into expanding the protocol and the engine to support various steering techniques.

https://github.com/simoneiacomino/xenolith

reply
ilaksh 18 minutes ago
I wish someone would add Intel support to ds4. And also improve AMD support.

Maybe Intel and AMD should help them with that.

reply
vlowther 2 hours ago
It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
reply
HoldOnAMinute 30 minutes ago
How is this different from other LLM runners?
reply
csmlab_notes 3 minutes ago
[flagged]
reply
pulkitsh1234 33 minutes ago
curious, why did antirez go with C instead of something like Rust ?
reply
ilaksh 20 minutes ago
Antirez has been writing C for a million years so is much more familiar with it than Rust.

Also the goal of the project is to squeeze the absolute maximum performance and capability possible out of limited hardware resources (compared to clusters of B200s or something).

Does Rust even give you good access to low-level code on different platforms? And if so, how much extra work do you need to do to make it acceptable to the compiler? And is that work worthwhile if you are not going to get the security guarantees of normal Rust code? Is it a worthwhile tradeoff when the goal is performance?

Those are real questions by the way, not rhetorical. If Rust could work well for this type of project then I would like to know.

reply
GTP 30 minutes ago
Personal preference of the author, he made at least one video on YouTube on why he dislikes Rust. I think he finds it too cumbersome and not worth it when the software isn't security-critical (not that I agree, just reporting what IIRC his stance is).
reply
doctorpangloss 4 hours ago
the problem is the dsv4 checkpoint so quantized isn't very good
reply
ilaksh 27 minutes ago
Which ds4 checkpoint for which model exactly did you test? Don't they have multiple different versions and quantization levels?
reply
c0rruptbytes 34 minutes ago
not my experience

the ds4 quants were very good beating the unsloth quants https://github.com/michaelasper/benchmarks/blob/main/deepsee...

reply
dotancohen 26 minutes ago
That's quite the statement - unsloth quants are amazing.
reply
123-11292 3 hours ago
[flagged]
reply
fierycatnet 2 hours ago
Random comment but the name is funny to me, reminds me of Silicon Valley.

What are we going to name the company, how about Dwarfism 2.0? What happened to 1.0 Jared?

reply
seemaze 24 minutes ago
Dwarf Star is better than Dirty Socks, or Dynamic Slinky.. definitely not the worst backronym.
reply
dools 35 minutes ago
Smallulator
reply