跳到主要內容

ds4評測:Redis之父打造的本地DeepSeek V4引擎 | ds4 Review: Redis Creator's Local DeepSeek V4 Engine

By Kit 小克 | AI Tool Observer | 2026-10-04

🇹🇼 ds4評測:Redis之父打造的本地DeepSeek V4引擎

ds4(別名 DwarfStar 4)是 Redis 原作者 Salvatore Sanfilippo(antirez)開發的開源本地 LLM 推理引擎,今天在 Hacker News 衝上熱門榜,討論聲量壓過不少新模型發布。它的賣點很單純:不當「什麼模型都能跑」的通用工具,只專心把 DeepSeek V4 Flash(以及高記憶體機器上的 DeepSeek V4 PRO、GLM 5.2)跑到最順,從模型載入、KV cache 管理到內建 coding agent 全部自己寫、自己測。

ds4 是什麼?為什麼本地LLM圈在討論它?

ds4 是一個單一用途的推理引擎:只支援少數幾個精選模型,但針對這些模型做深度優化。對比 llama.cpp、Ollama 這類「什麼 GGUF 都能跑」的通用方案,antirez 的邏輯是反過來的——窄而深,把 prompt 渲染、工具呼叫、記憶體分頁全部當成一個整合系統設計,而不是拼湊起來的外掛鏈。這種「反通用化」的設計哲學,正是它在 HN 上引發熱烈討論的原因。

ds4 支援哪些硬體?Mac也能跑嗎?

可以。ds4 用 make 直接編譯支援 Apple Silicon 的 Metal 後端,另外也有 CUDA(Ada、L40S、DGX Spark)與 ROCm(Strix Halo、Framework Desktop)的專屬編譯選項。換句話說,從 MacBook 到多卡 GPU 工作站都有對應的建置路徑,不用自己改 Makefile。

ds4 效能如何?實測數字告訴你

根據官方 benchmark 文件,在 8 張 L40S 的配置下,ds4 可以做到約 120 t/s 的聚合生成速度與 2000 t/s 的 prefill 速度。它內建 ds4-bench 工具,可以用自己的硬體在不同 context 長度下量測實際表現,不用只相信官方數字——這也符合「好不好用,試了才知道」的精神。

怎麼安裝 ds4?新手也能跟著做

  • 下載原始碼:git clone https://github.com/antirez/ds4.git
  • 下載模型權重:./download_model.sh q2(會自動從 Hugging Face 下載,支援斷點續傳)
  • 編譯:Apple Silicon 用 make,CUDA 卡用 make cuda-generic,DGX Spark 用 make cuda-spark
  • 啟動內建 API server 與 coding agent 即可開始使用

ds4 值得換掉現有工具嗎?

如果你只是想要「一個能跑各種模型的瑞士刀」,ds4 不是答案——它刻意不支援任意 GGUF。但如果你明確要跑 DeepSeek V4 系列,而且在意整合度(不想自己拼 llama.cpp + 自訂 agent 腳本),ds4 的「垂直整合」策略確實解決了一個真實痛點:本地 LLM 工具鏈長期碎片化的問題。由 Redis 作者主導,至少在工程品質上有一定的信任基礎,但它仍是早期專案(2026 年 5 月才首次發布),穩定性與模型支援廣度都還在成長中。

好不好用,試了才知道。


🇺🇸 ds4 Review: Redis Creator's Local DeepSeek V4 Engine

ds4 (aka DwarfStar 4) is an open-source local LLM inference engine built by Redis creator Salvatore Sanfilippo (antirez), and it shot to the top of Hacker News today, out-discussing several model launches. Its pitch is narrow by design: instead of trying to run every GGUF model under the sun, ds4 focuses entirely on running DeepSeek V4 Flash (plus DeepSeek V4 PRO on high-memory machines, and GLM 5.2) as well as possible — with model loading, KV cache handling, and a built-in coding agent all built and tested together as one system.

What Is ds4 and Why Is the Local LLM Crowd Talking About It?

ds4 is a single-purpose inference engine: it supports only a handful of curated models, but optimizes deeply for each one. Where tools like llama.cpp or Ollama bet on broad GGUF compatibility, antirez bet the opposite way — narrow and deep, treating prompt rendering, tool calls, and memory paging as one integrated system rather than a chain of bolted-on plugins. That anti-general-purpose philosophy is exactly why it's generating so much discussion.

What Hardware Does ds4 Support?

ds4 compiles directly to Metal for Apple Silicon via make, with dedicated build targets for CUDA (Ada, L40S, DGX Spark) and ROCm (Strix Halo, Framework Desktop). There's a build path from a MacBook to a multi-GPU workstation without hand-editing Makefiles.

How Fast Is ds4? The Real Numbers

According to the project's own performance docs, an 8x L40S setup hits roughly 120 tokens/sec aggregate generation and 2000 tokens/sec prefill. ds4 ships its own ds4-bench tool so you can measure these numbers on your own hardware at different context lengths — don't just trust the vendor's chart, test it yourself.

How Do You Install ds4?

  • Clone the repo: git clone https://github.com/antirez/ds4.git
  • Download weights: ./download_model.sh q2 (pulls from Hugging Face with resumable downloads)
  • Build: make on Apple Silicon, make cuda-generic for CUDA cards, make cuda-spark for DGX Spark
  • Launch the built-in API server and coding agent

Is ds4 Worth Switching To?

If you want a Swiss-army-knife runner for any model, ds4 isn't it — it deliberately refuses broad GGUF support. But if you specifically run the DeepSeek V4 family and care about integration (instead of duct-taping llama.cpp to custom agent scripts), ds4's vertical-integration approach solves a real pain point in the fragmented local-LLM toolchain. Having Redis's creator behind it buys some engineering credibility, but it's still an early project — first released in May 2026 — so stability and model breadth are still maturing.

好不好用,試了才知道。(Good or not — you only know after trying it.)

Sources / 資料來源

常見問題 FAQ

ds4是什麼?

ds4(DwarfStar 4)是Redis創辦人antirez開發的開源本地LLM推理引擎,專門優化DeepSeek V4 Flash、DeepSeek V4 PRO與GLM 5.2的本地執行效能。

ds4支援哪些硬體?

支援Apple Silicon(Metal)、CUDA(Ada、L40S、DGX Spark)與ROCm(Strix Halo、Framework Desktop),各平台都有專屬編譯指令。

ds4跟llama.cpp有什麼不同?

llama.cpp走通用路線,什麼GGUF模型都能跑;ds4刻意只支援少數精選模型,換取更深度的整合與優化。

ds4的效能如何?

官方benchmark顯示8張L40S配置下約可達120 t/s聚合生成速度與2000 t/s prefill速度,實際數字會因硬體而異,建議用內建的ds4-bench自行測試。

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code