Squeezing a 284B MoE into a 128GB Laptop
DwarfStar 4 (ds4), a C inference engine developed by Redis creator Salvatore Sanfilippo, fundamentally shifts the barrier to running frontier models locally: the limiting factor is no longer whether you can afford high-end accelerator cards, but whether your hardware has enough memory. DeepSeek V4 Flash, a colossal 284B Mixture-of-Experts (MoE) model, natively demands staggering amounts of video RAM. Mainstream consumer hardware historically stood no chance of handling a parameter footprint of that scale.
ds4 tackles this challenge with asymmetric 2-bit quantization, aggressively compressing routed experts while preserving core compute capability. Applying aggressive quantization across tens of billions of routed expert parameters proves that selective compression along critical execution paths does not equate to cognitive lobotomy. By shoehorning a 284B MoE into a 128GB consumer machine, this design—trading memory capacity for compute throughput—flattens the barrier of frontier inference from datacenter server racks directly to the desktop.
At the upper end of consumer silicon, a 512GB Mac Studio running V4 PRO achieves 150 t/s prefill and 10–13 t/s decode throughput—a performance envelope fully capable of powering deep, iterative analysis across complex codebases. Treating compute infrastructure as a one-time workstation capital expense of around $12,000 establishes a compelling financial baseline for repatriating large models from the cloud to local development environments. This hardware investment model completely eliminates the anxiety of unpredictable, token-metered cloud billing.
790 t/s Prefill Reveals the True Bottleneck
On an M5 Max with 128GB unified memory under a 2048-token context, ds4 clocked 790.2 t/s for prefill, dramatically cutting time-to-first-token. On the exact same machine, however, generation throughput dropped to 39.4 t/s, laying bare the physical memory bandwidth ceiling that throttles unified memory architectures during memory-bound decoding phases. This stark divergence across inference phases dictates that edge inference strategies must deliberately play to their strengths while mitigating inherent bandwidth constraints.
Comparing numbers between the M5 Max and DGX Spark side-by-side illustrates differing hardware priorities. In contrast to DGX Spark’s 825.8 / 18.1 t/s metrics, server-grade hardware maintains an advantage in raw prefill compute, but generation throughput shows virtually no divergence. Long-context workloads test these systems even more rigorously: at 65,536 tokens, the M5 Max posts 398.5 / 27.6 t/s. Robust prefill performance confirms that ballooning KV caches do not inevitably derail pipeline efficiency.
Under extensive contexts where prefill barely degrades while generation drops by half, the operating model for coding agents is redefined. Orchestration layers are compelled to avoid lengthy, continuous generation streams, leaning instead toward high-frequency inputs and short-turn interactions for debugging and editing. These underlying physical realities of hardware are steering developer tooling toward lightweight interactions grounded in rapid context reloading.
Streaming from SSD When Memory Falls Short
When memory limits are reached, forcing the entirety of model weights into physical RAM is no longer the only viable option. ds4’s architectural breakthrough pairs disk-backed KV caching with NVMe SSD weight streaming. Whenever parameter sizes exceed memory capacity, the engine parks inactive expert weights on the SSD, pulling them on demand to bypass physical RAM ceilings. The microsecond latency profile of modern NVMe drives paves the hardware runway for such dynamic streaming.
Offloading KV cache to disk supports restoration by prompt hash, eliminating the need to re-prefill prompts after restarts. Following any unexpected crash or session reset, loading cached states straight from NVMe bypasses long redundant recomputations. These low-level systems details often matter far more than model choice in deciding whether a workflow remains viable on resource-constrained devices.
Figure: Prefill and generation throughput curves on M5 Max from the ds4 repository speed-bench. Source: antirez/ds4 repo speed-bench/m5_max_ts.svg
For native coding agents, inference execution is tightly managed within a dedicated local process. Traditional cloud API latency, packet loss, and serialization overhead are eliminated. Treating high-speed NVMe storage as a direct extension of memory challenges legacy assumptions that once governed LLM runtime design.
Stitching Two 128GB Machines Together via RDMA
The physical boundaries of standalone hardware can be expanded horizontally across local networks. ds4 supports tensor parallelism across multiple nodes using Apple RDMA, creating a unified, heterogeneous memory pool across devices. By bypassing kernel syscall overhead in conventional network stacks, this communication layer reduces tensor-slice interchange latency to practical working thresholds.
For clustered deployments, an 8× L40S configuration delivers 126 t/s aggregate generation throughput—ample capacity to serve the concurrent interactive requests of an entire small engineering team. Older accelerator cards sidelined by official frameworks find second lives as usable compute nodes, extracting lingering utility via multi-tenant inference server configurations. Distributed collaboration fundamentally hinges on the surgical slicing and scheduling of memory and compute.
Figure: Generation throughput comparison across Qwen3.8 checkpoints from the ds4 repo speed-bench. Source: antirez/ds4 repo speed-bench/qwen38-checkpoints/generation-throughput.svg
The aggregate throughput gains unlocked by multi-user concurrency essentially trade single-request latency for maximized memory bus utilization. Whether using tensor slicing or pipeline parallelism, the objective is squeezing every drop of throughput from available silicon. This creates a viable operational space for idle, previous-generation accelerators well outside hyperscale cloud clusters.
A Narrow Implementation Deliberately Shunning Generic Ecosystems
ds4 explicitly rejects becoming a generic GGUF runner, recognizing only its own bespoke, minimalist GGUF layout. This focused implementation strips away the branching control-flow overhead required to support fragmented quantization schemes, reserving scarce on-chip cache lines for core compute loops. The absence of release tags reinforces its identity as a fast-iterating research vehicle rather than an enterprise runtime.
Project commit logs openly reveal deep collaboration with AI coding agents during development. AI assistants participated directly in refactoring their own underlying inference primitives, significantly compressing the low-level tuning cycles required for target hardware architectures. Released under the MIT license, this lightweight implementation can be integrated cleanly into proprietary software stacks without friction.
Yet the cost of such a narrow scope is equally undeniable: the engine’s long-term utility remains tied to the few labs willing to release weights under permissive open licenses. Should top-tier frontier model builders close their open-weight releases, the returns on such bespoke optimization could evaporate quickly. While it extracts peak performance from targeted hardware, it leaves its technological roadmap reliant on the open ecosystem’s continued vitality.
By trading away broad generality for extreme hardware optimization, ds4 redraws the accessibility map for frontier AI. It proves that when developers are willing to strip away abstraction layers on targeted silicon, the barrier to local frontier inference is no longer raw compute horsepower, but memory and storage bandwidth. Today, running high-parameter frontier models has arrived on the desktop; the deciding factor is simply the physical memory limit of your machine.
Reference links:
- HN Discussion Archive
- antirez Official Benchmark Report