Compact native inference engine "DwarfStar 4" designed exclusively for DeepSeek V4 Flash

TL;DR AI
2 min readKey summary
Salvatore Sanfilippo has open-sourced DwarfStar 4, a compact inference engine tuned specifically for DeepSeek V4 Flash.
The engine uses asymmetric 2-bit/8-bit quantization and llama.cpp-style optimizations to run large contexts efficiently on local machines.
Reports say Macs with around 128GB of RAM can handle substantial workloads, drawing strong interest from Hacker News and researchers.
The release highlights the practicality of model-specific local inference and expands options for high-performance AI without heavy cloud dependence.



