Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
ai
On Tuesday, a developer on Hacker News introduced Slotstream, a tool that runs a 125-billion-parameter AI model on a Mac with as little as 16 gigabytes of memory. The model, Qwen 3.8 Flash Next, would need over 100 gigabytes, but Slotstream uses expert offloading and SSD streaming to fit it into less. The developer reports speeds around 12 tokens per second and plans to add speculative decoding next.
Source: https://github.com/carloslfu/slotstream
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton