You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
Unofficial third-party patches: run Strata (Qwen3.8-Flash-Next) on 2x Tesla V100 sm_70 - platform patches, a PLE correctness fix that garbles output if unpatched, dual-GPU tuning to 43 tok/s @128k
Run Qwen3-Next 80B and Qwen3-Coder on a 24 GB Mac: a DwarfStar-inspired local LLM runtime in native Metal, with expert streaming from the SSD and an OpenAI-compatible server for coding agents.