Mac Studio · M3 Ultra · 96 GB
Joshua Goh
Running Qwen3.8 Flash-Next agents on oMLX, Unsloth and antirez's ds4: what each engine did with a 30-minute job and a 5-hour one.
Setup · how a request flows
The two tests · run first on the dense 27B, then on the Flash-Next MoE
oMLX, 11 min · Unsloth, 18 min
Short
Design one scannable poster with the school logo.
Long
Read scanned answer sheets for seven classes and reproduce the school's report.
Part 1 · Qwen3.8 27B, dense · poster test
27B MTPLX-Opt-Speed
27B 4bit
27B oQ4e-mtp
Dirk 27B oQ4e
Dirk 27B oQ4e
Dirk 27B oQ4e
Qwopus 27B Flash oQ4
27B oQ4e-mtp
27B GSQ IQ3_S · Unsloth
Part 1 · Qwen3.8 27B, dense · poster test
Part 1 · Qwen3.8 27B, dense · OMR test
memory guard events in five 27B OMR sessions
output tokens in the longest session, 438 steps
effective speed at 200k context, about half its poster speed
Part 2 · Qwen3.8 Flash-Next, MoE · poster test on oMLX and Unsloth
Run A: 75% of experts resident, MTP off. Run B: all experts resident, PLE on SSD, MTP off. Run C: as B with Lightning MTP on. Run D: Unsloth llama-server, UD-IQ4_XS GGUF, peak memory not logged. Same prompt each time.
Part 2 · Qwen3.8 Flash-Next, MoE · poster test on oMLX and Unsloth
A · oMLX, expert offload 75%
34 min · saved just before a GPU error
B · oMLX, full residency
31 min · completed
C · oMLX, full residency + MTP
11 min · completed
D · Unsloth llama-server
18 min · completed
Part 2 · Qwen3.8 Flash-Next, MoE · OMR test on oMLX
Turn 1 · "Prefill would require ~81.66 GB peak … but dynamic ceiling is 81.58 GB. Only 6.69 GB is reclaimable right now."
Turn 2 · "Request aborted: process memory limit exceeded (usage 85.3 GB, abort threshold 83.6 GB, dynamic ceiling 88.0 GB)."
Part 2 · Qwen3.8 Flash-Next, MoE · why oMLX stops
Raising the ceiling does not create memory. It went 88 → 90 → 91.2 GB, and the guard was then switched off. Prefill still peaked above 84 GB.
Expert offload would fit but was not usable. At 50% residency the model is 37 GB, but that disables MTP, halves decode, and run A shows it can still crash the GPU.
Part 3 · the other engines on the long test · ds4 by antirez
largest checkpoint run on this 96 GB Mac, streamed from SSD by ds4
images per request its vision server accepts. The OMR agent keeps every scanned sheet in context and passes that in under an hour.
memory guard events across every ds4 session
Part 3 · the other engines on the long test · Unsloth llama-server
one turn, 218 steps, ended "completed"
aborts. 20 prompt-cache evictions cost a re-prefill each, never the run
requests, 984k prompt tokens processed
context high-water mark inside a 200k slot
Part 3 · Unsloth run · output
Input · scanned answer sheet
Generated · out/Class Report.pdf
Reference · supplied by the school
class reports plus a results JSON for each class
Python modules in the omr package, including an agent-in-the-loop verify server
Part 3 · Unsloth run · output
Generated
Reference
Part 4 · every Flash-Next OMR attempt, grouped by engine
Takeaways
Next: retry the 200k run on oMLX with expert offload at 25 to 50%, and try Unsloth with its MTP draft GGUF to close the decode gap.