Mac Studio

Mac Studio · M3 Ultra · 96 GB

Fast is not the same
as finished

Joshua Goh

Running Qwen3.8 Flash-Next agents on oMLX, Unsloth and antirez's ds4: what each engine did with a 30-minute job and a 5-hour one.

Setup · how a request flows

Any device sends a task, one engine serves it, weights stream from the NVMe

Mac StudioM3 Ultra · 96 GB · runs everything MacBook Airjoins over Tailscale task tailnet only DeepSeek Harness coding agent bash · read_image · edit · write prompt + tool results one engine per session oMLXMLX · Lightning MTP · 64 t/s Unslothllama-server · GGUF · ~27 t/s on short runs ds4 by antireznative C · SSD streaming NVMe 1.1 TB of models weights Models never touch the Mac Studio's internal SSD: streaming experts and PLE pages wears a replaceable external drive.

The two tests · run first on the dense 27B, then on the Flash-Next MoE

Poster: under 2 hours. OMR build: 5 to 12 hours.

QR poster from oMLXQR poster from Unsloth

oMLX, 11 min · Unsloth, 18 min

Short

QR poster

Design one scannable poster with the school logo.

OMR report output

Long

OMR software

Read scanned answer sheets for seven classes and reproduce the school's report.

Part 1 · Qwen3.8 27B, dense · poster test

Every 27B variant finished the poster, on oMLX and on Unsloth

27B MTPLX-Opt-Speed poster

27B MTPLX-Opt-Speed

27B 4bit poster

27B 4bit

27B oQ4e-mtp poster

27B oQ4e-mtp

Dirk 27B oQ4e poster

Dirk 27B oQ4e

Dirk 27B oQ4e poster

Dirk 27B oQ4e

Dirk 27B oQ4e poster

Dirk 27B oQ4e

Qwopus 27B Flash oQ4 poster

Qwopus 27B Flash oQ4

27B oQ4e-mtp poster

27B oQ4e-mtp

27B GSQ IQ3_S · Unsloth poster

27B GSQ IQ3_S · Unsloth

Part 1 · Qwen3.8 27B, dense · poster test

27B quants decode the poster at 22 to 40 tokens per second. Flash-Next with MTP decodes at 64.

oMLXUnslothdecode speed excludes prefill and tool time

Part 1 · Qwen3.8 27B, dense · OMR test

The dense 27B never met the memory guard. It hit the 200k context cap, then finished after a handoff.

0

memory guard events in five 27B OMR sessions

526k

output tokens in the longest session, 438 steps

11 t/s

effective speed at 200k context, about half its poster speed

Part 2 · Qwen3.8 Flash-Next, MoE · poster test on oMLX and Unsloth

On Flash-Next, oMLX with MTP finished in 11 minutes. Unsloth took 18, expert offload 34.

Run A: 75% of experts resident, MTP off. Run B: all experts resident, PLE on SSD, MTP off. Run C: as B with Lightning MTP on. Run D: Unsloth llama-server, UD-IQ4_XS GGUF, peak memory not logged. Same prompt each time.

Part 2 · Qwen3.8 Flash-Next, MoE · poster test on oMLX and Unsloth

All four runs produced a scannable poster

Poster from run A

A · oMLX, expert offload 75%
34 min · saved just before a GPU error

Poster from run B

B · oMLX, full residency
31 min · completed

Poster from run C

C · oMLX, full residency + MTP
11 min · completed

Poster from run D

D · Unsloth llama-server
18 min · completed

Part 2 · Qwen3.8 Flash-Next, MoE · OMR test on oMLX

At 200k tokens the memory guard ended both turns and blocked every recovery

Turn 1 · "Prefill would require ~81.66 GB peak … but dynamic ceiling is 81.58 GB. Only 6.69 GB is reclaimable right now."

Turn 2 · "Request aborted: process memory limit exceeded (usage 85.3 GB, abort threshold 83.6 GB, dynamic ceiling 88.0 GB)."

Part 2 · Qwen3.8 Flash-Next, MoE · why oMLX stops

A 68 GB model plus a 200k-token context crosses the abort line on 96 GB

Raising the ceiling does not create memory. It went 88 → 90 → 91.2 GB, and the guard was then switched off. Prefill still peaked above 84 GB.

Expert offload would fit but was not usable. At 50% residency the model is 37 GB, but that disables MTP, halves decode, and run A shows it can still crash the GPU.

Part 3 · the other engines on the long test · ds4 by antirez

ds4 never ran out of memory. It ran out of images.

147 GB

largest checkpoint run on this 96 GB Mac, streamed from SSD by ds4

16

images per request its vision server accepts. The OMR agent keeps every scanned sheet in context and passes that in under an hour.

0

memory guard events across every ds4 session

Part 3 · the other engines on the long test · Unsloth llama-server

Unsloth decoded at a third of oMLX's speed and was the only engine to finish

278 min

one turn, 218 steps, ended "completed"

0

aborts. 20 prompt-cache evictions cost a re-prefill each, never the run

232

requests, 984k prompt tokens processed

168k

context high-water mark inside a 200k slot

Part 3 · Unsloth run · output

From scanned sheet to a report that matches the school's reference

Scanned answer sheet, student name blurred

Input · scanned answer sheet

→
Generated report page

Generated · out/Class Report.pdf

Reference report page

Reference · supplied by the school

7

class reports plus a results JSON for each class

14

Python modules in the omr package, including an agent-in-the-loop verify server

Part 3 · Unsloth run · output

The report's bar charts, rebuilt from the scans, match the reference

Generated

Generated item analysis chart Generated grade distribution chart

Reference

Reference item analysis chart Reference grade distribution chart

Part 4 · every Flash-Next OMR attempt, grouped by engine

Eleven sessions. Only the two Unsloth runs finished with usable output.

Takeaways

Choose the engine by context length, not by tokens per second

Next: retry the 200k run on oMLX with expert offload at 25 to 50%, and try Unsloth with its MTP draft GGUF to close the decode gap.

←→ navigate · F fullscreen ·