Chat
27BQwen 3.8 27B AXQ 6-bit MTP — `qwen3.8-27b:axq`
MAC-FIRST INFERENCE RUNTIME
Apple Silicon inference with speculative MTP decoding. One AXQ 6-bit pack, native Metal graph, OpenAI-compatible server on 127.0.0.1:31418.
Open source · Apache-2.0
Install in one command
brew install defai-digital/tap/ax-engineAX ENGINE / Models
AX Engine supports Qwen 3.5 / 3.6 / 3.8, Gemma 4, GLM, Nemotron, Whisper, and embedding models through a direct-first runtime contract. The qualification pack is Qwen 3.8 27B AXQ 6-bit MTP; other families stay supported but are not the first-run target.
Chat
27BQwen 3.8 27B AXQ 6-bit MTP — `qwen3.8-27b:axq`
Coding agent
MoEOrnith 1.5 35B-A3B coding agent
Vision chat
VLQwen3-VL 30B-A3B AXQ
AX ENGINE / Performance
AX Engine leads the loadable MTP peers on the Qwen 3.8 27B AXQ 6-bit MTP pack: MTPLX (draft depth 3), OMLX (Lightning depth 1), and the direct-AR `mlx-lm` baseline. Numbers are 20-run medians on the same checkpoint.
MTP Tier 2 is still pending. Decode moves past the DRAM ceiling by emitting several tokens per weight pass; the bandwidth ceiling on the qualification SKU is reached at ~266 GB/s of weight-stream bandwidth.
Full peer comparison, methodology, and Tiel / Cyber-Tiel MXFP4 MTP rowsAX Engine 7.5.x ships 6-bit AXQ MTP for Qwen 3.8 27B (SKU: Mac mini M4 Pro 64 GB) and Tiel/Cyber-Tiel MXFP4 MTP packs (campaign: M5 Max 128 GB).
See the roadmap →Dense 27B, MXFP4 MoE, native MTP, OpenAI-compatible endpoints. Run it on the machine you already own.
Install AX Engine