AX Engine 7.5.x ships dense 27B past the DRAM ceiling.See the numbers
AX Engine.

APPLE SILICON · LOCAL INFERENCE

Past the DRAM ceiling.On your Mac.

Local inference on your terms.

MAC-FIRST INFERENCE RUNTIME

Dense 27Bpast theDRAM ceiling.

Apple Silicon inference with speculative MTP decoding. One AXQ 6-bit pack, native Metal graph, OpenAI-compatible server on 127.0.0.1:31418.

Open source · Apache-2.0

Install in one command

brew install defai-digital/tap/ax-engine

macOS 26+ Apple SiliconTerminalInstallation guide

macOSmacOS 26+ · Apple Silicon (M2 or newer)5 supported model families76.90 tok/s decode · M5 Max 128 GBApache-2.0 open source

AX ENGINE / Models

Supported models.

AX Engine supports Qwen 3.5 / 3.6 / 3.8, Gemma 4, GLM, Nemotron, Whisper, and embedding models through a direct-first runtime contract. The qualification pack is Qwen 3.8 27B AXQ 6-bit MTP; other families stay supported but are not the first-run target.

Chat

27B

Qwen 3.8 27B AXQ 6-bit MTP — `qwen3.8-27b:axq`

Coding agent

MoE

Ornith 1.5 35B-A3B coding agent

Vision chat

VL

Qwen3-VL 30B-A3B AXQ

Full model matrix, certifications, and MTP pack notes

AX ENGINE / Performance

Performance.

AX Engine leads the loadable MTP peers on the Qwen 3.8 27B AXQ 6-bit MTP pack: MTPLX (draft depth 3), OMLX (Lightning depth 1), and the direct-AR `mlx-lm` baseline. Numbers are 20-run medians on the same checkpoint.

MTP Tier 2 is still pending. Decode moves past the DRAM ceiling by emitting several tokens per weight pass; the bandwidth ceiling on the qualification SKU is reached at ~266 GB/s of weight-stream bandwidth.

Full peer comparison, methodology, and Tiel / Cyber-Tiel MXFP4 MTP rows
On the horizon

AX Engine 7.5.x ships 6-bit AXQ MTP for Qwen 3.8 27B (SKU: Mac mini M4 Pro 64 GB) and Tiel/Cyber-Tiel MXFP4 MTP packs (campaign: M5 Max 128 GB).

See the roadmap →

Your inference on your Mac.

Dense 27B, MXFP4 MoE, native MTP, OpenAI-compatible endpoints. Run it on the machine you already own.

Install AX Engine