If you have been following our local-AI hardware series - the Mac mini deep dive, the Mac Studio deep dive, and the NVIDIA DGX Spark deep dive - this is the fourth and most price-aggressive chapter. AMD’s Ryzen AI Max family, codenamed Strix Halo, puts up to 128GB of unified LPDDR5X memory and a 40-compute-unit Radeon 8060S GPU in mini PCs that start well under the price of a DGX Spark or a fully loaded Mac Studio. The question, as always: how fast does it actually run large language models, and how big can those models be?
Every number below carries a link to its source. We read AMD’s official spec sheets, AMD’s own 2026 benchmark blog, the community Strix Halo benchmark repos, Framework’s published machine-learning figures, and storefront pricing pages, all crawled in September 2026.
What “Ryzen AI Max” Actually Is
Strix Halo comes in three main bins. The flagship is the AMD Ryzen AI Max+ 395: 16 Zen 5 cores / 32 threads up to 5.1 GHz, Radeon 8060S graphics with 40 RDNA 3.5 compute units at up to 2900 MHz, a 50-TOPS XDNA 2 NPU (up to 126 total TOPS), and 256-bit LPDDR5X-8000 memory, 128GB maximum, with a configurable TDP of 45-120W. All of these figures come straight from AMD’s official specification page.
The step-down Ryzen AI Max 385 pairs the same memory subsystem with 8 cores and 32 graphics cores, which is the configuration Framework sells as its 32GB entry tier at $1,269. AMD also sells its own reference box, the Ryzen AI Halo developer platform, a 150 x 150 x 45.4 mm machine with 128GB at 8000 MT/s, 256 GB/s of memory bandwidth, 10GbE, and a Linux-first ROCm software stack.
| Spec | Value | Source |
|---|---|---|
| CPU | 16x Zen 5, 32 threads, up to 5.1 GHz | AMD specs |
| GPU | Radeon 8060S, 40 CU RDNA 3.5 @ 2900 MHz | AMD specs |
| NPU | 50 TOPS XDNA 2 (126 TOPS total) | AMD specs |
| Memory | 256-bit LPDDR5X-8000, max 128GB | AMD specs |
| Memory bandwidth | 256 GB/s | AMD Halo page |
| TDP | 45-120W configurable | AMD specs |
| VRAM allocation | Up to 96GB on Windows, ~115GB GTT pool on Linux | Framework, seehiong blog |
The 256 GB/s Rule
Just like the DGX Spark’s 273 GB/s (NVIDIA), Strix Halo’s decode speed is governed by memory bandwidth. The theoretical 256 GB/s works out to roughly 215 GB/s measured in real inference runs - about a 16 percent gap - according to the community benchmark repository visorcraft/strix-halo-llm-perf.
The consequence is identical to every other bandwidth-limited box in this series: mixture-of-experts models with few active parameters fly, dense models crawl. A 30B MoE with 3B active parameters decodes at 86-100 tok/s, while a dense 70B at the same quantization manages 5.1 tok/s - the whole dense weight matrix streams from memory for every single token.
| Model type | Decode speed | Source |
|---|---|---|
| MoE 30B-A3B (3B active) | 86-100 tok/s | visorcraft, RunAIHome |
| MoE 120B (5.1B active) | 53-56 tok/s | visorcraft |
| Dense 32B Q4 | ~10 tok/s | llmrun |
| Dense 70B Q4_K_M | 5.1 tok/s | RunAIHome |
Decode Throughput: The Community Numbers
The single best public resource for this chip is the strix-halo-llm-perf GitHub repository, which benchmarks llama.cpp across GMKtec EVO-X2 and Beelink GTR9 Pro machines. Combined with Framework’s official LM Studio figures, llmrun’s device page, RunAIHome’s 2026 analysis, and llamaperf’s Strix Halo 128GB page, a consistent picture emerges on a 128GB box:
- gpt-oss-20b: 58 tok/s in Framework’s official LM Studio testing; the 8B-class LFM2.5 reaches 152.5 tok/s.
- Qwen3-30B-A3B: 86.1 tok/s Vulkan (visorcraft) and up to 100.04 tok/s at IQ4_XS on RADV drivers (RunAIHome). Qwen3-Coder-30B hits 96.8-98.5 tok/s.
- gpt-oss-120b: the headline act. Its MXFP4 weights occupy 63.39 GB, so it only fits on 128GB models. llama.cpp Vulkan runs deliver 53.4-55.57 tok/s; Framework’s official figure is 38 tok/s; GMKtec’s Ollama build manages 19.25 tok/s. Your backend choice matters more than your brand of mini PC.
- Qwen3-Coder-Next 80B-A3B: 42.7 tok/s. Qwen3.6 35B Q6: 50 tok/s.
- Qwen3.8 Flash-Next (~125B MoE): 38-45 tok/s with multi-token prediction.
- GLM-5.3-Flash: 14.63 tok/s on ROCm vs 8.57 on Vulkan - the ROCm/Vulkan split flips depending on the model.
- Frontier-adjacent MoE on one box: MiniMax M2.5 (228.7B, Q3) at 32.8 tok/s; Qwen3-235B-A22B Q3 at 17.2 tok/s; Llama 4 Scout 109B at 13.83 tok/s; Heretic2 at IQ4_XS in a 92GB footprint runs 24 tok/s.
- MLPerf Client v6.1 on Strix Halo 128GB decodes Atlas NVFP4 at 20.6 tok/s, and DeepSeek V4.1 Flash Q2 (streamed from SSD) decodes at 7.12 tok/s.
- Dense big models: Llama 3.1 70B Q4_K_M at 5.1 tok/s; dense 27-32B quantized models sit around 9.6-11.3 tok/s.
AMD’s own January 2026 blog frames the same story marketing-first: gpt-oss-120b runs roughly 10x faster than Llama 3 70B on the AI Max+ 395, and the platform delivers about 1.7x the tokens-per-dollar of DGX Spark across four models tested in LM Studio.
Prompt Processing: ROCm vs Vulkan
Backends matter on this chip. The strix-halo-llm-perf maintainers found that ROCm 7.x builds give the best prompt processing, while AMDVLK Vulkan builds give the best token generation (+16 percent in some tests), and the kyuz0 amd-strix-halo-toolboxes project packages prebuilt toolchains so you can A/B test both against your own workloads. On Linux, Ubuntu 24.04 with recent ROCm builds unlocks a ~115GB GTT memory pool for the GPU, versus the 96GB VRAM slider on Windows (seehiong blog, Framework).
Model Map by Cluster Size
| Setup | Biggest practical models | Speed | Source |
|---|---|---|---|
| 32GB (Max 385) | 8-14B dense, gpt-oss-20b Q4 | fast small models | Framework |
| 64GB (Max+ 395) | Qwen3-30B-A3B, Qwen3-Coder-30B, dense 32B Q4 | 86-100 / ~97 / ~10 tok/s | llmrun |
| 128GB (Max+ 395) | gpt-oss-120b, Qwen3.8 Flash-Next, MiniMax M2.5 Q3, Llama 4 Scout 109B, dense 70B Q4 | 53-56 down to 5.1 tok/s | visorcraft |
| 2 boxes (USB4/10GbE RPC) | MiniMax M2.5-REAP 228.7B, Qwen3.5-397B | 15.35 / ~12 tok/s | visorcraft |
| Strix Halo + Mac M5 Max | GLM-5.3 321B | 24 tok/s | llamaperf |
Note the 64GB catch: gpt-oss-120b’s 63.39 GB of MXFP4 weights cannot squeeze into a 64GB machine’s VRAM allocation. If 120B-class models are the goal, 128GB is the floor.
Clustering: USB4, 10GbE, and Mixed Mac+AMD
llama.cpp’s RPC mode works fine over Strix Halo’s USB4 ports, measuring about 9.4 Gbps of effective throughput. Two hosts split MiniMax M2.5-REAP 228.7B at 15.35 tok/s and Qwen3.5-397B at roughly 12 tok/s. Because RPC is vendor-agnostic, the community also runs hybrid clusters: a Strix Halo box plus a Mac Studio M5 Max ran GLM-5.3 321B at 24 tok/s - the same mixed-vendor trick we covered in the Mac Studio deep dive.
For rack ambitions, the Minisforum MS-S1 MAX ships with dual 10GbE, a PCIe 4.0 x16 slot, and explicit dual-unit-to-2U-rack cluster support. Framework documents llama.cpp RPC clustering as a supported workflow for its Desktop, and has announced a 192GB memory tier is coming.
Fine-Tuning
Fine-tuning on Strix Halo is a Linux-first affair: AMD’s ROCm 7.x stack supports gfx1151, and the Ryzen AI Halo developer platform ships ROCm preinstalled with a 120W power envelope, while the strix-halo-llm-perf and toolboxes projects document the ROCm builds that work. With 96-115GB of GPU-visible memory, QLoRA on 30-70B models and LoRA on small MoE models fit where a 16GB consumer GPU cannot. There is no CUDA ecosystem here - expect to debug - but for inference-first buyers that is a non-issue.
Ryzen AI Max vs DGX Spark vs Mac Studio
| Platform | Memory | Bandwidth | gpt-oss-120b decode | Price (128GB class) |
|---|---|---|---|---|
| Ryzen AI Max+ 395 | 128GB LPDDR5X | 256 GB/s | 38-56 tok/s | from $3,449 |
| DGX Spark (GB10) | 128GB LPDDR5X | 273 GB/s | ~55-60 tok/s | $3,999+ |
| Mac Studio M5 Max | up to 128GB | far higher | see our deep dive | from ~$4,999+ |
Sources: AMD Halo page, NVIDIA, visorcraft, Framework. The bandwidth gap to Apple silicon (M5 Max/Ultra reach 460-1,228 GB/s, documented in our Mac Studio post) is real - but so is the price gap. AMD’s own blog claims 1.7x tokens-per-dollar over DGX Spark, and the community numbers above back the direction of that claim: you trade roughly 6 percent bandwidth and a CUDA license for hundreds of dollars and full x86 tooling.
Where to Buy (2026 Prices)
- Framework Desktop: DIY Max 385 32GB $1,269 (currently out of stock), Max+ 395 64GB $1,959, Max+ 395 128GB $3,449. Mini-ITX mainboard, repairable, 5GbE plus dual USB4.
- GMKtec EVO-X2: $2,199 (64GB) / $3,649 (128GB), re-verified September 2026 amid the DRAM shortage.
- BOSGAME M5 AI: 128GB/2TB at $3,499.99 on Newegg.
- Minisforum MS-S1 MAX: 128GB “Max AI Compute” edition at $3,799 (regular $4,749), with dual 10GbE and PCIe x16.
- AMD Ryzen AI Halo: official developer platform, $4,699.99 at Newegg, Linux + ROCm preinstalled.
Conclusion
The Ryzen AI Max family is the value play of the 2026 local-AI scene. For $3,449-$3,649 you get a 128GB box that runs gpt-oss-120b at 38-56 tok/s, Qwen3-30B MoE at ~100 tok/s, and clusters into two-box MiniMax/Qwen3.5-397B territory or hybrid Mac+AMD rigs running GLM-5.3 321B. It loses the bandwidth race to Apple silicon and the software race to NVIDIA, but per dollar of decode throughput on big-memory MoE models, it is the cheapest door into frontier-adjacent local AI today. Pair the purchase decision with the DGX Spark analysis in our Spark deep dive and pick by ecosystem, not spec sheets.
All Sources
- AMD Ryzen AI Max+ 395 official specifications
- AMD Ryzen AI Halo developer platform
- AMD blog: Ryzen AI Max AI PCs deliver exceptional intelligence (Jan 2026)
- visorcraft/strix-halo-llm-perf community benchmark repo
- kyuz0/amd-strix-halo-toolboxes
- llamaperf: Strix Halo 128GB
- RunAIHome: Ryzen AI Max 395 Strix Halo local LLM analysis
- llmrun: Framework Desktop 128GB device page
- Framework Desktop store and pricing
- Framework Desktop machine learning benchmarks
- Framework community: running gpt-oss-120b on AI Max 395 128GB
- GMKtec EVO-X2 product page
- Minisforum MS-S1 MAX product page
- seehiong: running llama.cpp on AMD Strix Halo (Linux GTT)
- Newegg: AMD Ryzen AI Halo developer platform
- Newegg: BOSGAME M5 AI 128GB
- NVIDIA DGX Spark official page
- PyShine: Best Mac Studio for local AI deep research
- PyShine: Best DGX Spark for local AI deep research
- PyShine: Best Mac mini for local AI deep research
Enjoyed this post? Never miss out on future posts by following us