Unsloth publishes GGUF quantizations of Qwen3.8-27B that run through Ollama and llama.cpp. I tested the UD-Q5_K_XL build on an M4 Max with 64GB of unified memory. After the first download, I could run inference locally without sending prompts to a hosted model API.

About the Model

The model card describes Qwen3.8-27B as a 27B vision-language model with configurable thinking and improved tool use. Unsloth reports that its Dynamic V3.0 quantization improves accuracy at the same file size compared with the providers it tested.

Runtime support varies, so confirm that your chosen Ollama or llama.cpp build exposes the capability you need.

Install with Ollama

Install and open Ollama:

brew install --cask ollama

Then run a quant that fits your machine:

ollama run hf.co/unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL

Ollama downloads the model on the first run and reuses the cached model file later.

Install with llama.cpp

brew install llama.cpp
llama-cli -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL

Choose a Quant

The highlighted GGUF files are about 8.4GB for UD-IQ2_S, 14.3GB for UD-IQ4_XS, 17.6GB for UD-Q4_K_XL, and 20.9GB for UD-Q5_K_XL. The runtime needs additional memory. Choose a file that leaves headroom, then test it on your machine.

For my 64GB M4 Max test, I used:

ollama run hf.co/unsloth/Qwen3.8-27B-GGUF:UD-Q5_K_XL

Demo

Sources