The Hallway Track
Engineering Insights

Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things

Simon Willison · Aug 16, 2026 · Engineering Insights

Qwen 3.8 27B defaults to extreme overthinking, consuming 22K tokens for simple tasks

“This is a hilarious default. It's absolutely not a good way to run the model, especially on consumer hardware.”

Alibaba's Qwen 3.8 27B is a capable Apache 2 licensed vision LLM that outperforms its predecessors on benchmarks, but ships with a default reasoning effort of 'xhigh' that causes severe overthinking—one SVG generation took 21 minutes and 22K reasoning tokens. The model is promising for local inference at 17GB, but the default configuration makes it impractical without tuning reasoning effort down.

open-source-models local-inference reasoning-models qwen alibaba on-device-ai

Watch / read the original source →