Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Qwen 3.8 27B defaults to extreme overthinking, consuming 22K tokens for simple tasks
“This is a hilarious default. It's absolutely not a good way to run the model, especially on consumer hardware.”
Alibaba's Qwen 3.8 27B is a capable Apache 2 licensed vision LLM that outperforms its predecessors on benchmarks, but ships with a default reasoning effort of 'xhigh' that causes severe overthinking—one SVG generation took 21 minutes and 22K reasoning tokens. The model is promising for local inference at 17GB, but the default configuration makes it impractical without tuning reasoning effort down.