Posted by stared 20 hours ago
/data/llm/llama.cpp/build/bin/llama-server
--threads 4
--threads-batch 8
--batch-size 4096
--ubatch-size 256
--port 9999
--temp "1.0"
--top-p "0.95"
--top-k "20"
--min-p "0.0"
--presence-penalty "0.0"
--reasoning auto
--reasoning-preserve
--reasoning-budget 4096
--gpu-layers-draft all
--spec-type draft-mtp,ngram-map-k4v,ngram-mod
--spec-draft-n-max 3
--spec-draft-p-min 0.75
--spec-ngram-mod-n-match 24
--spec-ngram-mod-n-min 4
--spec-ngram-mod-n-max 16
--spec-ngram-map-k4v-size-n 8
--spec-ngram-map-k4v-size-m 16
--spec-ngram-map-k4v-min-hits 1
--n-gpu-layers all
--ctx-size 131072
--repeat-penalty 1.0
--jinja
--metrics
--model /data/llm/models/unsloth/Qwen3.8-27B-UD-IQ3_S.gguf
--chat-template-file /data/llm/models/qwen3.6-chat-template.jinja
--fit off
--flash-attn on
--cors-origins localhost
--mmproj /data/llm/models/unsloth/Qwen3.8/mmproj-BF16.gguf
--no-mmproj-offload
--parallel 1
--kv-unified
--cache-type-k q4_0
--cache-type-v q4_0
--cache-type-k-draft q4_0
--cache-type-v-draft q4_0The chat template is froggeric's fixed qwen template, v22.5 as of today.
I use a combination of a Claude Max subscription and local inference, including qwen3.8-27b, 4bit. I have found qwen to be absolutely useless at anything but very specific, surgical code changes. In my experience, for anything even remotely nuanced, a frontier model is required.
Index methodologies here: https://artificialanalysis.ai/evaluations/artificial-analysi...
Also see some specific benchmarks here: https://huggingface.co/Qwen/Qwen3.8-27B e.g. qwen scores 61.7 on swe bench pro, while opus 4.6 scores 53.4.
If you want to argue with the benchmarks, go for it. Fwiw i am not saying qwen3.8-27b is better or as good as the frontier. But i am saying it has crossed the threshold and is now a useful tool for coding and debugging. From my experience, Qwen3.6-35b-a3b was what you describe - it could do surgical edits only.
I think you're referring to a 5060Ti 16GB, yes?
32k context is easily done there. 64k can work with a more aggressive quant, but you lose a bit of speed.
But I don’t quite follow you - how does a more aggressive quant slow it down? Less bits per token means faster inference not slower.
qwen3.8-27b 4bit has a following specifically for being exceptionally gifted for such a small model.
It does clear the useful enough to be worth it bar though.
Zero interest in remote models but local ones if they offer utility, sure.
Runs pretty well on a 7900XTX as well.