Underlying it all is that any architecture can be trained to the same convergence just difference in compute utilization both in training and inference
I suffix everything with "Reply in English", and even so I‘m getting lots of Chinese.
seriously
Oh interesting, I can assume what the benefits is for including the Encoder, but whats the downside? I’m thinking GPT (which is decoder only) ruled out Encoder for a reason?
In terms of coding and command line capabilities I'm also very interested to see a head-to-head of it vs. qwen 3.8-flash-next Q8 which is something like 190GB of memory used when loaded into llama-server. It fits very well in all sorts of 256GB or under class machines.
In Pi (pi.dev), it tells me it's definitely Claude by Anthropic, via the API via curl it tells me it's "probably ChatGPT", its very funny.
Thinking: > The user is asking what model I am. According to my system prompt, I'm powered by "deepseek-flash" with model ID "opencode-go/deepseek-flash".
>I'm powered by the model opencode-go/deepseek-flash.
I was hoping for a bit more, but it's still 100% faster for a very good price, so I won't complain.