An important factor is that fine tuning existing open models is incredible cheap. You can easily change any cultural biases if you want a model to be 'sovereign'. And Mistral could combine that with their custom data sets for their enterprise customer needs. Mistral still trains their own models, but they also seem to offer fine tuning existing models.
With model weights being commoditized, another differentiator could be deploying efficient inference chips, especially if you combine it with a developer ecosystem for vendor lock-in. That is why it is interesting that both Samsung and ASML are investors, since they are companies that could make a difference in this area.
It may make sense to train a frontier model on an existing architecture if the base model is not available and the instruction trained version doesn't fit with what you want. There are techniques like ablation, but those could have other effects on the model, and there can still be lingering effects of the instruction training in the model that surface less frequently (e.g. on an input not covered by the ablation training).
Otherwise, fine tuning is definitely the way to go. However, you need to be careful not to over-tune the model such that it is only tuned to the data you are training it on.
Seriously "thought to be" is such a baseless statement. Thought to be by whom? And on what basis?
You might as well pretend there's no money in search or social media.
1999. That Google, it has no serious business model, it'll never make real money. Hello GoTo.
2006. That Facebook, it has no serious business model, it'll never make real money.
One more: 2020. That Uber, I could clone an MVP and kill them over the weekend if I really wanted to. That's a side project at best. They're bleeding to death anyway, losing money hand over fist, they'll never make it. They have no moat at all.
The bubble is going to burst soon.
Perhaps Mistral has too high moral standards to keep up?
They would if they could, which means they can‘t, even though they want to.