I had a great experience with llama-cpp with Nvidia backend on NixOS.
(Sorry for being that guy.)
https://github.com/cptskippy/battlemage-llm-gateway
It's designed so that you can re-run the scripts to pull the latest updates. When Muse Glimmer was released the other day I just ran the 02 script to build the latest version of llama.cpp with support for it.
Must be tough not to be able to monitor your own models!
(The odds that that tagline was AI-generated seem high.)