Building llama.cpp from source against CUDA 13.3 on Fedora, and getting a 35B-parameter MoE model running comfortably on a single 4080.