Llamacpp hasn't been merged yet,
The person who made Dflash released DFlash2 as a PR a while ago.
I tried it on Qwen3.8 27B UD-Q5-KXL, which was merged this time by the person who made llamacpp for strixhalo, and it's godflash -0-;
Even with a context of over 100K, it clocks around 15t/s. In an empty state, it sometimes hits close to 30.
As a result, countless models that were in the SSD have been purged -_-;
Qwen 3.6 35B A3B was also immediately purged because Orinith 1.5 came out.
Gemma4 26B is kept alive for now because it speaks well, while 31B is purged.
Ling 3.0 Tiny is 8B A1B and does tool calls surprisingly well for its size!?
The results vary depending on where you use it, but at unsloth studio, it did a great job of summarizing news and organizing data with this model. Naturally, it's fast too.
So I kept it alive for now..; Those who haven't used it yet, please try it once.