https://github.com/Nathanw1014/strix-halo-llamacpp
This is a version where the developer fixed some inefficient operations between Strix Halo and llamacpp.
I heard that a merge request was made to llamacpp, but it seems that it hasn't been approved yet because there wasn't enough proof,
but I think it's significantly improved.
Based on Qwen 3.6 27B Q6K, with 180K Prefill, it achieves 130t/s and about 10tps in normal operation.