DeepSpeech V4 0731 was rolled back to Unslos Q2KL
It goes up to IQ3XXS, but it's too tight and the speed is a bit slow,
Anyway, when I run it with ROCm, the prefill seems to stay around 100 seconds, and the tokens stay around 12-13 per second.
If you run it on ds4, the prefill is 150-130 and the tokens are 15-16 more
It's a hassle to go back and forth with ds4, so I switched it to llama for now.
The picture is the result of running the prompt on the link.
There seems to be something wrong with vulkan, so I think some modifications will be needed. If you use the strix halo kv cache modified version, it will generally increase by about 30%, but llama hasn't merged it yet.ㅜㅡㅠ

▶ Original source: https://sunsetcompare.web.app/