
I did it with nvfp4 + vllm, it's the mtp3 version.

I installed it with mtp 2.
It seems that token generation has improved a bit, but there isn't a big difference.
Still, mtp2 would be more stable than 3, right????????
The optimal recipe hasn't been created yet,
qwen 122b-a10b came out with 60 tokens through optimization. I believe that if the optimization is done well, token generation can be further improved.