
When I first did it with 1 spark, the bench came out around 24~26, and in actual use it came out around 16~20.
And

When I did it with 2 sparks, it went up almost twice as much.
However, in actual use, it came out to about 25~33 tokens, so there was a difference from the bench.
And today I tried setting it up with sglang.

There is a difference between vLLM and sglang.
It seems that sglang is currently making a difference of about 10~20%.
However, in actual use, DeepSeek feels faster. The difference is bigger than the number of tokens.