Let me introduce a good article.
I think it would be nice to have even one DGX Spark, but I also wonder what it would be like to actually have two.
These days the service prices of Chinese open-weight models keep getting cheaper, so from an individual's standpoint I sometimes wonder whether it is worth going to the trouble of obtaining expensive high-end equipment. Reading the article, the author has organized it really well and written down the questions in an easy-to-understand way. Thinking simply, the reason API calls are processed quickly is probably the effect of batch size and efficiency on large-scale distributed compute resources.
Running a large model at home on two DGX Sparks probably does not suit impatient Koreans like us. For the sake of mental health, using a commercial service seems better. (This is me consoling myself.)
▶ Original source: https://seapy.com/deepseek-v4-flash-api-vs-2x-dgx-spark/