(www.tomshardware.com)

521 points hd4 | 2 comments | 20 Oct 25 12:31 UTC | HN request time: 0.001s | source

Paper: https://dl.acm.org/doi/10.1145/3731569.3764815

Show context

kilotaras ◴[20 Oct 25 15:05 UTC] No.45644776[source]▶

Alibaba Cloud claims to reduce Nvidia GPU used for serving unpopular models by 82% (emphasis mine)

> 17.7 per cent of GPUs allocated to serve only 1.35 per cent of requests in Alibaba Cloud’s marketplace, the researchers found

Instead of 1192 GPUs they now use 213 for serving those requests.

replies(5): >>45645037 #>>45647752 #>>45647863 #>>45651559 #>>45653363 #

bee_rider ◴[20 Oct 25 19:06 UTC] No.45647863[source]▶

>>45644776 #

I’m slightly confuse as to how all this works. Do the GPUs just sit there with the models on them when the models are not in use?

I guess I’d assumed this sort of thing would be allocated dynamically. Of course, there’s a benefit to minimizing the number of times you load a model. But surely if a GPU+model is idle for more than a couple minutes it could be freed?

(I’m not an AI guy, though—actually I’m used to asking SLURM for new nodes with every run I do!)

replies(6): >>45648058 #>>45648291 #>>45648653 #>>45649219 #>>45650208 #>>45653517 #

svachalek ◴[20 Oct 25 20:53 UTC] No.45649219[source]▶

>>45647863 #

Models take a lot of VRAM which is tightly coupled to the GPU so yeah, it's basically sitting there with the model waiting for use. I'm sure they probably do idle out but a few minutes of idle time is a lot of waste--possibly the full 82% mentioned. In this case they optimized by letting the GPUs load multiple models and sharing the load out by token.

replies(2): >>45650833 #>>45651174 #

1. andy_ppp ◴[20 Oct 25 23:46 UTC] No.45650833[source]▶

>>45649219 #

How does this work with anything but trivially small context sizes!?

replies(1): >>45651181 #

2. jychang ◴[21 Oct 25 00:39 UTC] No.45651181[source]▶

>>45650833 (TP) #

Tensor parallelism, so you only need to store a fraction of kv cache per gpu.

↑

Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system