vLLM keeps model weights in GPU memory across engine restarts
vLLM added a preload command that keeps weights in GPU memory through restarts, gated multimodal request settings behind a flag and removed several options.
vLLM added a preload command that keeps weights in GPU memory through restarts, gated multimodal request settings behind a flag and removed several options.