vLLM 0.31.0 shipped on October 5, 2026, with fast-restart features that can reduce repeated initialization work for model-serving engines. The release introduces a preload CLI for cached weights and experimental initialized-engine snapshots. These mechanisms address different parts of startup and should not be treated as interchangeable. The official release notes establish the version and feature status.
Weight caching keeps resources alive across engine restarts
The release describes a daemon that retains post-quantized weights in GPU memory across engine restarts, with support for data parallelism, draft models and readiness checks. The preload implementation merged September 25 before inclusion in the October release.
Retaining weights can avoid repeating expensive work, but it also changes memory ownership. Stopping the serving engine may no longer release every resource associated with the model. Operators need to understand which process holds the cache, how it is stopped and how another workload obtains the GPU when the model should be removed.
Initialized-engine snapshots remain experimental
The notes describe CRIU-based snapshots that restore a fully initialized TP1 engine. That is a narrower capability than general distributed snapshot recovery. A supported single-engine path should not be extrapolated to every tensor-parallel configuration or deployment environment.
Test recovery after the failures that matter in the actual service. Restarting a process on the same healthy host differs from losing the machine or changing the driver environment. Record whether the recovery restores only initialization state or any application-level work; an inference server’s startup optimization does not resume a business workflow automatically.
Measure readiness, not just process launch
Compare cold start, cached restart and snapshot restore separately. The useful endpoint is when the engine can accept representative traffic successfully, not when a process exists or an HTTP port opens. Include a request that exercises the selected model’s normal context and output behavior.
Verify memory usage after repeated restart cycles and when switching checkpoints. Keep the release revision and configuration with the results, and stage the upgrade where it can be reversed. vLLM 0.31 offers meaningful tools for reducing restart overhead, but their benefit depends on a clear resource lifecycle and recovery behavior in the deployed environment.