Keep the request alive while compute is absent
The gateway accepts a prompt, writes a job to Redis and returns an ID immediately. A client polls for pending, completed or failed status instead of holding an HTTP connection through node provisioning and model load. This separates request intake from GPU availability, but Redis and the gateway remain online and continue to cost money while the GPU is off.
The CPU worker moves each job into a processing list before calling vLLM's completions endpoint. KEDA watches both the waiting and processing lists for the worker and vLLM Deployments. That second trigger matters: once the worker takes the last item, an empty waiting list must not cause the model to disappear during inference.
Give an interrupted job a way back
On restart, the single worker requeues unfinished items from the processing list. It writes the result and acknowledges the item in one Redis transaction. A crash after vLLM generates a response can still repeat inference, so this is at-least-once processing, not exactly-once execution. The worker is deliberately limited to one replica until recovery ownership is designed for multiple consumers.
Redis uses append-only persistence on a volume to improve recovery, but it is not a guarantee against every storage failure. Results expire after five minutes; clients need to poll within that window.
Separate pod scaling from node scaling
KEDA scales the CPU worker and GPU-backed vLLM pods from zero when a job appears. On a configured GKE cluster, the pending vLLM pod's GPU request is what should prompt the node pool autoscaler to add a GPU VM. Kubernetes manifests alone cannot create that pool, install drivers or prove a scale-to-zero node cycle; the repository includes the prerequisite and deployment steps rather than hiding them in an application diagram.
The model cache lives on a persistent volume so weights can survive pod and node churn after the initial download. Separating the approximately 3.5 GB of Qwen weights from the approximately 8 GB vLLM runtime also makes the runtime image suitable for a GKE Secondary Boot Disk cache.
Measure the complete cycle before optimizing
The repository includes a concurrent submit-and-poll driver, a bounded cold/warm/full-zero capture script, Redis queue metrics and Prometheus scrape targets for vLLM. The capture refuses to call a run cold unless both pods and the GPU pool begin at zero. It records cluster events and returns nonzero if jobs fail or the node does not scale down.
Local tests exercise the request/result contract, failed jobs and recovery from an interrupted item. The measured T4 Spot run in us-east1-d reduced queue-spike-to-first-token time from 659 seconds to 338 seconds. The PV-only improvement is an estimate, not an isolated benchmark; the two optimizations were measured together. The remaining time is mostly GPU VM bring-up and loading weights from the network-attached PVC into VRAM. The full build and deployment record is in the repository's cold-start research note, while DCGM data and other hardware-specific results remain environment-specific.
A GPU node at zero is only one part of the bill. A useful demonstration shows the queue, the pod, the node and the remaining always-on services on the same clock.