Start with workload shape, not a GPU count

Training jobs are elastic, long-running and often checkpointable. Online inference is latency-sensitive, bursty and sensitive to noisy neighbors. A cluster that gives both workloads the same queue and the same priority is making an architectural decision; it is just making it invisibly.

Name a small set of workload classes first: interactive experiments, batch training, and production inference are a useful starting point. For each class, decide its latency objective, preemption policy, accelerator profile and maximum queue time. These are service-level choices, not scheduler trivia.

Make hardware pools explicit

Separate nodes by the hardware they actually offer: memory size, architecture, interconnect and partitioning mode matter as much as the accelerator model. Use labels for selection and taints to keep general workloads from consuming reserved capacity. The policy should be managed with the node-pool definition so it stays reviewable alongside the cluster.

Where the hardware supports partitioning, publish those profiles as distinct capacity. A small inference slice and a full training device are different products, even when they share a card. Keep the device plugin and driver versions pinned to the node image and roll them as a tested unit.

Queues need fairness and a way out

A queue prevents a burst of experiments from turning into a thundering herd, but FIFO alone can strand short jobs behind a large training run. Add per-team quotas, bounded priorities and a documented preemption policy. Reserve a floor for production inference, then let training borrow unused capacity only when the reclaim path is understood.

Autoscaling must agree with the queue. Scale from pending accelerator requests and provisioning delay, not CPU utilization on nodes that are already full. Track unschedulable time separately from provisioning time; they point to different fixes.

Measure the whole request path

Collect accelerator utilization and memory alongside queue age, time-to-first-token, batch size, pod placement and model revision. A GPU at 95% utilization is not automatically healthy if the serving queue is growing or tail latency is breaching its objective.

Put the signals on one dashboard with workload class and team dimensions. That makes capacity conversations concrete: the question becomes which workload is waiting, for what profile, and for how long.

A rollout checklist

  • Validate drivers, device plugin and node labels on a canary pool before admitting production jobs.
  • Set quota and priority defaults that prevent one namespace from consuming the fleet.
  • Exercise node loss, image pull failure and model warm-up in a non-production environment.
  • Alert on queue age and inference latency, not just node health.
A useful GPU platform makes the right placement the easy default and makes scarce capacity visible before it becomes an incident.