Kubernetes · AI/ML · 7 min read
Make the GPU a first-class citizen in Kubernetes
A practical scheduling model for mixed training and inference workloads, from device discovery to queue policy.
Practical notes about putting models into production, shaping Kubernetes platforms and making infrastructure changes safer to operate.
A practical scheduling model for mixed training and inference workloads, from device discovery to queue policy.
Treat model rollout as a control loop: version the artifact, shift traffic deliberately and let service signals decide what happens next.
Validated zone files are a start; address ownership, drift detection and rollback turn DNS automation into an operational control.