Search results for: kubernetes


Whose GPUs are these, anyway? Secure, self-service metrics for multi-tenant Kubernetes

Posted on September 9, 2026 | Bingi Narasimha Karthik (Golden Kubestronaut, Adobe) and Ramkumar Nagaraj (Golden Kubestronaut, Adobe)

The question that stopped the meeting It was a routine cost review. The slide showed the month’s GPU spend, the biggest line on the whole infrastructure bill, and someone asked a five-word question: “Are we using…


Kubernetes access via an identity provider: Public client, not confidential

Posted on September 8, 2026 | Kolawole Olowoporoku | CNCF Ambassador and Senior Platform Engineer

Access control belongs on the same day-zero checklist as networking and storage. On most on-prem clusters, it never makes the list. The Identity Gap Managed cloud Kubernetes ships IAM or SSO integration out of the box….


China Merchants Bank Wins CNCF End User Case Study Contest for Unifying AI Training and Inference on Kubernetes

Posted on September 7, 2026

New cloud native platform lifted average accelerator compute utilization from 35% to more than 60% and cut inference cost per 1 million tokens by more than 60% Key Highlights SHANGHAI, China – KubeCon + CloudNativeCon +…


Kubernetes isn’t new, but AI makes It scary again

Posted on September 4, 2026 | Andy Suderman, CTO Fairwinds

Kubernetes isn’t brand new anymore. Yet, for many teams, adopting it still feels intimidating. Even if you’ve watched Kubernetes become the default foundation for production software and AI workloads, it can still feel like a big…


Migrating a critical Kubernetes deployment from the default namespace without any downtime

Posted on September 3, 2026 | George Sims, Downtherabbithole.dev

Somewhere in your cluster there’s probably a deployment sitting in the default namespace that everyone knows shouldn’t be there. Nobody put it there maliciously, it just happened, early on, before anyone had opinions about namespace hygiene,…


Observability in Kubernetes: From metrics to meaning

Posted on August 31, 2026 | Neel Shah, Stackgen

Kubernetes made infrastructure more programmable, scalable, and resilient. It also made production systems harder to reason about. Workloads move, replicas churn, dependencies multiply, and a single user request can cross ingress, services, queues, storage, and background…


Scale before the spike: Predictive autoscaling for GPU workloads on Kubernetes

Posted on August 28, 2026 | Ramkumar Nagaraj (Golden Kubestronaut, Adobe) and Bingi Narasimha Karthik (Golden Kubestronaut, Adobe)

The 3 AM Call We got paged one Tuesday morning. A critical production service had crashed under traffic—not gradually degraded, but crashed. Hundreds of pending pods. Users were seeing 15–20% error rates. The incident postmortem was…


Your Kubernetes platform is ready for containers. Is it ready for AI?

Posted on August 28, 2026 | Kasia Hilborne, Vultr

Kubernetes has given platform teams a consistent way to deploy, scale, and operate containerized applications. Now, many of those same teams are being asked to support AI. The transition is already underway. According to the CNCF…


Network World: “Kubernetes 1.37 advances workload-aware scheduling and cluster networking”

Posted on August 27, 2026

New scheduling capabilities were built for AI and machine-learning training jobs that need groups of pods to start and scale together. 


Building an AI factory on Kubernetes

Posted on August 27, 2026 | Hrittik Roy | CNCF Ambassador and Platform Advocate at vCluster

An AI factory is not just a model or a cluster. It is a pool of GPUs that many teams draw from at once: one team fine-tuning, another serving inference, a third running evaluations, all on…