Kubernetes

Platform Engineering for Developers, Architects & the Rest of Us • Daniel Bryant • GOTO 2025

Platform Engineering for Developers, Architects & the Rest of Us • Daniel Bryant • GOTO 2025

Daniel Bryant discusses platform engineering for software developers and architects, emphasizing treating platforms as internal products with developers as customers. He outlines a three-layered architecture, the evolution from monolithic systems to microservices, and the importance of 'golden bricks' over 'golden paths' for composability. Key takeaways include API-first design, minimizing cognitive load, avoiding leaky abstractions, and measuring success through frameworks like DORA and DevEx to achieve speed, safety, and scale.

Agentic SDLC at Uber — Uday Kiran Medisetty & Adam Huda, Uber

Agentic SDLC at Uber — Uday Kiran Medisetty & Adam Huda, Uber

Uber has transformed its software development with an agentic AI-powered factory, leading to a dramatic increase in engineer productivity. The presentation details six key infrastructure components: a unified model gateway with strict PII and safety guardrails, an MCP gateway for streamlined agent tool access and token optimization, agentified dev pods for rapid execution, a managed skills marketplace, a comprehensive context graph, and the Cortana AI assistant. Adam Huda then demonstrates an end-to-end feature development workflow, highlighting a critical shift to inner-loop validation (stopping short of CI) and automated, managed maintenance loops. The ultimate takeaway is that the bottleneck has moved from technical execution to strategic decision-making: "should we build it?" rather than "can we build it?"

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai

Krea's Gabriel Jorge Menezes shares critical insights into building the infrastructure for large-scale ML model training and serving. Key takeaways include the necessity of custom metrics beyond standard GPU utilization, aggressive checkpointing on ultra-fast storage to counter frequent cluster crashes, and a dynamic Kubernetes-based system using gang scheduling, virtual-kubelet, and taints to seamlessly shift inference workloads to external providers when training consumes on-prem GPUs. The approach highlights practical solutions for silent failures, thermal management, and optimizing resource utilization in a unified production and training environment.

Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face

Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face

Arek Borucki details how Hugging Face scales its infrastructure to serve millions of models and users, focusing on the evolution of search architecture using MongoDB Atlas and Apache Lucene, robust database scaling with a seven-node cluster and sharding, and dynamic frontend autoscaling with Kubernetes and KEDA to ensure an instant, seamless user experience.

WebAssembly on Kubernetes • Nicolas Frankel • YOW! 2025

WebAssembly on Kubernetes • Nicolas Frankel • YOW! 2025

Nicolas Fränkel explores the evolution of WebAssembly (Wasm) beyond its web origins, showcasing its potential to revolutionize application deployment on Kubernetes. The talk demonstrates how Wasm enables incredibly small container sizes (down to 2MB for an HTTP server) by integrating specific Wasm runtimes with Kubernetes' extensible architecture. However, Fränkel also provides a candid assessment of the ecosystem's rapid, often unstable, development, recommending Wasm on Kubernetes for agile startups seeking competitive advantage but cautioning traditional enterprises due to the inherent risks and maintenance challenges.

The Platform Engineer’s Handbook • Ajay Chankramath & Kaspar von Grünberg • GOTO 2026

The Platform Engineer’s Handbook • Ajay Chankramath & Kaspar von Grünberg • GOTO 2026

This conversation with Ajay Chankramath, author of 'The Platform Engineer’s Handbook,' delves into why practical, code-first guidance is essential for building Internal Developer Platforms. He argues that developer adoption failures stem from a "product discipline gap," not a technology one, emphasizing developer experience as a first-class outcome. The discussion covers the book's arc from foundations to enterprise-grade features and its focus on 100% open-source, vendor-agnostic tooling. Crucially, it highlights how agentic AI raises the stakes for platform engineering, requiring new IDP layers for agent context, memory, and guardrails, asserting that these must be built, owned, and operated internally for safe and productive AI adoption.