"We Hit 10x Query Volume With Zero Latency Spikes": An SRE's AI Infrastructure Story
An SRE's real-world story of scaling AI inference infrastructure from baseline to 10x query volume — with zero latency spikes. Covers observability, KEDA autoscaling, semantic caching, and warm pool strategies.