
Agentic tools and AI-assisted code review have made shipping faster than ever. The reliability practices that keep those systems running are under pressure to match.
The systems themselves are changing too. Most teams are already running multiple models in production. When something breaks, the cause is often a rate limit, a prompt update, or a model that changed upstream rather than anything that would show up in a deploy log.
Ajuna Kyaruzi, Manager of SRE and Platform Advocacy at Datadog, will share how we can keep reliability in step with development velocity. The observability signals that help give us the complete picture, how incident response changes when systems can drift without a deployment, and what teams operating at scale are learning.
Ajuna Kyaruzi leads the SRE & Platform Advocacy team at Datadog. She cares about using software to help people sustainably run large-scale systems, focusing on Incident Response and SLOs. She loves community building and volunteers with multiple mentorship programs aimed at helping early career folks break into tech, and ensuring they have successful careers. Previously she worked at Google as a Software Engineer on Google Maps and as a Site Reliability Engineer on Google Cloud.