
Cloud infrastructure incidents are often hard to troubleshoot because the problem is not always visible from the application layer. A service may look healthy, the instances may be running, and the dashboards may not clearly show the issue, but connectivity can still fail because of routing, security rules, DNS, or cloud configuration problems. In this talk, I will show how AI-assisted workflows can support SREs and cloud engineers during infrastructure troubleshooting. The focus is not on using AI to replace engineering judgment, but on using it to investigate faster, ask better questions, review configuration, interpret CLI output, and reduce the time spent jumping between documentation, consoles, and logs. As a practical demo, I will use AWS MCP servers to investigate a real-world style cloud connectivity issue and work toward the root cause. We will look at what context is useful to give the AI assistant, how MCP can help connect the assistant to AWS documentation and environment information, and where the engineer still needs to validate the output. The goal is to give attendees a realistic view of what AI can and cannot do during real infrastructure incidents. The session will focus on practical troubleshooting patterns, safe usage of AI during production-like investigations, and the importance of human validation when dealing with cloud reliability and infrastructure security.
Pooria Ghaedi is a Senior Cloud Engineer focused on cloud infrastructure, reliability, infrastructure security, and automation. He works on designing and improving scalable cloud platforms, with a strong interest in practical troubleshooting and operational excellence. He is an AWS Community Builder in the Networking & Content Delivery category and writes about cloud engineering, infrastructure security, and AI-assisted engineering workflows.