# Agentic AI for Cloud Reliability

## Abstract

Cloud systems are becoming ever more critical; yet failures are the norms in the cloud. Outages and incidents occur every day, and downtime for large-scale systems can cost hundreds of thousands of dollars per hour. Despite massive investments, cloud system reliability today still relies heavily on human engineers. This raises a fundamental research challenge: how can we embed intelligence into systems so they can autonomously detect, diagnose, and recover from failures safely and efficiently?

In this talk, I will present my research on the design of AI-operable cloud systems that are built with AI across the incident lifecycle. I will first introduce my work on AI-augmented root cause analysis, where large language models can reason over heterogeneous telemetry to localize failures. I will then turn to failure mitigation, where I design reliable and safety-aware agentic systems capable of executing recovery actions automatically. I will conclude by outlining my future research agenda on improving the reliability, efficiency, and security of AI and systems.

## Bio

Yinfang Chen is an Assistant Professor at Arizona State University. He got his Ph.D. in Computer Science at the University of Illinois Urbana-Champaign, advised by Prof. Tianyin Xu. His research sits at the intersection of artificial intelligence, computer systems and software engineering. He works in both directions, leveraging AI to operate computer systems, e.g., detecting bugs, doing root cause analysis, and autonomous mitigation from failures, and building systems that make AI more reliable and efficient.
