Get in Touch

Course Outline

Foundations of Agentic AI in Operations

  • The shift from static runbooks to reasoning agents: the progression of IT automation
  • Component analysis: reasoning loops, tool utilization, memory, and planning
  • Determining when to automate versus when to retain human involvement

Agent Frameworks and Structural Designs

  • Single-agent methodologies: ReAct, Plan-and-Execute, and tool-calling cycles
  • Multi-agent structures: supervisor, hierarchical, and swarm configurations
  • Framework evaluation: LangGraph, CrewAI, AutoGen, and custom agent development
  • Creating your initial operational agent: querying monitoring, diagnosing issues, and proposing solutions

Tool Integration for IT Operations

  • Connecting agents to APIs for Prometheus, Grafana, Datadog, and PagerDuty
  • Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk
  • Leveraging infrastructure tools: executing kubectl, Terraform, and Ansible via agent actions
  • Designing secure tool interfaces featuring parameter validation and idempotency

Automating Incident Response

  • Automated incident triage: classifying severity and routing responses
  • Generating root cause hypotheses and collecting supporting evidence
  • Automated corrective actions: restarting, scaling, rolling back, and failing over
  • Constructing an incident runbook agent with progressive levels of autonomy

Safety, Guardrails, and Human Oversight

  • Classifying actions: read-only, low-risk, high-risk, and destructive
  • Establishing approval gates and escalation protocols for critical operations
  • Guardrail strategies: action allowlists, blast radius constraints, and rollback assurances
  • Maintaining audit trails and decision provenance for compliance adherence

Orchestrating Multi-Agent Systems for Complex Incidents

  • Coordinating specialized agents: triage, diagnosis, and remediation units
  • Managing inter-agent communication and shared context
  • Resolving conflicts when agents suggest opposing actions
  • Simulating end-to-end major incidents with multi-agent responses

Observability and Performance Evaluation

  • Tracing agent reasoning chains for debugging and audit purposes
  • Assessing agent decision quality: precision, recall, and resolution time
  • Implementing feedback loops: learning from operator overrides and outcomes
  • Tracking costs and analyzing token economics for operational agents

Production Deployment and Operational Management

  • Deploying agents as services: utilizing APIs, webhooks, and scheduled tasks
  • Phased autonomy rollout: transitioning from shadow mode to full auto-remediation
  • Runbooks for agent failures: procedures when the agent itself malfunctions
  • Developing business cases and measuring ROI for autonomous operations

Requirements

  • Practical experience in IT operations, DevOps, or SRE methodologies.
  • Proficiency in Python scripting and REST API interactions.
  • Fundamental knowledge of Large Language Model (LLM) capabilities and prompt engineering.

Target Audience

  • SRE and DevOps engineers interested in AI-driven automation.
  • Platform engineers focused on creating self-healing infrastructure.
  • IT operations leaders evaluating agentic AI solutions for incident management.
 14 Hours

Upcoming Courses

Related Categories