Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Agentic AI in Operations
- The shift from static runbooks to reasoning agents: the progression of IT automation
- Component analysis: reasoning loops, tool utilization, memory, and planning
- Determining when to automate versus when to retain human involvement
Agent Frameworks and Structural Designs
- Single-agent methodologies: ReAct, Plan-and-Execute, and tool-calling cycles
- Multi-agent structures: supervisor, hierarchical, and swarm configurations
- Framework evaluation: LangGraph, CrewAI, AutoGen, and custom agent development
- Creating your initial operational agent: querying monitoring, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to APIs for Prometheus, Grafana, Datadog, and PagerDuty
- Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: executing kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces featuring parameter validation and idempotency
Automating Incident Response
- Automated incident triage: classifying severity and routing responses
- Generating root cause hypotheses and collecting supporting evidence
- Automated corrective actions: restarting, scaling, rolling back, and failing over
- Constructing an incident runbook agent with progressive levels of autonomy
Safety, Guardrails, and Human Oversight
- Classifying actions: read-only, low-risk, high-risk, and destructive
- Establishing approval gates and escalation protocols for critical operations
- Guardrail strategies: action allowlists, blast radius constraints, and rollback assurances
- Maintaining audit trails and decision provenance for compliance adherence
Orchestrating Multi-Agent Systems for Complex Incidents
- Coordinating specialized agents: triage, diagnosis, and remediation units
- Managing inter-agent communication and shared context
- Resolving conflicts when agents suggest opposing actions
- Simulating end-to-end major incidents with multi-agent responses
Observability and Performance Evaluation
- Tracing agent reasoning chains for debugging and audit purposes
- Assessing agent decision quality: precision, recall, and resolution time
- Implementing feedback loops: learning from operator overrides and outcomes
- Tracking costs and analyzing token economics for operational agents
Production Deployment and Operational Management
- Deploying agents as services: utilizing APIs, webhooks, and scheduled tasks
- Phased autonomy rollout: transitioning from shadow mode to full auto-remediation
- Runbooks for agent failures: procedures when the agent itself malfunctions
- Developing business cases and measuring ROI for autonomous operations
Requirements
- Practical experience in IT operations, DevOps, or SRE methodologies.
- Proficiency in Python scripting and REST API interactions.
- Fundamental knowledge of Large Language Model (LLM) capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers interested in AI-driven automation.
- Platform engineers focused on creating self-healing infrastructure.
- IT operations leaders evaluating agentic AI solutions for incident management.
14 Hours