AI-Driven Observability: From Logs to LLM-Powered Insights Training Course
Conventional observability depends heavily on dashboards, threshold-based alerts, and manual log inspection. AI-driven observability revolutionizes this approach by enabling natural language queries of telemetry data, utilizing LLMs for root cause analysis, detecting anomalies via foundation models, and generating context-aware automated incident summaries.
This instructor-led, live training (available online or onsite) is designed for observability and SRE engineers seeking to integrate LLMs and AI into their monitoring, alerting, and incident analysis processes.
Upon completion of this training, participants will be able to:
- Develop natural language interfaces for querying Prometheus, Elasticsearch, and SQL-based observability repositories.
- Implement log analysis and anomaly detection pipelines powered by LLMs.
- Create automated incident summaries and draft postmortems from raw telemetry data.
- Design AI-assisted root cause analysis workflows incorporating evidence chaining.
- Incorporate foundation models for time-series anomaly detection and forecasting.
- Deploy an AI-enhanced on-call experience featuring intelligent alert enrichment.
Course Format
- Interactive lectures and discussions.
- Extensive exercises and practice sessions.
- Hands-on implementation within a live-lab environment.
Customization Options
- To request customized training, please contact us to make arrangements.
Course Outline
The AI Observability Landscape
- Moving from dashboards to conversations: the transition toward AI-augmented observability.
- LLM capabilities relevant to observability: summarization, reasoning, and pattern matching.
- Architecture patterns for embedding AI into existing observability stacks.
Natural Language Telemetry Querying
- Text-to-PromQL: translating natural language into monitoring queries.
- NL querying for Elasticsearch, OpenSearch, and Loki log stores.
- SQL generation from natural language for structured telemetry.
- Building a query assistant agent with tool use and context awareness.
LLM-Powered Log Analysis
- Automated log parsing and structuring using LLMs.
- Anomaly detection in log streams via embedding similarity.
- Log clustering and pattern discovery at scale.
- Generating human-readable explanations from raw log sequences.
Intelligent Alerting and Incident Enrichment
- Alert correlation and deduplication with semantic understanding.
- Automated gathering of incident context from runbooks, past incidents, and documentation.
- Smart alert routing based on content understanding and team expertise.
- Reducing alert fatigue through AI-driven noise reduction.
AI-Assisted Root Cause Analysis
- Hypothesis generation from multi-source telemetry correlation.
- Evidence chaining: linking symptoms across metrics, logs, and traces.
- Guided troubleshooting with interactive AI diagnosis sessions.
- Building a root cause analysis agent with progressive investigation.
Automated Incident Response and Communication
- Generating incident summaries and status updates from telemetry data.
- Automated postmortem drafting with timeline reconstruction.
- Tailored stakeholder communication for technical and executive audiences.
- Runbook suggestions and automated remediation recommendations.
ML for Observability
- Time-series forecasting for capacity planning and anomaly prediction.
- Foundation models for zero-shot anomaly detection on metrics.
- Embedding-based service dependency mapping and topology discovery.
- Training and deploying lightweight ML models alongside observability pipelines.
Production Deployment and Ethics
- Latency and cost considerations for real-time AI observability.
- Data privacy: ensuring LLMs do not leak sensitive telemetry.
- Human oversight: when AI diagnosis needs operator validation.
- Measuring impact: MTTD, MTTR, and on-call experience metrics.
Requirements
- Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Familiarity with log management and metrics concepts.
- Basic proficiency in Python scripting for data processing.
Audience
- SRE and observability engineers adopting AI-enhanced tooling.
- Platform engineers building next-generation monitoring pipelines.
- DevOps leads evaluating LLM integration into incident workflows.
Open Training Courses require 5+ participants.
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Booking
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Enquiry
AI-Driven Observability: From Logs to LLM-Powered Insights - Consultancy Enquiry
Upcoming Courses
Related Courses
Agentic Development with Gemini 3 and Google Antigravity
21 HoursGoogle Antigravity serves as an agentic development environment engineered to create autonomous agents that can plan, reason, code, and execute actions, leveraging the multimodal capabilities of Gemini 3.
This live, instructor-led training session, available both online and onsite, is tailored for advanced technical professionals seeking to design, construct, and deploy autonomous agents using Gemini 3 within the Antigravity ecosystem.
Upon completion of this course, participants will be equipped to:
- Construct autonomous workflows that leverage Gemini 3 for reasoning, planning, and execution tasks.
- Create agents within Antigravity capable of analyzing assignments, generating code, and interacting with various tools.
- Integrate Gemini-powered agents with enterprise systems and external APIs.
- Enhance agent behavior, ensuring safety and reliability in complex operational environments.
Course Structure
- Expert-led demonstrations paired with interactive discussions.
- Practical experimentation focused on developing autonomous agents.
- Hands-on implementation utilizing Antigravity, Gemini 3, and associated cloud tools.
Customization Options
- Should your team require domain-specific agent behaviors or bespoke integrations, please reach out to us to tailor the program to your needs.
Advanced Antigravity: Feedback Loops, Learning & Long-Term Agent Memory
14 HoursGoogle Antigravity serves as a sophisticated framework for exploring long-lived agents and the emergent interactive behaviors they exhibit.
This instructor-led live training, available both online and on-site, is tailored for advanced professionals seeking to design, analyze, and optimize agents that can retain memory, improve through feedback, and evolve over extended operational periods.
Upon completion of this course, participants will acquire the skills necessary to:
- Architect long-term memory structures for agent persistence.
- Implement robust feedback loops to refine agent behavior.
- Evaluate learning trajectories and monitor model drift.
- Integrate memory mechanisms within complex multi-agent ecosystems.
Course Format
- Expert-led discussions complemented by technical demonstrations.
- Hands-on exploration via structured design challenges.
- Application of concepts in simulated agent environments.
Customization Options
- If your organization requires tailored content or case-specific examples, please reach out to us to customize this training.
Advanced Mastra Integrations: APIs, Tools, Enterprise Data & External Systems
21 HoursMastra provides a robust framework for achieving deep integration between AI agents, APIs, enterprise applications, and external data systems.
This instructor-led training, available online or onsite, is designed for intermediate-level engineers seeking to establish reliable, secure, and scalable connections between Mastra agents and the broader enterprise ecosystem.
Upon completing this training, participants will be equipped to:
- Execute API-driven integrations connecting Mastra agents with external services.
- Link enterprise data systems and tools to automated agent workflows.
- Implement best practices for secure data exchange and authentication.
- Architect integration layers that are scalable, maintainable, and ready for production.
Course Format
- Interactive lectures and discussions.
- Practical integration engineering and API exercises.
- Live-lab implementations utilizing real-world enterprise scenarios.
Customization Options
- Custom API scenarios, enterprise system mappings, or data-integration workshops can be arranged upon request.
Interactive AI Agents: AgentCore Memory, Code Interpreter & Browser Tool in Action
14 HoursAgentCore equips AI agents with the ability to retain persistent memory, execute secure code, and interact with browser environments, enabling them to deliver highly interactive, dynamic, and context-aware user experiences.
Designed for intermediate to advanced technical practitioners, this instructor-led live training (available online or on-site) focuses on designing and deploying AI agents that can maintain long-term context, perform real-time calculations, and engage directly with web-based user interfaces.
Upon completing this program, participants will be equipped to:
- Implement AgentCore memory to support stateful, context-aware workflows.
- Utilize the secure code interpreter for dynamic calculations and data transformations.
- Integrate the browser tool to facilitate real-time data retrieval and UI interaction.
- Architect interactive agents for analytics, customer support, and research applications.
Course Format
- Interactive lectures and facilitated discussions.
- Practical lab exercises featuring AgentCore memory and associated tools.
- Analytical case studies covering analytics, automation, and customer support scenarios.
Customization Options
- Please reach out to arrange a customized training experience tailored to your specific requirements.
Accelerating AI Agent Deployment with AgentCore Runtime & Gateway
14 HoursAgentCore Runtime & Gateway provides a complementary pair of AWS services designed to streamline the packaging, deployment, and secure exposure of AI agents, featuring simplified integrations with external systems.
This instructor-led live training, available both online and onsite, targets intermediate-level engineering teams aiming to transition their agents from prototype stages to production. Participants will gain expertise in using the AgentCore Runtime for deployment and the Gateway for secure connectivity and API integration.
Upon completing this training, participants will be able to:
- Provision AgentCore Runtime environments and package agents for deployment.
- Expose agents via the Gateway using authenticated, rate-limited endpoints.
- Incorporate external tools and APIs into agent workflows through stable contracts.
- Implement observability, logging, and usage monitoring for production operations.
Course Format
- Interactive lectures and discussions.
- Hands-on labs covering Runtime deployments and Gateway integrations.
- Practical exercises emphasizing reliability, security, and deployment strategies.
Customization Options
- To request a customized training session for this course, please contact us to arrange.
Antigravity for Developers: Building Agent-First Applications
21 HoursAntigravity is a dedicated development platform engineered to support the creation of AI-driven, agent-first applications.
Delivered as instructor-led, live training (available online or on-site), this course is tailored for intermediate-level developers looking to engineer practical applications utilizing autonomous AI agents within the Antigravity ecosystem.
Upon completion, participants will possess the skills to:
- Construct applications powered by autonomous and coordinated AI agents.
- Leverage the Antigravity IDE, editor, terminal, and browser for comprehensive end-to-end development.
- Orchestrate multi-agent workflows using the Agent Manager.
- Embed agent capabilities into robust, production-grade software architectures.
Course Format
- Blended instructional sessions featuring detailed demonstrations.
- Substantial hands-on practice supported by guided exercises.
- Practical implementation tasks conducted directly within the live Antigravity environment.
Customization Options
- To align content specifically with your development stack, please reach out to arrange a customized training version.
Getting Started with Antigravity: An Introduction to Agent-First IDEs
14 HoursGoogle Antigravity is an agent-first development environment designed to streamline engineering workflows through intelligent automation.
This instructor-led, live training (online or onsite) is aimed at beginner-level practitioners who wish to explore the fundamentals of Antigravity and understand how agent-driven coding environments enhance productivity.
Upon completion of this training, participants will be able to:
- Install and configure Google Antigravity.
- Navigate and understand both the Editor View and Manager View.
- Work effectively with agents to automate simple development tasks.
- Use Antigravity to generate, refine, and manage project files.
Format of the Course
- Instructor explanations supported by real-time demonstrations.
- Guided exercises focused on hands-on use of agents.
- Practical exploration of core Antigravity features in a controlled lab environment.
Course Customization Options
- If you require a tailored version of this training, please contact us to arrange a customized program.
Antigravity for Web Automation & Browser-Based Tasks
21 HoursGoogle Antigravity serves as a robust platform for developing agents that seamlessly interact with web applications, browser environments, and complex multi-surface workflows.
Designed for intermediate-level professionals, this instructor-led live training—available online or onsite—focuses on constructing, automating, and rigorously testing browser-based workflows utilizing Google Antigravity.
By the conclusion of the program, participants will possess the capability to:
- Develop agents capable of interacting with web applications within a browser surface.
- Streamline end-to-end workflows across various browser contexts.
- Thoroughly validate and troubleshoot agent performance in UI-driven environments.
- Deploy advanced cross-surface automation strategies using Antigravity.
Course Structure and Delivery
- Expert-led instruction complemented by practical demonstrations.
- Immersive, hands-on activities and scenario-based practical exercises.
- Deployment of agent workflows within an interactive laboratory setting.
Tailored Learning Options
- To align the curriculum with specific organizational objectives, please reach out to discuss customized training solutions.
Building Fully Managed AI Agents with AgentCore: From Concept to Production
14 HoursAgentCore streamlines the creation, refinement, and oversight of fully managed AI agents by offering a cohesive suite of services designed for scalable implementation.
This live, instructor-led session, available online or on-site, is tailored for practitioners ranging from beginners to intermediate levels who seek practical experience in developing production-grade AI agents using AgentCore.
Upon completion, participants will be equipped to:
- Grasp the fundamental capabilities of AgentCore for AI agent creation.
- Design and set up basic AI agents leveraging managed services.
- Incorporate workflows to expand agent capabilities.
- Deploy and oversee AI agents within production settings.
Instructional Approach
- Engaging lectures and collaborative discussions.
- Practical labs focusing on AgentCore services.
- Supervised exercises guiding the journey from conceptualization to deployment.
Tailored Learning Options
- To explore bespoke training options for this curriculum, please reach out to us to discuss arrangements.
AI Agent Development with Mastra
14 HoursThis instructor-led, live training (available online or onsite) is designed for intermediate-level software developers and engineering teams seeking to build scalable, observable AI systems using Mastra.
By the end of this training, participants will be capable of:
- Understanding Mastra’s architecture and how it connects with LLMs and external APIs.
- Designing and implementing AI agents and workflows using TypeScript.
- Utilizing Mastra’s observability and memory tools to oversee and refine agent performance.
- Deploying production-ready AI applications by harnessing Mastra’s framework features.
Mastra Debugging, Evaluation & Quality Assurance for AI Agents
21 HoursMastra is a framework that provides structured tools for evaluating, debugging, and assuring the reliability of AI agents operating across complex workflows.
This instructor-led, live training (online or onsite) is aimed at intermediate-level practitioners who wish to rigorously test agent behavior, improve reliability, and implement measurable evaluation processes.
At the end of this training, participants will confidently:
- Apply debugging techniques to identify and correct agent behavior issues.
- Evaluate agents using structured metrics, benchmarks, and quality scores.
- Implement tooling and workflows that track reliability, drift, and hallucinations.
- Design QA strategies that ensure consistent and predictable agent performance.
Format of the Course
- Interactive lecture and discussion.
- Hands-on debugging and evaluation exercises.
- Live-lab analysis of agent behaviors using observability tools.
Course Customization Options
- Customized reliability testing scenarios and industry-specific QA methods can be arranged upon request.
Mastra Ops & Production Engineering: Deploying and Scaling AI Agents
21 HoursMastra serves as an operational framework aimed at streamlining the deployment, scaling, and lifecycle management of AI agents within production environments.
This instructor-led live training, available either online or onsite, is designed for intermediate to advanced technical professionals who need to reliably and efficiently operationalize AI agents across their production systems.
Upon completing this training, participants will be able to:
- Deploy Mastra-based AI agents into controlled, production-grade environments.
- Scale agents both horizontally and vertically by leveraging platform-native primitives.
- Establish observability pipelines to monitor agent behavior and performance.
- Optimize runtime configurations to minimize latency, costs, and operational risks.
Course Format
- Interactive lectures and discussions.
- Hands-on exercises centered around real-world deployment scenarios.
- Live lab implementation within containerized and orchestrated environments.
Customization Options
- Topics, hands-on labs, or industry-specific scenarios can be customized upon request.
Mastra Workflow Automation & Multi-Agent Orchestration
21 HoursMastra is a framework designed to enable sophisticated workflow automation and coordination across multiple AI agents operating within distributed systems.
This instructor-led live training, available either online or onsite, targets intermediate-level practitioners aiming to design, orchestrate, and manage multi-agent workflows at scale.
Upon completion of this training, participants will acquire the skills necessary to:
- Design complex workflows leveraging Mastra’s orchestration capabilities.
- Coordinate multiple agents executing parallel or dependent tasks.
- Implement monitoring and debugging tools for workflow execution.
- Optimize orchestration logic to enhance reliability, throughput, and automation efficiency.
Format of the Course
- Interactive lecture and discussion.
- Hands-on workflow design and automation exercises.
- Practical implementation within a containerized live-lab environment.
Course Customization Options
- Customized automation scenarios, enterprise integrations, or workflow patterns can be provided upon request.
Managing Agent Workflows in Google Antigravity: Orchestration, Planning and Artifacts
14 HoursGoogle Antigravity serves as an agent-centric development platform, enabling the orchestration, supervision, and coordination of AI-driven coding and automation workflows.
Designed for intermediate professionals, this instructor-led training—available online or onsite—focuses on designing, managing, and optimizing multi-agent workflows within the Google Antigravity ecosystem.
By the end of this program, participants will be equipped to:
- Define agent responsibilities and establish orchestration pipelines through the Manager interface.
- Create and analyze Antigravity artifacts, such as task lists, strategic plans, logs, and browser recordings.
- Apply verification strategies that ensure agent actions are both transparent and auditable.
- Enhance multi-agent collaboration to tackle complex development and operational challenges.
Course Delivery Format
- Interactive presentations combined with practical live demonstrations.
- Scenario-based exercises addressing real-world workflow complexities.
- Direct experimentation within an active Antigravity workspace.
Customization Availability
- For a tailored version of this curriculum, please reach out to discuss specific customization needs.
Testing & Verifying Agent-Driven Code: Quality Assurance in Antigravity
14 HoursAntigravity represents a modern framework designed to facilitate advanced agent-driven development processes.
This live, instructor-led session, available both online and on-site, is tailored for intermediate to advanced professionals seeking to validate, secure, and assess the outputs generated by AI agents operating within Antigravity-based environments.
By the end of this training, participants will have the capability to:
- Evaluate the precision and safety of code artifacts created by agents.
- Leverage structured methodologies to confirm the execution of agent tasks.
- Effectively analyze browser recordings and track agent activities.
- Implement QA and security standards to guarantee the stability of agent workflows.
Course Structure
- Guided technical discussions and briefings led by the instructor.
- Practical exercises centered on verifying authentic agent workflows.
- Hands-on testing and validation conducted within a controlled lab setup.
Customization Options
- Tailoring of scenarios, workflows, and testing examples can be arranged upon request.