SRE
The SRE group in the Agentic Transformation Platform (ATP) sidebar holds 5 assessment pages. Each page is run by the same agent, Dana (site reliability engineering agent). Use this page to choose the right assessment and prepare the evidence it asks for.
Before you start
- Every page in this group uses the same assessment workspace. To learn how to upload evidence, run an analysis and ask follow-up questions, read Run an assessment.
- Upload metadata and exports only. Each page's Instructions panel says what not to upload.
Open a page in this group
- In ATP, open the menu at the left of the navigation bar.
- Expand SRE.
- Select the page you need.
When a page from this group is open, the navigation bar shows the group's pages as tabs, so you can move between them. Pages that don't fit are under More.
Pages in this group
| Page | What it covers |
|---|---|
| Critical Workload Ops | Critical workload operations and management. |
| Critical Incident Hypercare | Critical incident response and hypercare support. |
| Workload Operational Optimization & Auto. | Workload optimization and automation. |
| Chaos Engineering | Chaos engineering and resilience testing. |
| SRE Observability | Site reliability engineering observability. |
Critical workload operations
Select Critical Workload Ops to open this page. The page header shows the badge Assessment workspace and the line Expert Consultation with Dana. The sidebar describes it as: Critical workload operations and management.
The Instructions panel is titled How to Generate Your Critical Workload Operations Report. It says:
To generate your critical workload operations report, Dana only needs metadata (workload characteristics and performance metrics) — not business data. Upload raw exports from your workload monitoring systems, no JSON wrapping or special formatting is required.
It lists the evidence to upload:
- Workload Inventory: Critical workload definitions, application dependencies, customer volume data, tier classifications, and resource requirements.
- Performance Data: SLA performance records, throughput metrics, latency data, error rates, and availability trend data.
- SLA Requirements: SLA definitions, SLO targets, SLI measurements, compliance records, and breach history data.
- Incident Impacts: Workload-related incident records, customer impact assessments, downtime data, and recovery time measurements.
- Capacity & Scaling: Capacity planning records, auto-scaling configurations, resource headroom data, and peak demand profiles.
The panel also shows this note:
Important: Do not upload highly sensitive infrastructure configurations — Dana analyzes workload characteristics and performance data, not infrastructure configurations.
Critical incident hypercare
Select Critical Incident Hypercare to open this page. The page header shows the badge Assessment workspace and the line Expert Consultation with Dana. The sidebar describes it as: Critical incident response and hypercare support.
The Instructions panel is titled How to Generate Your Critical Incident Hypercare Report. It says:
To generate your critical incident hypercare report, Dana only needs metadata (incident characteristics and response patterns) — not business data. Upload raw exports from your incident management systems, no JSON wrapping or special formatting is required.
It lists the evidence to upload:
- Incident History: Past incidents, severity classifications, SRT data, post-mortem actions, and SLA violation records.
- Response Procedures: Escalation procedures, on-call notification configurations, communication templates, and response runbooks.
- Post-Mortem Analysis: Post-mortem reports, action item records, root cause findings, contributing factor data, and learning outcome documentation.
- On-Call Configuration: On-call schedules, escalation chains, rotation structures, coverage analysis, and paging integration data.
- Monitoring Integrations: Alert configurations, dependency mapping data, monitoring setup records, and alert correlation rules.
The panel also shows this note:
Important: Do not upload highly sensitive infrastructure configurations — Dana analyzes incident patterns and response workflows, not infrastructure configurations.
Workload operational optimization and automation
Select Workload Operational Optimization & Auto. to open this page. The page header shows the badge Assessment workspace and the line Expert Consultation with Dana. The sidebar describes it as: Workload optimization and automation.
The Instructions panel is titled How to Generate Your Workload Optimization Report. It says:
To generate your workload optimization report, Dana only needs metadata (resource utilization and automation configurations) — not business data. Upload raw exports from your cloud management platforms, no JSON wrapping or special formatting is required.
It lists the evidence to upload:
- Resource Profile: CPU, memory, network, and disk utilization data, workload bottleneck reports, and scaling pattern records.
- Automation Configurations: Automation rule definitions, runbook automation scripts, workflow configurations, and self-healing policy data.
- Optimization Rules: Rightsizing recommendations, scaling policy definitions, scheduling configurations, and resource optimization findings.
- Automation Rules: Event-driven automation triggers, remediation playbooks, auto-scaling thresholds, and orchestration workflow data.
- Cost Data: Resource cost efficiency metrics, optimization savings data, waste identification records, and cost-per-workload data.
The panel also shows this note:
Important: Do not upload highly sensitive infrastructure configurations — Dana analyzes workload optimization patterns and automation configurations, not infrastructure configurations.
Chaos engineering
Select Chaos Engineering to open this page. The page header shows the badge Assessment workspace and the line Expert Consultation with Dana. The sidebar describes it as: Chaos engineering and resilience testing.
The Instructions panel is titled How to Generate Your Chaos Engineering Report. It says:
To generate your chaos engineering report, Dana only needs metadata (system architecture and failure scenarios) — not business data. Upload raw exports from your resilience testing frameworks, no JSON wrapping or special formatting is required.
It lists the evidence to upload:
- System Architecture: Architecture diagrams, system topology data, failure domain definitions, circuit breaker configurations, and dependency maps.
- Experiment History: Past chaos experiment records, hypothesis definitions, blast radius data, experiment outcomes, and finding summaries.
- Failure Scenarios: Failure mode catalogs, fault injection configurations, scenario playbooks, and resilience hypothesis documentation.
- Resilience Findings: Weakness discovery records, remediation action items, resilience improvement tracking, and re-test outcome data.
- Observability Data: Monitoring configurations, alerting setups, SLO impact measurements, and observability coverage during experiments.
The panel also shows this note:
Important: Do not upload highly sensitive infrastructure configurations — Dana analyzes resilience patterns and failure scenarios, not live infrastructure configurations.
SRE observability
Select SRE Observability to open this page. The page header shows the badge Assessment workspace and the line Expert Consultation with Dana. The sidebar describes it as: Site reliability engineering observability.
The Instructions panel is titled How to Generate Your SRE Observability Report. It says:
To generate your SRE observability report, Dana only needs metadata (monitoring configurations and telemetry patterns) — not business data. Upload raw exports from your observability platforms, no JSON wrapping or special formatting is required.
It lists the evidence to upload:
- Monitoring Coverage: Monitoring service inventories, metrics collection configurations, dashboard setups, SLO/SLI monitoring data, and coverage gap reports.
- Telemetry Stack: Logging platform configurations, distributed tracing setups, metrics pipeline data, and telemetry integration records.
- Alert Configuration: Alerting rules, threshold definitions, routing policies, on-call integration data, and escalation configuration records.
- Observability Gaps: Coverage gap analysis, blind spot records, missing instrumentation reports, and observability maturity assessments.
- Data Retention & Cost: Log retention policies, metrics sampling configurations, storage cost data, and observability cost optimization records.
The panel also shows this note:
Important: Do not upload highly sensitive infrastructure configurations — Dana analyzes observability configurations and monitoring patterns, not infrastructure configurations.