The Cloud Resiliency Specialist will support customers through focused technical engagements aimed at improving the reliability, continuity, and recoverability of their cloud environments. The specialist will combine deep technical expertise in Microsoft Azure with a strong command of disaster recovery, business continuity, and operational resiliency processes, including the design and evaluation of Major Incident Response Plans (MIRPs).
This role requires the ability to assess both cloud infrastructure and resiliency processes end-to-end, identifying risks, aligning to best practices, and delivering clear, actionable recommendations that strengthen the customer’s ability to withstand and recover from disruptions.
Assess the resiliency posture of customer environments from both technical and operational perspectives.
Evaluate the effectiveness of disaster recovery (DR) plans, business continuity (BC) strategies, and Major Incident Response Plans (MIRPs).
Conduct in-depth technical analysis of cloud infrastructure, application dependencies, observability configurations, and failover capabilities.
Review and enhance MIRPs, including escalation protocols, incident communication plans, recovery sequencing, and impact containment strategies.
Identify risks, vulnerabilities, and single points of failure across workloads and operational processes.
Recommend improvements aligned with the Azure Well-Architected Framework, SRE principles, and ITIL practices.
Engage customer teams to understand RTO/RPO targets, recovery workflows, and coordination models for major incidents.
Deliver professional, customer-facing documentation summarizing technical findings and process maturity recommendations.