Senior Site Reliability Engineer
Job Overview
Connect to your opportunity Step into a highly impactful role where you will help us build and operate our brand-new centralized command center for public finance in Cairo. As a Site Reliability Engineer (SRE) you will ensure our enterprise infrastructure remains highly available, scalable, and fully optimized. You will go beyond basic troubleshooting to proactively manage and resolve complex technical challenges across diverse cloud environments. You will compose our support level 2 and 3 team for a national scale product that has been deployed in 3 GCC counties, the knowledge of the product will be transferred in training sessions. Monitor and proactively troubleshooting to prevent outage and service interruption Investigate and resolve infrastructure and application incidents coming from support level 1 or from the observability tools. Root cause analysis and problem management, conduct post incident reviews and update runbooks and knowledge base articles Performance and capacity management, monitor cloud utilization and optimize multi-cloud environments to ensure continuous high system reliability and uptime. Automation and efficiencies, develop scripts to generate effort reductions and standardization, maintain infrastructure as a code templates and repeatable procedures. Security and compliance, assist with client s audits and incidents responses Back up and recovery, validate and monitor back up jobs, test recovery procedures Document and knowledge sharing, maintain accurate documents for configuration, train support level 1 to increase first call resolution Customer communication and SLAs/KPIs tracking and reporting Vendor and tool liaison, coordinate with cloud providers and track open tickets to ensure timely resolution On call shift responsibilities, participate in on call rotations
Desired Candidate Profile
Connect to your skills and professional experience Analytical Thinking enables you to break down complex system behaviors to find the direct root cause of infrastructure issues. Adaptability allows you to seamlessly transition between different cloud platforms and tools in a fast-paced environment. Collaborative Problem-Solving ensures you work effectively with specialized engineering teams to build resilient, long-term technical solutions. Essentials: Extensive, highly advanced expertise in multi-cloud operations and Site Reliability Engineering, reflecting senior-level industry tenure. Demonstrated multi-cloud expertise with hands-on technical capabilities across platforms. Possession of at least two of the following certifications (Note: if you hold two, Deloitte will train and support you in acquiring the third): Amazon Web Services (AWS) Certification Google Cloud Platform (GCP) Certification Oracle Cloud Infrastructure (OCI) Certification Experience with observability tools Experience with integration flows Experience with DevOps tools Familiarity with foundational Site Reliability Engineering (SRE) practices and automation frameworks. 3 to 7 years minimum prior experience operating within an enterprise-level IT command center.
Ready to apply?
You are viewing this role on JobSphere AI. Applications are completed on the original employer / source website.
Apply NowOpens the employer's site in a new tab
- CompanyDeloitte
- LocationEgypt
- CategoryCybersecurity
- SourceNaukrigulf
- Listed1 month ago
Related Cybersecurity jobs
Application Support – Enterprise Loyalty
عربي JOBS SERVICES LOGIN REGISTER Interview Q&As EMPLOYERS?
UAE National | Service Advisor | Toyota/Lexus | Sharjah
عربي JOBS SERVICES LOGIN REGISTER Interview Q&As EMPLOYERS?
Mobile Developer
Design, develop, and maintain high-quality Android applications using Kotlin Build and optimize map-based features (e.g., location tracking, routing, markers…
Senior iOS Developer
Client Expectation : We are seeking a Senior iOS Developer to lead the design and development of native applications using SwiftUI. You will architect scalable…