Site Reliability Engineer/ Expert/ Specialist (DevOps Eng)
Job Description Roles & Responsibilities Define, build, and maintain support systems to ensure high availability and performance. Handle complex cases for the PSO. Implement automation for system provisioning, self-healing, auto-recovery, deployment, and monitoring. Perform incident response and root cause analysis (RCA) for critical system failures. Monitor system performance and establish Service-Level Indicators (SLIs) and Service-Level Objectives (SLOs). Collaborate with Development and Operations to integrate reliability best practices, including zero-downtime architecture. Proactively identify and remediate performance issues. Work closely with Product T&E, ICE, and Service Architects for new product productization as SGS technical expert. Coordinate with internal and external stakeholders to improve service performance and ensure high availability. Ensure Operations readiness to support new products. Accountable within SGS for in-scope product availability and performance. Problem Management Conduct thorough problem investigations and root cause analyses to diagnose recurring incidents and service disruptions. Coordinate with Incident Management teams and collaborate with PSOs and Engineering/Product teams to implement permanent solutions. Monitor effectiveness of problem resolution activities and provide regular reporting to ensure continuous improvement. Event Management Define, build, and maintain an event catalog specifying active events, thresholds, and remediation actions; optimize it for efficiency. Develop event response protocols, provide training, and ensure efficient incident handling. Customer & Operational Support Collaborate with Customer Success Managers to implement initiatives that enhance customer satisfaction and retention. Prepare reports, documentation, and communication materials covering customer metrics, updates, and product changes. Identify and implement improvements in internal processes and workflows. Contribute to knowledge management resources such as FAQs and training materials. Data Steward Responsibilities Implement data governance policies defined by the Data Owner and ensure adherence to standards. Monitor data quality, consistency, and compliance on an ongoing basis. Act as a Subject Matter Expert (SME) for data within the assigned area, providing guidance and answering queries. Desired Candidate Profile Bachelor s degree in Computer Science, Information Technology, Engineering, or a related field. 5+ years of experience in IT operations, service management, or infrastructure management, including roles such as Site Reliability Engineer, Problem Manager, or DevOps Manager. Proven experience managing high-availability systems and ensuring operational reliability. Extensive experience in root cause analysis (RCA), incident management, and developing permanent solutions for recurring service disruptions. Hands-on experience with CI/CD pipelines, automation, system performance monitoring, and infrastructure as code (IaC). Strong background in collaborating with cross-functional teams (Development, Operations, Engineering, etc.) to improve operational processes and service delivery. Experience managing deployments, conducting risk assessments, and optimizing event and problem management processes. Familiarity with cloud technologies, containerization, and scalable architectures, including zero-downtime deployment strategies. Technical Skills (Must-to-Have): Strong AKS & On prem K8s skills and experience. Scripting (Ansible & Bash, Python - combination of anything would be great), Automation, CI/CD pipeline, Terraform exposure, Azure (or) AWS skill. Basic DB skills. Strong problem-solving skills & quick learner. SRE mindset. Company Industry IT - Software Services Department / Functional Area IT Software Keywords Site Reliability Engineer/ Expert/ Specialist (DevOps Eng) Get real-time job updates only on our App
Ready to apply?
You are viewing this role on JobSphere AI. Applications are completed on the original employer / source website.
Apply NowOpens the employer's site in a new tab
- CompanySITA
- LocationEgypt
- CategoryDevOps
- SourceNaukrigulf
- Listedjust now
Related DevOps jobs
Maintenance Supervisor
The Maintenance Supervisor is responsible for the smooth operation of all equipment, supervising all kinds of maintenance activities such as preventive…
Lead Planning Engineer Mechanical
Position Summary: The Lead Planner – Mechanical is accountable for planning and scheduling, cost control, invoicing, progress reporting, client coordination…
SQL & SSRS Report Developer – Banking Domain
Design, develop, and maintain enterprise-level SSRS reports for banking and regulatory reporting. Develop interactive Power BI dashboards for business and…
Test Lead Automation
du poste En tant que Test Lead Automation , vous serez responsable de la strat gie d automatisation des tests et de l encadrement des quipes QA. Vous…