Expert Site Reliability Engineer
Job Description Roles & Responsibilities To drive the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices. Main Duties and Responsibilities: Define and implement advanced reliability engineering practices across critical technology services. Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets. Design automation to reduce manual operational activities and improve system resilience. Develop and enhance monitoring, observability, alerting, and incident detection capabilities. Lead technical analysis and resolution of complex production incidents. Conduct root-cause analysis and drive permanent corrective and preventive actions. Design solutions to improve system availability, scalability, capacity, and disaster resilience. Identify reliability risks and recommend architectural and engineering improvements. Drive performance engineering and capacity planning for critical services. Provide advanced technical guidance and mentorship on SRE practices. Promote automation and engineering approaches that reduce operational toil and improve service reliability. Desired Candidate Profile Bachelor s degree in Computer Science, Software Engineering, IT, or a related field . 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles. Strong experience in cloud platforms, Kubernetes, and production environments . Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics . Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform) . Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA) . Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery . Experience driving reliability improvements and reducing operational toil through automation . Strong analytical, problem-solving, and technical leadership skills. Experience in Banking, FinTech, or Payment environments is preferred. Company Industry IT - Software Services Department / Functional Area IT Software Keywords Expert Site Reliability Engineer Get real-time job updates only on our App
Ready to apply?
You are viewing this role on JobSphere AI. Applications are completed on the original employer / source website.
Apply NowOpens the employer's site in a new tab
- CompanyTAWANTECH
- LocationRiyadh, Saudi Arabia
- CategoryDevOps
- SourceNaukrigulf
- Listed3 days ago
Related DevOps jobs
FINIQ Consultant
Write, debug, and tune complex MS SQL queries, stored procedures, and triggers. Ensure database integrity, high availability, and optimal performance. Build…
FINIQ Application Developer & Support Engineer (.NET / MS SQL)
We are looking for a hands-on Application Developer & Support Engineer to own and optimize our FINIQ core platform environment. This role blends backend…
Agricultural Engineer
Conduct feasibility studies and cost analysis for new agricultural projects to ensure economic viability. Collaborate with agronomists and farmers to implement…
Senior Leasing Manager | Retail
The Retail Leasing Manager will be part of the Al-Futtaim Retail Leasing team, tasked with managing the retail leasing portfolio across the MENA region. The…