- Home
- Jobs
- TakaMoL Holding
- Site Reliability Engineering Officer
Site Reliability Engineering Officer
Job Description Roles & Responsibilities Provide support for application incidents across digital platforms, working closely with Platform Engineering, Application Development, and customer support teams to ensure timely resolution according to established SLAs and escalation procedures. Operate and monitor the Elastic Observability stack including Elasticsearch cluster health, Kibana, Fleet Server, APM Server, and Elastic Agent deployed and managed via ECK on OKE. Assist with day-to-day Elasticsearch operations such as index lifecycle management (ILM), snapshot lifecycle management (SLM), data tier housekeeping (hot, warm, cold, frozen), and capacity monitoring. Troubleshoot telemetry ingestion issues across logs, metrics, traces, and synthetic monitors, ensuring consistent data collection from all platforms. Maintain and update Kibana dashboards, alerting rules, and saved objects under the guidance of the SRE Manager. Perform root cause analysis and participate in blameless post-incident reviews to improve system reliability and reduce recurrence. Collaborate with Platform Engineering to automate repetitive tasks, improve deployment pipelines, and enhance observability coverage using Terraform, Helm charts, and scripting. Develop and maintain support documentation, runbooks, and knowledge base articles aligned to standardized incident response procedures. Manage and prioritize incidents and requests via the ticketing system (Jira/ServiceNow), ensuring all incidents, requests, and resolutions are documented in the service management system. Participate in an on-call rotation and help reduce operational toil through automation and tooling. Monitor and report on key performance metrics related to incident management, including mean time to detect (MTTD) and mean time to resolve (MTTR). Collaborate with cross-functional teams and vendor partners to improve overall system reliability, observability maturity, and security posture. Desired Candidate Profile Bachelor s degree in Computer Science, IT, Engineering, or related field (or equivalent experience). 1 3 years of experience in IT operations, system administration, application support, DevOps, or SRE. Familiarity with Observbility tools such as Elastic Stack (Elasticsearch, Kibana, etc.), including basic querying and dashboard usage. Knowledge of Linux systems and scripting (Bash, Python, or Go). Understanding of monitoring, logging, and alerting concepts. Experience with ITSM tools (ServiceNow, Jira, Zendesk) and ITIL practices. Strong grasp of incident, problem, and change management. Basic experience with cloud native enviroments and containers such as Docker and Kubernetes. Strong critical thinking, troubleshooting, and communication skills. Company Industry ConsultingManagement ConsultingAdvisory Services Department / Functional Area Engineering Keywords Site Reliability Engineering Officer Get real-time job updates only on our App
Ready to apply?
You are viewing this role on JobSphere AI. Applications are completed on the original employer / source website.
Apply NowOpens the employer's site in a new tab
- CompanyTakaMoL Holding
- LocationRiyadh, Saudi Arabia
- CategoryCybersecurity
- SourceNaukrigulf
- Listed1 week ago
Related Cybersecurity jobs
Senior Support Engineer Digital Channels
The Senior Support Engineer Digital Channels is responsible for end-to-end production support, maintenance and enhancement of the Bank's mobile banking…
Technical Project Manager
We are looking for an experienced Technical Project Manager to lead the planning, execution, and successful delivery of software development and digital…
Registered Midwife
Midwives are recognized as a responsible and accountable professional who works in partnership with women to give the necessary support, care and advice during…
Senior Consultant PSP
This is a position focused on power systems analysis. The skill shall be applied to a large span of issues such as renewables integration, grid code compliance…