- Home
- Jobs
- DALIL INFORMATION TECHNOLOGY
- Infrastructure & Platform Operations Engineer
Infrastructure & Platform Operations Engineer
Job Description Roles & Responsibilities bout the Platform Project: Morocco Steady-State Operations Environment: Government-grade, on-premises bare-metal infrastructure. Three security-zone Kubernetes clusters (DMZ, Core, Red Zone) underpinned by a Fortinet security fabric (FortiGate 901G HA pair, FortiSwitch, FortiManager, FortiPAM, FortiWeb WAF, FortiADC, FortiAnalyzer) and Dell PowerEdge R670 compute nodes with Ceph distributed storage. Backup is delivered via Dell PowerProtect. Operations model: The platform has been deployed and accepted. These roles cover Day 2 steady-state operations — monitoring, incident response, change management, backup validation, and continuous improvement. No deployment or project delivery responsibilities. Role Overview The Infrastructure & Platform Operations Engineer is the primary Day 2 owner of the compute, Kubernetes, storage, and backup layer of the platform. You maintain cluster health, respond to node and workload incidents, manage the backup and recovery lifecycle, and ensure the observability pipeline is accurate and actionable. You work alongside the Network & Security Operations Engineer, who owns the Fortinet fabric. Together, the two roles provide full-stack operational coverage. When a platform incident has a network/security dimension — for example, a pod failing to reach an external endpoint — you collaborate across the boundary and escalate correctly. Day-to-Day Responsibilities Kubernetes Cluster Operations Monitor health of all three Kubernetes clusters (DMZ, Core, Red Zone): node readiness, control plane component status, etcd health, and API server responsiveness. Investigate and resolve pod scheduling failures, CrashLoopBackOff events, OOMKill incidents, and persistent volume claim binding issues. Manage node maintenance: cordon, drain, perform maintenance (firmware, OS patch, hardware swap), and return nodes to service. Monitor and enforce namespace resource quotas, limit ranges, and network policy posture — flag and remediate any drift from the accepted baseline. Manage RBAC: onboard/offboard cluster users and service accounts, review permissions quarterly, maintain least-privilege posture. Renew and rotate Kubernetes cluster certificates, ingress TLS certificates, and any expiring secrets before expiry. Coordinate with the application team on deployment readiness: validate ingress routing, service endpoints, and resource availability before application releases. Ceph Storage Operations Monitor Ceph cluster health dashboard daily: OSD status, PG states, replication factor compliance, and capacity utilisation against defined thresholds. Respond to OSD failure alerts: investigate disk health (SMART data), replace failed OSDs, and monitor rebalancing to completion before any second failure. Monitor Ceph pool capacity and raise alerts when utilisation exceeds planned thresholds — coordinate capacity planning reviews with the client. Validate PVC provisioning for new workloads: confirm correct pool assignment, access mode, and reclaim policy. Dell Compute & Node Operations Monitor Dell PowerEdge R670 nodes via iDRAC: hardware alerts (PSU, DIMM, NIC, disk, thermal), BIOS event logs, and remote console readiness. Manage firmware lifecycle: track available BIOS, iDRAC, NIC, and PERC firmware updates; schedule and apply via approved maintenance windows. Maintain Linux OS baseline on all Kubernetes nodes: security patches, kernel updates, containerd and kubelet version alignment with the cluster compatibility matrix. Maintain hardware inventory register: track serial numbers, firmware versions, and warranty status per node. Backup & Recovery Operations Monitor Dell PowerProtect backup jobs daily: confirm job completion, investigate and remediate failures, and maintain the backup success rate SLA. Execute and document periodic restore tests (minimum quarterly): restore at least one VM and one data scenario, record restore time, verify data integrity. Manage backup policy changes through change control: scope adjustments, retention updates, and schedule changes. Maintain the restore runbook with current, tested procedures — update after every restore test. Observability & Alerting Maintain the observability pipeline: ensure log collection, metrics scraping, and alerting are functioning across all nodes and clusters. Triage platform-level alerts: distinguish noise from genuine incidents, tune alerting thresholds, and escalate per the defined severity model. Produce weekly platform health summaries: cluster utilisation, backup status, storage capacity, and outstanding incidents. Desired Candidate Profile Must-Have Requirements Kubernetes operations experience in a production environment: node management, workload troubleshooting, RBAC, PVC/storage, certificate rotation. CKA required. Linux systems administration: OS patching, kernel management, systemd, journald, process and disk troubleshooting. Ceph or comparable distributed storage operations: OSD monitoring, failure response, capacity management. Backup platform operations (Dell PowerProtect, Veeam, or equivalent): job monitoring, failure remediation, restore testing and evidence production. Bare-metal server operations: iDRAC or iLO monitoring, firmware management, hardware fault response. Observability tooling: experience with Prometheus, Grafana, Loki, or equivalent for cluster-level monitoring. On-site availability in Morocco on a permanent or long-term basis. Good-to-Have CKS (Certified Kubernetes Security Specialist) — strongly preferred given the government security classification of the platform. Dell PowerEdge-specific experience (PowerEdge R670 or R-series equivalent, iDRAC 9/10, OMSA). Dell PowerProtect-specific experience vs generic backup platform knowledge. Kubernetes upgrade experience: minor version upgrades of control plane and worker nodes on a live cluster. Familiarity with Fortinet/network layer concepts (enough to collaborate on cross-domain incidents with the Network & Security Ops Engineer). French or Arabic language capability. Employment Type Full Time Company Industry IT - Software Services Department / Functional Area IT Hardware SupportIT Hardware Repair & Maintenance Keywords Network EngineeringInfrastructurePlatform Operations EngineerSite Reliability EngineerDevOps EngineerCloud Operations SpecialistAutomationAutomation Engineer Get real-time job updates only on our App
Ready to apply?
You are viewing this role on JobSphere AI. Applications are completed on the original employer / source website.
Apply NowOpens the employer's site in a new tab
- CompanyDALIL INFORMATION TECHNOLOGY
- LocationMorocco
- CategoryCybersecurity
- SourceNaukrigulf
- Listed1 month ago
Related Cybersecurity jobs
Senior Cloud Engineer
Key Responsibilities Provision and administer approved cloud resources, platform services and access configurations. Monitor workloads, performance…
Sr. DevOps Engineer
Our Mission is to Simplify Life. We are looking to Simplify and automate complex decision-making for customer centric industries, like Utilities, Financial…
Senior DevOps Engineer
The Senior DevOps Engineer executes and refines the DevOps strategy and is instrumental in implementing automation and integration across IT Operations…
Commercial Drone Pilot & Geospatial Data Analyst
Execute complex aerial data acquisition missions, ensuring adherence to flight plans, safety protocols, and regulatory requirements. Operate advanced drone…