Site Reliability Engineer

منذ يوم

Abu Dhabi, Abu Dhabi Emirate, الإمارات العربية المتحدة Innovations Global دوام كامل ‏350,000 € - ‏550,000 € عقد
Position Summary Innovations Global is urgently seeking an experienced, proactive Site Reliability Engineer (SRE) to champion the continuous availability, scalability, and performance of mission-critical banking and financial technology platforms in Abu Dhabi, United Arab Emirates. Operating in a full-time onsite capacity for a 1-year renewable contract, you will bridge the divide between software engineering and systems operations to ensure enterprise infrastructure resilience. With a minimum of five years of hands-on experience, you will spearhead observability, drive end-to-end automation, lead rapid incident recovery, and enforce stringent Site Reliability Engineering principles across hybrid cloud environments. This position provides an exceptional opportunity for a technical engineering professional to make a transformative impact on premier banking infrastructure in the UAE. Detailed

Job Description
As a Site Reliability Engineer (SRE) at Innovations Global supporting a premier enterprise banking environment in Abu Dhabi, you will hold operational accountability for the health, availability, performance, and efficiency of high-throughput transactional applications and distributed systems. You will collaborate closely with cross-functional software engineering teams, DevOps squads, database administrators, and cyber security teams to establish resilient deployment pipelines and maintain robust production ecosystems. Your core technical mandate involves architecting and administering enterprise Linux and Unix servers, orchestrating microservices utilizing Docker and Kubernetes, and managing scalable workloads across leading cloud platforms (AWS, Azure, or GCP). You will implement proactive monitoring and observability frameworks, define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs), and eliminate operational toil through Python and Bash automation scripting. Additionally, you will direct incident response triage, execute root cause analysis (RCA), optimize CI/CD release workflows, and troubleshoot complex TCP/IP enterprise networking bottlenecks. This role requires rigorous diagnostic discipline, deep systems acumen, and the capability to maintain zero-downtime reliability within a fast-paced financial services environment.

Key Responsibilities
- Ensure the maximum reliability, availability, performance, and operational efficiency of enterprise banking applications and underlying cloud infrastructure.
- Implement, tune, and manage full-stack monitoring, telemetry, and observability platforms to capture proactive operational insights and real-time alerts.
- Eliminate operational toil by designing, building, and maintaining automated workflows and operational scripts using Python, Bash, or Shell scripting.
- Lead rapid incident management, triage system outages, conduct detailed root cause analysis (RCA), and implement permanent corrective remediations.
- Deploy, configure, manage, and scale containerized application workloads utilizing Kubernetes clusters and Docker environments.
- Administer, optimize, and maintain high-performance enterprise Linux and Unix server operating systems in Tier-compliant hosting environments.
- Establish, monitor, and report on core reliability engineering metrics, including Service Level Agreements (SLAs), Service Level Objectives (SLOs), and Error Budgets.
- Optimize and support automated CI/CD deployment pipelines, ensuring secure, reliable, and frictionless software releases into production. Required Qualifications & Skills
- Minimum 5+ years of dedicated professional experience in Site Reliability Engineering (SRE), DevOps, or Linux/Unix Systems Engineering.
- Strong technical expertise in Linux and Unix system administration, operating system internals, kernel parameters, and performance tuning.
- Hands-on expertise deploying, managing, and maintaining enterprise workloads on major public cloud platforms such as AWS, Microsoft Azure, or GCP.
- Demonstrated technical proficiency with containerization and orchestration platforms, specifically Docker and Kubernetes.
- Extensive experience implementing modern monitoring, logging, and observability tools (e.g., Prometheus, Grafana, ELK Stack, Datadog, or Dynatrace).
- Proven proficiency in scripting and automation utilizing Python, Bash, or Shell scripting for operational workflows.
- Solid understanding of core TCP/IP networking, routing, DNS, load balancing, SSL/TLS, and enterprise perimeter network security.
- Demonstrated experience in incident management, blameless post-mortem investigations, and reliability engineering practices (SLA/SLO/SLI frameworks). Nice-to-Have Skills
- Prior hands-on engineering experience within the banking, financial services, fintech, or large-scale transactional enterprise domains.
- Recognized professional certifications such as AWS Certified Solutions Architect, Azure Solutions Architect Expert, or