Site Reliability Engineer

منذ 2 أيام

Abu Dhabi, الإمارات العربية المتحدة Talent Arabia دوام كامل

Urgent requirement for  Site Reliability Engineer -Brokerage Management is required for our banking clients in Abu Dhabi ,UAE

Candidate focuses on ensuring high availability, performance, and stability of critical applications including Prospero, Advent, and related PL/SQL based systems.--Must

Monitor, maintain, and improve availability, latency, and performance of Brokerage Management--Must

SRE/Ops Experience with monitoring tools like Datadog, Splunk, Grafana, AppDynamics. Good understanding of ITIL, incident, and problem management--Must

 Cloud exposure- AWS/Azure basics, especially for hybrid wealth deployments.--Must

 

Banking Domain --Must

 

 

Role Overview

We are looking for a Site Reliability Engineer with strong experience supporting Brokerage  Management platforms. The role focuses on ensuring high availability, performance, and stability of critical applications including Prospero, Advent, and related PL/SQL based systems. You will bridge development and operations, driving automation and proactive support for business-critical trading and portfolio systems.

 

Key Responsibilities

 

- System Reliability: Monitor, maintain, and improve availability, latency, and performance of Brokerage Management applications like Prospero or  Advent

- Production Support: Provide L2/L3 technical support for Advent suite, Prospero, and backend PL/SQL processes. Handle incidents, root cause analysis, and permanent fixes

- Database & Batch Jobs: Troubleshoot PL/SQL stored procedures, packages, triggers, and batch jobs. Optimize queries and ensure smooth EOD/BOD cycles

- Automation & Tooling: Develop automation scripts for monitoring, deployment, and incident response using Python, Shell, or Ansible. Reduce manual toil

- Incident Management: Own major incidents end-to-end. Drive war rooms, coordinate with app vendors, infra, and business teams. Ensure SLAs are met

- Capacity & DR: Perform capacity planning, failover testing, and DR drills for wealth platforms

- Change Management: Review releases, assess risk, and ensure zero-downtime deployments for critical wealth applications

- Documentation: Maintain runbooks, SOPs, and knowledge base for support and escalation workflows

 

Required Skills & Qualifications

 

- Domain: 5+ years supporting Brokerage Management applications. 

- Database: Strong PL/SQL – ability to debug complex packages, performance tune, and analyze logs

- SRE/Ops: Experience with monitoring tools like Datadog, Splunk, Grafana, AppDynamics. Good understanding of ITIL, incident, and problem management

- Scripting: Python, Shell, or PowerShell for automation

- Infra: Working knowledge of Windows/Linux servers, job schedulers like Control-M/Autosys, and middleware

- Soft Skills: Excellent communication to work with traders, portfolio managers, and business users. Strong analytical and ownership mindset