AI Quality

منذ يوم

Abu Dhabi, الإمارات العربية المتحدة Tech Junction Ltd دوام كامل

Position Summary:

Halian is seeking an experienced, highly technical, and accomplished AI Quality & Reliability Engineer (QA/SRE) with 6+ years of professional experience to join our banking industry client on a full-time, permanent basis in Abu Dhabi, UAE. In this specialized engineering role at the intersection of AI, automated quality assurance, cloud engineering, and site reliability engineering (SRE), you will spearhead driving the quality, performance, resilience, and reliability of enterprise cloud-based AI solutions and LLM-driven applications. You will work closely with software architects, machine learning engineers, and banking IT operations squads to establish banking-grade production environments. Ideal candidates bring a robust academic degree in computer science or engineering, deep practical mastery of Python testing frameworks, cloud platforms (Azure/AWS), chaos engineering, and immediate availability for deployment in Abu Dhabi.

Detailed Job Description:

As an AI Quality & Reliability Engineer (QA/SRE) at Halian in Abu Dhabi, UAE, you will take full ownership of testing AI/LLM systems, validating RAG pipelines, managing cloud observability, and executing rigorous site reliability engineering workflows. Your day-to-day responsibilities encompass developing automated test suites utilizing PyTest, Selenium, and Robot Framework, performing advanced load testing and chaos engineering using K6 and JMeter, and monitoring system telemetry across Grafana, Prometheus, Azure Monitor, and CloudWatch. You will ensure strict compliance with banking security standards, evaluate prompt engineering accuracy, and maintain CI/CD infrastructure validation pipelines. Working in a fast-paced banking technology environment, you will drive operational resilience and AI excellence.

Key Responsibilities:

  • Design, build, and execute comprehensive automated testing frameworks for cloud-based AI solutions and LLM applications.
  • Spearhead site reliability engineering (SRE) initiatives, observability monitoring, and incident response runbooks in banking environments.
  • Conduct rigorous load testing, stress testing, chaos engineering, and scalability evaluations using K6 and JMeter.
  • Perform specialized AI/LLM testing, prompt engineering validation, and Retrieval-Augmented Generation (RAG) evaluation.
  • Configure, maintain, and analyze infrastructure monitoring telemetry via Grafana, Prometheus, Azure Monitor, and AWS CloudWatch.
  • Implement continuous integration and continuous deployment (CI/CD) pipelines with automated infrastructure validation checks.
  • Validate banking-grade security baselines, data compliance protocols, and high-availability production architectures.
  • Write robust test automation scripts utilizing Python, PyTest, Selenium, and Robot Framework.
  • Collaborate proactively with machine learning developers, cloud architects, and financial stakeholders.
  • Maintain comprehensive technical documentation, test coverage reports, failure analysis logs, and SRE post-mortem reviews.

Required Qualifications & Skills:

  • Bachelor’s or Master’s degree in Computer Science, Software Engineering, Information Technology, or a related technical discipline.
  • Minimum 6+ years of professional experience spanning software quality assurance, cloud engineering, and site reliability engineering (SRE).
  • Strong programming proficiency in Python and hands-on expertise with testing frameworks (PyTest, Selenium, Robot Framework).
  • Practical hands-on experience deploying and managing enterprise workloads on Azure and/or AWS cloud platforms.
  • Proven experience conducting load testing, performance tuning, and chaos engineering (K6, JMeter).
  • Specialized expertise in AI/LLM testing, prompt evaluation, and RAG pipeline accuracy verification.
  • Solid operational knowledge of observability platforms including Grafana, Prometheus, Azure Monitor, and CloudWatch.
  • Familiarity with security, compliance, and auditing requirements in banking-grade production environments.
  • Professional availability for full-time employment in Abu Dhabi, UAE.

Nice-to-Have Skills:

  • Professional cloud certifications (AWS Certified DevOps Engineer, Microsoft Certified: Azure DevOps Engineer Expert).
  • Hands-on experience with infrastructure as code (Terraform, Ansible) and Kubernetes container orchestration.
  • Prior working experience in financial services, tier-1 banking institutions, or enterprise IT consultancy firms.
  • Advanced knowledge of machine learning operations (MLOps) pipelines and model deployment governance.
  • Bilingual proficiency in Arabic and English.

Recruitment Pro Tip:

When applying for specialized AI SRE and QA roles in banking, ensure your resume explicitly highlig