CUDA Performance Engineer

منذ 23 ساعات

Al Ruways Industrial City, Abu Dhabi Emirate, الإمارات العربية المتحدة YO IT Consulting دوام كامل ‏279,000 د.إ.‏ - ‏502,000 د.إ.‏ عقد

CUDA Performance Engineer - GPU Kernels Job Snapshot Role: CUDA Performance Engineer - GPU Kernels Location: Abu Dhabi Emirate, United Arab Emirates Industry: Computer Software Function: IT-Software Development Experience: Advanced experience in CUDA, C , GPU programming, and performance optimization Job Type: Contractor Position Overview CUDA Performance Engineer - GPU Kernels in Abu Dhabi Emirate, United Arab Emirates is a remote Computer Software opportunity for an experienced engineer specializing in CUDA development, C , GPU kernel profiling, and performance optimization. YO IT Consulting is hiring a technical specialist to investigate GPU workloads, improve kernel efficiency, develop GLSL and WebGPU shader logic, and deliver measurable performance gains across high-performance computing and accelerated software environments.

Job Details

Country: United Arab Emirates

City: Abu Dhabi Emirate

Industry: Computer Software

Function: IT-Software Development

Salary: 25000-45000 Estimated salary range based on similar jobs in the job city; please confirm the final offer with the employer.

Gender: Any

Candidate Nationality: Any

Job Type: Contractor

Role Context

The CUDA Performance Engineer will focus on understanding how GPU workloads behave at runtime and translating profiling evidence into targeted engineering improvements. The position requires strong awareness of GPU architecture, parallel execution, memory behavior, and kernel design so that inefficient code paths can be isolated and corrected. The engineer will also contribute to shader development and architecture discussions, helping improve the speed, scalability, and maintainability of performance-critical GPU software.

Key Responsibilities

  • Profile CUDA kernels and GPU workloads using NVIDIA Nsight or comparable performance-analysis tools.
  • Diagnose execution bottlenecks involving memory access, occupancy, synchronization, thread organization, or computational throughput.
  • Design targeted optimization strategies based on measurable runtime behavior.
  • Refactor CUDA and C code to improve efficiency while preserving readability and maintainability.
  • Tune kernel implementations to make more effective use of available GPU resources.
  • Evaluate memory-transfer patterns and reduce unnecessary host-device or device-memory overhead where practical.
  • Develop GPU shader logic using GLSL and WebGPU for compute-oriented or graphics-related workloads.
  • Benchmark alternative implementations and quantify performance improvements using repeatable testing methods.
  • Investigate performance regressions and determine whether issues originate in kernel design, memory utilization, workload structure, or architectural limitations.
  • Adapt optimization approaches for different GPU architectures when workload behavior varies across hardware.
  • Document profiling results, optimization rationale, code changes, and observed performance gains.
  • Contribute technical insight during discussions involving GPU architecture, accelerated computing, and software performance.
  • Coordinate with remote engineering and cross-functional teams on complex optimization tasks.

Ideal Profile

  • Strong professional expertise in CUDA programming and GPU kernel optimization.
  • Advanced C development skills with experience in performance-sensitive software.
  • Hands-on capability with GLSL and WebGPU.
  • Practical experience using NVIDIA Nsight or similar GPU profiling and debugging platforms.
  • Detailed understanding of GPU execution models, parallelism, memory hierarchy, occupancy, and performance characteristics.
  • Ability to identify root causes of poor GPU performance through profiling rather than assumptions.
  • Experience optimizing workloads across different generations or classes of GPU hardware is highly desirable.
  • Background in high-performance computing, GPU-accelerated software, scientific computing, graphics, or compute shaders is valuable.
  • Exposure to AI or machine learning acceleration is advantageous but not required.
  • Strong analytical judgment when balancing raw performance, code maintainability, and implementation complexity.
  • Clear technical communication skills for explaining profiling evidence and optimization decisions.
  • Comfortable working independently within a distributed contractor environment.
  • Prior AI experience is not required.

Skills Set

  • CUDA programming
  • CUDA kernel optimization
  • CUDA C
  • Advanced C
  • NVIDIA GPUs
  • NVIDIA Nsight
  • GPU profiling
  • GPU architecture
  • Parallel computing
  • GPU memory optimization
  • Kernel occupancy
  • Thread optimization
  • Memory coalescing
  • Host-device optimizatio