Technical Site Reliability Engineer
احفظ هذه الوظيفة وحافظ على تنظيم بحثك
قم بإنشاء حساب مجاني لحفظ الوظائف وإنشاء التنبيهات والعودة إلى هذه القائمة من لوحة التحكم الخاصة بك.
بالمتابعة، فإنك توافق على الشروط & سياسة الخصوصية.
Anduril Industries is a defense technology company building AI-powered military systems. As the founding Site Reliability Engineer for the Advanced Capabilities division, you will design, build, and operate infrastructure for next-generation wargaming simulation facilities that run massive-scale simulations of autonomous systems in contested environments.
What You'll Do
- Maintain the simulation software stack including installation, configuration, updates, version management, and day-to-day functionality across simulation center tools and environments
- Own underlying infrastructure—compute, networking, storage, and environment configuration—keeping systems provisioned, patched, and performant
- Build and maintain post-release test suites with regression and smoke-test automation that runs after every software release to surface integration issues immediately
- Root-cause errors and failures in the system, drive them to permanent resolution, and implement guardrails, monitoring, or process changes to prevent recurrence
- Partner with development teams to review changes for reliability risk, surface concerns early, and implement mitigation strategies before releases
What You Need
- Proficiency in Python for automation, tooling, and test development
- Working knowledge of C++ to read, debug, build, and trace issues in simulation codebases
- Solid general networking fundamentals including TCP/IP, UDP, multicast, DNS, routing, and firewall configuration with ability to diagnose latency, packet loss, and connectivity problems
- Experience maintaining production or production-adjacent systems and troubleshooting under time pressure
- Strong written and verbal communication skills for escalation, root cause explanation, and runbook documentation
- Eligibility to pass security and background check requirements for sensitive information systems
Nice to Have
- Experience with modeling and simulation, wargaming, or distributed simulation standards such as DIS, HLA, TENA and platforms like AFSIM or VBS
- Test automation and CI/CD experience including building automated validation pipelines
- Infrastructure-as-code and configuration management tools such as Terraform, Ansible, Docker, and Kubernetes
- On-premises and cloud deployment experience
- Observability tooling such as Prometheus, Grafana, or ELK
- Linux systems administration depth with comfort in mixed Linux and Windows environments
- Prior work in defense, aerospace, or classified environments
- Active security clearance
Highly competitive equity grants are included in the majority of full‑time offers and are considered part of total compensation package. Top-tier benefits for full‑time employees including health and recovery support.