On-site Full Time
--
Discovered MENA

Job Details

Principal AI Ops Engineer – Abu Dhabi

Discover the Opportunity:

We’re partnering with a leading organisation in Abu Dhabi that is building and operating advanced AI systems at significant scale.


They’re looking for a Principal AI Ops Engineer to take ownership of how AI and LLM systems operate in production, covering inference and model serving, deployment, observability, reliability and performance.


This is a Principal-level individual contributor role for someone who combines deep AI infrastructure expertise with strong software and reliability engineering fundamentals. You’ll set the operational standards that allow engineering teams to deploy and run production AI systems safely, reliably and efficiently.


Discover the Responsibilities:

  • Design, operate and optimise GPU-based inference and model-serving infrastructure for production AI and LLM workloads.
  • Optimise model serving across latency, throughput, batching, quantisation, autoscaling and infrastructure cost.
  • Build automated release pipelines for models, prompts and agent configurations, including canary deployments, regression gates and rollback strategies.
  • Establish AI-specific observability across model, retrieval and orchestration layers, including tracing, latency, cost and quality monitoring.
  • Define and maintain SLOs across availability, latency and AI system quality, alongside automated alerting and incident response processes.
  • Own capacity planning and cost optimisation across GPU and AI infrastructure.
  • Build secure, scalable Kubernetes environments and Infrastructure-as-Code patterns for production AI workloads.
  • Develop reusable deployment patterns, tooling and operational standards that enable engineering teams to ship AI systems reliably.
  • Lead complex production incidents, load testing and root-cause analysis across AI infrastructure and applications.
  • Provide technical leadership and help establish engineering standards for operating AI systems at scale.


Discover the Requirements:

  • Proven experience operating at Staff, Principal or equivalent senior IC level, with a track record of running production ML or LLM systems at scale.
  • Deep hands-on experience with GPU-based inference and model serving, including technologies such as vLLM, TGI, TensorRT-LLM or similar.
  • Strong understanding of batching, quantisation, latency, throughput, autoscaling and the performance trade-offs involved in production LLM serving.
  • Strong experience with AI/LLM observability, including tracing, quality monitoring, drift and regression detection.
  • Strong reliability engineering fundamentals across SLOs, incident response, capacity planning and post-mortems.
  • Strong Python engineering skills with experience building production-grade automation and infrastructure tooling.
  • Deep experience with Kubernetes, Docker and Infrastructure as Code, ideally Terraform, across cloud environments.
  • Experience with observability technologies such as Langfuse, LangSmith, Arize Phoenix, Grafana or Prometheus.
  • Experience integrating AI evaluation and regression testing into CI/CD and production release processes.
  • Experience with cloud infrastructure, ideally Azure, and production environments with strong security, data residency or compliance requirements.
  • Experience with GPU/AI infrastructure cost optimisation and FinOps would be advantageous.
  • A highly hands-on approach, with the ability to set technical standards while remaining close to engineering and production systems.

Similar Jobs

About Discovered MENA
UAE, Abu Dhabi
Staffing and Recruiting