Principal AI Ops Engineer – Abu Dhabi
Discover the Opportunity:
We’re partnering with a leading organisation in Abu Dhabi that is building and operating advanced AI systems at significant scale.
They’re looking for a Principal AI Ops Engineer to take ownership of how AI and LLM systems operate in production, covering inference and model serving, deployment, observability, reliability and performance.
This is a Principal-level individual contributor role for someone who combines deep AI infrastructure expertise with strong software and reliability engineering fundamentals. You’ll set the operational standards that allow engineering teams to deploy and run production AI systems safely, reliably and efficiently.
Discover the Responsibilities:
- Design, operate and optimise GPU-based inference and model-serving infrastructure for production AI and LLM workloads.
- Optimise model serving across latency, throughput, batching, quantisation, autoscaling and infrastructure cost.
- Build automated release pipelines for models, prompts and agent configurations, including canary deployments, regression gates and rollback strategies.
- Establish AI-specific observability across model, retrieval and orchestration layers, including tracing, latency, cost and quality monitoring.
- Define and maintain SLOs across availability, latency and AI system quality, alongside automated alerting and incident response processes.
- Own capacity planning and cost optimisation across GPU and AI infrastructure.
- Build secure, scalable Kubernetes environments and Infrastructure-as-Code patterns for production AI workloads.
- Develop reusable deployment patterns, tooling and operational standards that enable engineering teams to ship AI systems reliably.
- Lead complex production incidents, load testing and root-cause analysis across AI infrastructure and applications.
- Provide technical leadership and help establish engineering standards for operating AI systems at scale.
Discover the Requirements:
- Proven experience operating at Staff, Principal or equivalent senior IC level, with a track record of running production ML or LLM systems at scale.
- Deep hands-on experience with GPU-based inference and model serving, including technologies such as vLLM, TGI, TensorRT-LLM or similar.
- Strong understanding of batching, quantisation, latency, throughput, autoscaling and the performance trade-offs involved in production LLM serving.
- Strong experience with AI/LLM observability, including tracing, quality monitoring, drift and regression detection.
- Strong reliability engineering fundamentals across SLOs, incident response, capacity planning and post-mortems.
- Strong Python engineering skills with experience building production-grade automation and infrastructure tooling.
- Deep experience with Kubernetes, Docker and Infrastructure as Code, ideally Terraform, across cloud environments.
- Experience with observability technologies such as Langfuse, LangSmith, Arize Phoenix, Grafana or Prometheus.
- Experience integrating AI evaluation and regression testing into CI/CD and production release processes.
- Experience with cloud infrastructure, ideally Azure, and production environments with strong security, data residency or compliance requirements.
- Experience with GPU/AI infrastructure cost optimisation and FinOps would be advantageous.
- A highly hands-on approach, with the ability to set technical standards while remaining close to engineering and production systems.
مهندس رئيسي لعمليات الذكاء الاصطناعي – أبوظبي
اكتشف الفرصة:
نحن نتشارك مع مؤسسة رائدة في أبوظبي تقوم ببناء وتشغيل أنظمة ذكاء اصطناعي متقدمة على نطاق واسع.
إنهم يبحثون عن مهندس رئيسي لعمليات الذكاء الاصطناعي ليتولى مسؤولية تشغيل أنظمة الذكاء الاصطناعي ونماذج اللغات الكبيرة (LLM) في بيئة الإنتاج، بما يشمل الاستدلال وتقديم النماذج والنشر وإمكانية الملاحظة والموثوقية والأداء.
هذا دور مساهم فردي على مستوى رئيسي (Principal) لشخص يجمع بين الخبرة العميقة في البنية التحتية للذكاء الاصطناعي وأساسيات البرمجيات وهندسة الموثوقية القوية. ستحدد المعايير التشغيلية التي تتيح لفرق الهندسة نشر وتشغيل أنظمة الذكاء الاصطناعي في بيئة الإنتاج بأمان وموثوقية وكفاءة.
اكتشف المسؤوليات:
- تصميم وتشغيل وتحسين البنية التحتية للاستدلال وتقديم النماذج القائمة على وحدات معالجة الرسومات (GPU) لأحجام العمل الخاصة بالذكاء الاصطناعي ونماذج اللغات الكبيرة في بيئة الإنتاج.
- تحسين تقديم النماذج عبر زمن الانتقال، ومعدل الإنتاجية، والمعالجة بالدفعة، والتكميم، والتوسع التلقائي، وتكلفة البنية التحتية.
- بناء خطوط أنابيب إصدار مؤتمتة للنماذج والمطالبات وتكوينات الوكلاء، بما في ذلك عمليات النشر التجريبي، وبوابات التراجع، واستراتيجيات الاسترداد.
- إنشاء إمكانية ملاحظة خاصة بالذكاء الاصطناعي عبر طبقات النماذج والاسترجاع والتنسيق، بما في ذلك التتبع، وزمن الانتقال، والتكلفة، ومراقبة الجودة.
- تحديد وصيانة أهداف مستوى الخدمة (SLOs) عبر التوافر وزمن الانتقال وجودة نظام الذكاء الاصطناعي، إلى جانب التنبيهات المؤتمتة وعمليات الاستجابة للحوادث.
- تولي التخطيط للقدرة الاستيعابية وتحسين التكلفة عبر البنية التحتية لوحدات معالجة الرسومات والذكاء الاصطناعي.
- بناء بيئات Kubernetes آمنة وقابلة للتوسع وأنماط البنية التحتية ككود (Infrastructure-as-Code) لأحجام عمل الذكاء الاصطناعي في بيئة الإنتاج.
- تطوير أنماط نشر وأدوات ومعايير تشغيلية قابلة لإعادة الاستخدام تمكّن الفرق الهندسيّة من إطلاق أنظمة الذكاء الاصطناعي بموثوقية.
- قيادة حوادث الإنتاج المعقدة، واختبارات الحمل، وتحليل السبب الجذر عبر البنية التحتية والتطبيقات الخاصة بالذكاء الاصطناعي.
- تقديم القيادة الفنية والمساعدة في إرساء المعايير الهندسية لتشغيل أنظمة الذكاء الاصطناعي على نطاق واسع.
اكتشف المتطلبات:
- خبرة مثبتة في العمل بمستوى Staff أو Principal أو ما يعادلها من مستويات المساهم الفردي الخبير (senior IC)، مع سجل حافل في تشغيل أنظمة تعلم الآلة أو نماذج اللغات الكبيرة في بيئة الإنتاج على نطاق واسع.
- خبرة عملية عميقة في الاستدلال وتقديم النماذج القائمة على وحدات معالجة الرسومات (GPU)، بما في ذلك تقنيات مثل vLLM أو TGI أو TensorRT-LLM أو ما شابه ذلك.
- فهم قوي للمعالجة بالدفعة والتكميم وزمن الانتقال ومعدل الإنتاجية والتوسع التلقائي والمفاضلات في الأداء المتعلقة بتقديم نماذج اللغات الكبيرة في بيئة الإنتاج.
- خبرة قوية في إمكانية ملاحظة الذكاء الاصطناعي/نماذج اللغات الكبيرة، بما في ذلك التتبع، ومراقبة الجودة، واكتشاف الانحراف والتراجع.
- أساسيات قوية في هندسة الموثوقية عبر أهداف مستوى الخدمة (SLOs)، والاستجابة للحوادث، والتخطيط للقدرة الاستيعابية، وتحليلات ما بعد الحوادث.
- مهارات هندسية قوية في لغة Python مع خبرة في بناء أتمتة وأدوات بنية تحتية مخصصة لبيئات الإنتاج.
- خبرة عميقة في Kubernetes وDocker والبنية التحتية ككود (Infrastructure as Code)، ويفضل Terraform، عبر بيئات السحابة.
- خبرة في تقنيات إمكانية الملاحظة مثل Langfuse أو LangSmith أو Arize Phoenix أو Grafana أو Prometheus.
- خبرة في دمج تقييم الذكاء الاصطناعي واختبار التراجع في عمليات التكامل المستمر/النشر المستمر (CI/CD) وإصدارات الإنتاج.
- خبرة في البنية التحتية السحابية، ويفضل Azure، وبيئات الإنتاج ذات المتطلبات الصارمة في الأمان، وإقامة البيانات، والامتثال.
- تعتبر الخبرة في تحسين تكلفة البنية التحتية للذكاء الاصطناعي/وحدات معالجة الرسومات وممارسات FinOps ميزة إضافية.
- أسلوب عمل عملي وميداني بدرجة عالية، مع القدرة على وضع المعايير الفنية مع البقاء على مقربة من الأنظمة الهندسية وأنظمة الإنتاج.