عن معهد النماذج التأسيسية (IFM)
معهد النماذج التأسيسية هو مختبر أبحاث مخصص لبناء ونظم الذكاء الاصطناعي واسعة النطاق وفهمها ونشرها وإدارة مخاطرها. ونحن نقود الابتكار في النماذج التأسيسية وتشغيلها، مما يعزز البحث والتعليم والاعتماد الصناعي من خلال بنية تحتية قابلة للتوسع وتطبيقات في العالم الحقيقي.
كجزء من فريق الهندسة لدينا، ستعمل عند نقطة التقاطع بين تعلم الآلة وتصميم الأنظمة — لبناء طبقات السحابة والتنسيق والنشر التي تشغل الجيل القادم من التطبيقات الذكية في جامعة محمد بن زايد للذكاء الاصطناعي (MBZUAI). ستعمل إلى جانب باحثين ومهندسين عالميين في مجال الذكاء الاصطناعي لتحويل نماذج اللغات الكبيرة (LLMs)، والنماذج الصوتية، والأنظمة متعددة الوسائط إلى الإنتاج على نطاق واسع.
الدور الوظيفي
بصفتك
كبير مهندسي MLOps، ستصمم وتبني وتحافظ على بنية تحتية قوية لتعلم الآلة (ML) عبر خطوط التدريب والاستنتاج والنشر. ستتولى مسؤولية دورة حياة النموذج — بدءًا من استيعاب البيانات وحتى تقديم الخدمة في الوقت الفعلي — وتضمن نشر نماذج اللغات الكبيرة (LLM) والنماذج الصوتية بكفاءة وأمان وبشكل قابل للتكرار في بيئات قائمة على Kubernetes.
يتطلب هذا المنصب خبرة عملية عميقة في
Kubernetes (EKS)، و
Helm، و
البنية التحتية السحابية لـ AWS، و
سلاسل أدوات MLOps الحديثة (مثل
vLLM، و
SGLang، و
OpenWebUI، و
Weights & Biases، و
MLflow). كما تُعد المعرفة بـ
أطر عمل الذكاء الاصطناعي الصوتي/الكلامي مثل
ElevenLabs، و
Whisper، و
RVC ذات قيمة كبيرة.
المسؤوليات الرئيسية
- تصميم وإدارة بنية تحتية قابلة للتوسع لتعلم الآلة على AWS باستخدام EKS وEC2 وRDS وS3 والتحكم في الوصول القائم على IAM
- بناء وصيانة عمليات نشر Kubernetes لاستنتاج LLM وTTS باستخدام Helm وArgoCD ومراقبة Prometheus/Grafana
- تنفيذ وتحسين خطوط تقديم النماذج باستخدام vLLM أو SGLang أو TensorRT أو أطر عمل مماثلة للاستنتاج عالي الإنتاجية
- تطوير أتمتة CI/CD وMLOps لتحديد إصدارات البيانات، والتحقق من صحة النماذج، والنشر (GitHub Actions أو Jenkins أو AWS CodePipeline)
- دمج OpenWebUI أو Gradio أو واجهات المستخدم المماثلة لعروض النماذج التوضيحية الموجهة للمستخدم وأدوات التقييم الداخلي
- التعاون مع باحثي تعلم الآلة لتحويل النماذج إلى منتجات — بما في ذلك TTS (مثل ElevenLabs API)، وASR (Whisper)، وأنظمة الدردشة القائمة على LLM
- ضمان إمكانية الملاحظة، وتحسين التكلفة، وموثوقية الموارد السحابية عبر بيئات متعددة
- المساهمة في الأدوات الداخلية لتنظيم مجموعات البيانات، ومراقبة النماذج، وخطوط إعادة التدريب
- الحفاظ على البنية التحتية كرمز باستخدام Terraform ومخططات Helm لضمان إمكانية التكرار والحوكمة
- دعم أعباء العمل متعددة الوسائط في الوقت الفعلي (الصوت، النص، الرؤية) عبر مجموعات الاستنتاج
المؤهلات الأكاديمية
- خبرة تزيد عن 4 سنوات في MLOps أو DevOps أو هندسة البنية التحتية السحابية لأنظمة تعلم الآلة
- كفاءة عالية في Kubernetes وHelm وتنسيق الحاويات
- خبرة في نشر نماذج تعلم الآلة عبر vLLM أو SGLang أو TensorRT أو Ray Serve
- إتقان خدمات AWS (EKS وEC2 وS3 وRDS وCloudWatch وIAM)
- خبرة قوية في Python وDocker وGit وخطوط CI/CD
- فهم قوي لإدارة دورة حياة النماذج، وخطوط البيانات، وأدوات الملاحظة (Grafana، Prometheus، Loki)
- مهارات تعاون ممتازة مع باحثي تعلم الآلة ومهندسي البرمجيات
الخبرة المهنية – المفضل
- خبرة واسعة في vLLM وK8s وElevenlabs وWhisper وGradio/OpenWebUI، أو استضافة نماذج TTS/ASR المخصصة
- إلمام بجدولة وحدات المعالجة الرسومية المتعددة (multi-GPU)، وتحسين NCCL، ودمج مجموعات الحوسبة عالية الأداء (HPC)
- معرفة بالأمان، وإدارة التكاليف، وسياسة الشبكة في مجموعات Kubernetes متعددة المستأجرين وأنظمة Cloudflare
- خبرة سابقة في نشر LLM، أو خطوط الضبط الدقيق، أو أبحاث النماذج التأسيسية
- خبرة بحوكمة البيانات وعمليات الذكاء الاصطناعي المسؤول في البيئات البحثية أو المؤسسية
About The Institute Of Foundation Models (IFM)
The Institute of Foundation Models is a dedicated research lab for building, understanding, deploying, and risk-managing large-scale AI systems. We drive innovation in foundation models and their operationalization, empowering research, education, and industry adoption through scalable infrastructure and real-world applications.
As part of our engineering team, you will operate at the intersection of machine learning and systems design — building the cloud, orchestration, and deployment layers that power the next generation of intelligent applications at MBZUAI. You’ll work alongside world-class AI researchers and engineers to productionize LLMs, voice models, and multimodal systems at scale.
The Role
As a
Senior MLOps Engineer, you will design, build, and maintain robust ML(Machine Learning) infrastructure across training, inference, and deployment pipelines. You will take ownership of the model lifecycle — from data ingestion to real-time serving — and ensure our LLM and speech models are deployed efficiently, securely, and reproducibly in Kubernetes-based environments.
This position requires deep hands-on experience with
Kubernetes (EKS),
Helm,
AWS cloud infrastructure, and
modern MLOps toolchains (e.g.,
vLLM,
SGLang,
OpenWebUI,
Weights & Biases,
MLflow). Familiarity with
speech/voice AI frameworks like
ElevenLabs,
Whisper, and
RVC is also valuable.
Key Responsibilities
- Design and manage scalable ML infrastructure on AWS using EKS, EC2, RDS, S3, and IAM-based access control
- Build and maintain Kubernetes deployments for LLM and TTS inference using Helm, ArgoCD, and Prometheus/Grafana monitoring
- Implement and optimize model serving pipelines using vLLM, SGLang, TensorRT, or similar frameworks for high-throughput inference
- Develop CI/CD and MLOps automation for data versioning, model validation, and deployment (GitHub Actions, Jenkins, or AWS CodePipeline)
- Integrate OpenWebUI, Gradio, or similar UIs for user-facing model demos and internal evaluation tools
- Collaborate with ML researchers to productize models — including TTS (e.g., ElevenLabs API), ASR (Whisper), and LLM-based chat systems
- Ensure observability, cost optimization, and reliability of cloud resources across multiple environments
- Contribute to internal tools for dataset curation, model monitoring, and retraining pipelines
- Maintain infrastructure-as-code using Terraform and Helm charts for reproducibility and governance
- Support real-time multimodal workloads (voice, text, vision) across inference clusters
Academic Qualifications
- 4+ years of experience in MLOps, DevOps, or Cloud Infrastructure Engineering for ML systems
- Strong proficiency in Kubernetes, Helm, and container orchestration
- Experience deploying ML models via vLLM, SGLang, TensorRT, or Ray Serve
- Proficiency with AWS services (EKS, EC2, S3, RDS, CloudWatch, IAM)
- Solid experience with Python, Docker, Git, and CI/CD pipelines
- Strong understanding of model lifecycle management, data pipelines, and observability tools (Grafana, Prometheus, Loki)
- Excellent collaboration skills with ML researchers and software engineers
Professional Experience – Preferred
- Extensive Experience with vLLM, K8s, Elevenlabs, Whisper, Gradio/OpenWebUI, or custom TTS/ASR model hosting
- Familiarity with multi-GPU scheduling, NCCL optimization, and HPC cluster integration
- Knowledge of security, cost management, and network policy in multi-tenant Kubernetes clusters and cloudflare systems
- Prior work in LLM deployment, fine-tuning pipelines, or foundation model research
- Exposure to data governance and responsible AI operations in research or enterprise settings