الملخص الوظيفي:
نحن نبحث عن مهندس دعم أول ليكون نقطة التصعيد لمشكلات الإنتاج، مما يضمن التشخيص السريع والحل والتواصل الواضح أثناء الحوادث الحرجة عبر المنصة.
المسؤوليات الرئيسية:
• فرز حوادث الإنتاج المصعدة والتحقيق فيها وحلها عبر أنظمة الويب والهاتف المحمول والأنظمة الخلفية
• مراقبة سلامة النظام وتحديد المشكلات المحتملة استباقيًا قبل أن تؤثر على المستخدمين
• صيانة وتحسين أدلة التشغيل وقاعدة المعرفة وتوثيق استكشاف الأخطاء وإصلاحها باستمرار
• التنسيق مع فرق الهندسة وDevOps أثناء الحوادث الكبرى وانقطاعات الخدمة
• تقديم تحليل السبب الجذر وتقارير مفصلة ما بعد الحوادث
• توجيه ورفع كفاءة موظفي الدعم المبتدئين، وتحسين عمليات الدعم الإجمالية وأوقات الاستجابة
• إدارة وتحديد أولوية تذاكر الدعم الواردة وفقًا لشدتها والتزامات اتفاقية مستوى الخدمة (SLA)
• إبلاغ حالة الحوادث والحلول بوضوح لأصحاب المصلحة الداخليين والخارجيين
• المساهمة في أتمتة مهام التشخيص والمعالجة الروتينية
• المشاركة في المناوبة على أهبة الاستعداد (On-call) لدعم الإنتاج الحرجة
المؤهلات والمهارات المطلوبة:
• 7 - 10+ سنوات من الخبرة في الدعم الفني أو هندسة دعم التطبيقات
• مهارات قوية في استكشاف الأخطاء وإصلاحها عبر طبقات التطبيق وقواعد البيانات والبنية التحتية
• خبرة في أدوات التسجيل والمراقبة (مثل Grafana وELK وDatadog)
• معرفة عملية بلغة SQL والبرمجة النصية للتشخيص
• عقلية قوية للشعور بالمسؤولية والقدرة على العمل تحت الضغط أثناء الحوادث
• مهارات تواصل كتابية وشفهية جيدة باللغة الإنجليزية
المؤهلات المفضلة:
• خبرة في البنية التحتية السحابية (AWS أو Azure أو GCP)
• دراية بأطر إدارة الحوادث (ITIL)
• التعامل مع البرمجة النصية/الأتمتة (Python وBash) لأدوات الدعم
• خبرة في قيادة أو تنسيق الاستجابة للحوادث الكبرى
• درجة علمية في هندسة البرمجيات وخبرة في تحليلات البيانات
JOB SUMMARY :
We are looking for a Senior Support Engineer to serve as the escalation point for production issues, ensuring rapid diagnosis, resolution, and clear communication during critical incidents across the platform.
KEY RESPONSIBILITIES :
• Triage, investigate, and resolve escalated production incidents across web, mobile, and backend systems
• Monitor system health and proactively identify potential issues before they impact users
• Maintain and continuously improve runbooks, knowledge base, and troubleshooting documentation
• Coordinate with Engineering and DevOps teams during major incidents and outages
• Provide root-cause analysis and detailed post-incident reports
• Mentor and upskill junior support staff, improving overall support processes and response times
• Manage and prioritize incoming support tickets according to severity and SLA commitments
• Communicate incident status and resolutions clearly to internal and external stakeholders
• Contribute to automation of routine diagnostic and remediation tasks
• Participate in on-call rotation for critical production support
REQUIRED QUALIFICATIONS & SKILLS:
• 7 -10 + years of technical support or application support engineering experience
• Strong troubleshooting skills across application, database, and infrastructure layers
• Experience with logging/monitoring tools (e.g. Grafana, ELK, Datadog)
• Working knowledge of SQL and scripting for diagnostics
• Strong ownership mindset and ability to work under pressure during incidents
• Good written and verbal communication skills in English
PREFERRED QUALIFICATIONS :
• Experience with cloud infrastructure (AWS, Azure, or GCP)
• Familiarity with incident management frameworks (ITIL)
• Exposure to scripting/automation (Python, Bash) for support tooling
• Experience leading or coordinating major incident response
• Software Engineering Degree and exp in Data Analytics