وصف الوظيفة
نظرة عامة على الدور نحن نبحث عن مهندس بيانات ذو خبرة (من 5 إلى 9 سنوات) مع خبرة عملية عميقة في Apache Spark، Databricks، Delta Lake، وUnity Catalog.
المسؤوليات الأساسية تصميم وبناء والحفاظ على خطوط أنابيب بيانات دفعة وبث على Databricks باستخدام Spark (SQL، PySpark، أو Scala) تنفيذ نماذج بيانات قوية وعمليات ETL/ELT على رأس Delta Lake، بما في ذلك تطور المخطط، السفر عبر الزمن، وأنماط CDC إعداد وإدارة Unity Catalog (مساحات العمل، الكتالوجات، المخططات، الأذونات، الأساليب الوراثية) لفرض حوكمة البيانات وأمان عبر lakehouse تحسين أداء خطوط الأنابيب والاستعلامات باستخدام معرفة بنى Spark الداخلية، بما في ذلك: التقسيم، Z-Ordering، التجميع السائل، وضبط تخطيط الملفات استراتيجيات التخزين المؤقت، الانضمام البثّي، وإدارة shuffle بشكل فعال قياس الكتلة/التوسع الآلي وتبادل التكاليف/الأداء العمل مع تنسيقات الجدول والملف المتعددة (Delta، Iceberg، Parquet، ORC، Avro، JSON، وغير ذلك) واختيار التنسيق المناسب بناءً على عبء العمل واحتياجات الحوكمة المساهمة في والحفاظ على خطوط CI/CD لعمليات Databricks، الملاحظات، وسير العمل، و/أو Databricks Asset Bundles (DABs) المؤهلات المطلوبة خبرة عملية قوية في بناء خطوط بيانات باستخدام Apache Spark مع خبرة لا تقل عن 5 سنوات (PySpark/Scala/SQL) في الإنتاج خبرة عملية مع Databricks على الأقل في سحابة رئيسية واحدة (AWS، Azure، أو GCP) فهم عميق لـ: Spark internals (نموذج التنفيذ، DAGs، المراحل، المهام، shuffle، Catalyst optimizer) تحسين الأداء وحل المشكلات (مثلاً تقليل التحريف، الانسكاب، تعزيز التبديل، استراتيجيات الانضمام) خبرة صلبة مع Delta Lake (معاملات ACID، فرض المخطط، السفر عبر الزمن، OPTIMIZE/VACUUM) المعرفة بأنظمة الجداول والتنسيقات بما في ذلك Delta، Iceberg، وParquet، ومتى تستخدم أي منها خبرة عملية في: Liquid Clustering وZ-Ordering لتحسين أداء الاستعلام وكفاءة التكلفة تقنيات تحسين التخزين/التخطيط الأخرى (التقسيم، الانضغاط، معالجة الملفات الصغيرة)
Job description
Role Overview We are looking for an experienced Data Engineer (5 to 9 years) with deep hands-on expertise in Apache Spark, Databricks, Delta Lake, and Unity Catalog Key Responsibilities Design, build, and maintain batch and streaming data pipelines on Databricks using Spark (SQL, PySpark, or Scala) Implement robust data models and ETL/ELT workflows on top of Delta Lake, including schema evolution, time travel, and CDC patterns Configure and manage Unity Catalog (workspaces, catalogs, schemas, permissions, lineages) to enforce data governance and security across the lakehouse Optimize pipeline and query performance using Spark internals knowledge, including: Partitioning, Z-Ordering, Liquid Clustering, and file layout tuning Caching strategies, broadcast joins, and efficient shuffle management Cluster sizing/auto-scaling and cost/performance trade-offs Work with multiple table and file formats (Delta, Iceberg, Parquet, ORC, Avro, JSON, etc.
) and choose the right format based on workload and governance needs Contribute to and maintain CI/CD pipelines for Databricks jobs, notebooks, and workflows ,and/or Databricks Asset Bundles (DABs) Required Qualifications Strong hands-on experience building data pipelines using Apache Spark with Min 5 years (PySpark/Scala/SQL) in production Practical experience with Databricks on at least one major cloud (AWS, Azure, or GCP) Deep understanding of: Spark internals (execution model, DAGs, stages, tasks, shuffles, Catalyst optimizer) Performance tuning and troubleshooting (e.
g., skew mitigation, spill, shuffle tuning, join strategies) Solid experience with Delta Lake (ACID transactions, schema enforcement, time travel, OPTIMIZE/VACUUM) Knowledge of table and file formats including Delta, Iceberg, and Parquet, and when to use which Hands-on experience with: Liquid Clustering and Z-Ordering for improving query performance and cost efficiency Other storage/layout optimization techniques (partitioning, compaction, small-file handling)