Lead Data Engineer
Role Overview
Beinex is looking for an experienced Lead Data Engineer to design, develop, and lead enterprise-scale data engineering solutions. The role requires strong hands-on expertise in Databricks, SQL, Python, PySpark, Apache Spark, Delta Lake, data lakehouse architecture, ETL/ELT pipelines, data modelling, and enterprise data quality. The candidate should also have experience in defining and implementing data quality rules aligned with ISO data quality principles and the DAMA framework.
Roles and Responsibilities Design scalable data lake, data warehouse, and lakehouse architectures. Build and optimize data pipelines using Databricks, SQL, Python, PySpark, and Spark. Develop batch, incremental, near-real-time, and streaming data solutions. Implement bronze, silver, and gold data layers using medallion architecture. Design reusable, metadata-driven, and configuration-driven ETL/ELT frameworks. Develop data ingestion solutions for databases, APIs, files, cloud storage, and enterprise systems. Implement Delta Lake capabilities including schema enforcement, schema evolution, time travel, and merge processing. Optimize Databricks clusters, Spark jobs, SQL queries, file structures, and data-processing workloads. Develop advanced SQL queries, views, transformations, reconciliation logic, and data validation controls. Build reusable Python and PySpark libraries for ingestion, transformation, validation, and profiling. Define and implement data quality rules for completeness, accuracy, consistency, validity, uniqueness, integrity, and timeliness. Generate data quality rules using profiling results, business definitions, source constraints, ISO principles, and DAMA guidelines. Define rule thresholds, severity levels, validation outcomes, and exception-handling processes. Develop automated data quality scorecards, exception datasets, and reconciliation reports. Implement data profiling to identify null values, duplicates, invalid formats, anomalies, outliers, and inconsistencies. Design source-to-target reconciliation controls using record counts, totals, hashes, and business measures. Design dimensional, normalized, and analytical data models. Implement metadata management, technical lineage, data classification, and access controls. Apply data masking, row-level security, column-level security, and secure data-access practices. Monitor pipeline performance, data freshness, quality failures, and processing reliability. Troubleshoot complex production issues and perform root-cause analysis. Review architecture, code, SQL, Python, PySpark, data models, and data quality rules. Define engineering standards and reusable technical frameworks. Mentor and guide data engineers across technical workstreams. Maintain architecture documents, source-to-target mappings, data quality catalogues, and technical documentation.
Required Skills
Advanced expertise in Databricks. Strong knowledge of Delta Lake and lakehouse architecture. Advanced SQL development and performance optimization. Strong Python and PySpark programming skills. Strong understanding of Apache Spark and distributed data processing. Experience with ETL and ELT architecture. Experience with batch, incremental, and streaming pipelines. Experience with metadata-driven data engineering frameworks. Strong understanding of data profiling, cleansing, validation, and reconciliation. Practical experience in enterprise data quality rule generation. Knowledge of ISO-aligned data quality practices. Working knowledge of the DAMA Data Management Body of Knowledge. Experience with dimensional modelling and analytical data structures. Knowledge of metadata, lineage, data governance, and access controls. Experience with pipeline, SQL, Spark, and Databricks performance optimization. Strong troubleshooting and technical problem-solving skills. Experience reviewing code and mentoring data engineering teams.
Qualifications
Bachelor’s degree in Computer Science, Information Technology, Engineering, Data Science, or a related field. Eight or more years of experience in data engineering, data integration, big data, or cloud data platforms. Proven experience in a senior or lead data engineering role. Hands-on experience with Databricks, SQL, Python, PySpark, and Apache Spark. Experience designing and implementing enterprise-scale data platforms. Experience developing scalable ETL and ELT pipelines. Experience implementing data quality rules and validation frameworks. Experience with ISO-aligned data quality and DAMA-based data management practices. Experience with data profiling, reconciliation, metadata, lineage, and data modelling. Experience supporting and troubleshooting production data platforms.
Preferred Certifications Databricks Certified Data Engineer. Microsoft Azure Data Engineer certification. AWS or Google Cloud data engineering certification. DAMA Certified Data Management Professional. Relevant certifications in data quality, governance, Spark, or cloud platforms.
Preferred Experience
Experience with Unity Catalog and Databricks governance capabilities. Experience with Delta Live Tables and structured streaming. Experience with cloud platforms such as Azure, AWS, or Google Cloud. Experience in government, banking, healthcare, telecom, or other data-intensive industries.
Key Competencies
Strong technical leadership. Analytical and problem-solving ability. Strong ownership and accountability. Focus on data quality, performance, security, and reliability. Ability to guide and mentor engineering teams. Ability to design practical and scalable technical solutions. Strong documentation and engineering discipline.