Back to board
unisongroup·8 days ago

Sr Databricks Engineer

Abu Dhabi, United Arab EmiratesFull-timeMid · 5-9 yearsH1B likely

About this role

Role Overview
We are looking for an experienced Data Engineer (5 to 9 years) with deep hands-on expertise in Apache Spark, Databricks, Delta Lake, and Unity Catalog

Key Responsibilities

Design, build, and maintain batch and streaming data pipelines on Databricks using Spark (SQL, PySpark, or Scala)
Implement robust data models and ETL/ELT workflows on top of Delta Lake, including schema evolution, time travel, and CDC patterns
Configure and manage Unity Catalog (workspaces, catalogs, schemas, permissions, lineages) to enforce data governance and security across the lakehouse
Optimize pipeline and query performance using Spark internals knowledge, including: Partitioning, Z-Ordering, Liquid Clustering, and file layout tuning Caching strategies, broadcast joins, and efficient shuffle management Cluster sizing/auto-scaling and cost/performance trade-offs
Work with multiple table and file formats (Delta, Iceberg, Parquet, ORC, Avro, JSON, etc.) and choose the right format based on workload and governance needs
Contribute to and maintain CI/CD pipelines for Databricks jobs, notebooks, and workflows ,and/or Databricks Asset Bundles (DABs)

Requirements

Required Qualifications
Strong hands-on experience building data pipelines using Apache Spark with Min 5 years
(PySpark/Scala/SQL) in production
Practical experience with Databricks on at least one major cloud (AWS, Azure, or GCP)
Deep understanding of: Spark internals (execution model, DAGs, stages, tasks, shuffles, Catalyst optimizer)
Performance tuning and troubleshooting (e.g., skew mitigation, spill, shuffle tuning, join strategies)
Solid experience with Delta Lake (ACID transactions, schema enforcement, time travel, OPTIMIZE/VACUUM)
Knowledge of table and file formats including Delta, Iceberg, and Parquet, and when to use which
Hands-on experience with:
Liquid Clustering and Z-Ordering for improving query performance and cost efficiency
Other storage/layout optimization techniques (partitioning, compaction, small-file handling)