Databricks Data Engineer
Hyderabad, IndiaFull-timeMid · 5+ years
About this role
Job Title: Databricks Data Engineer
Experience: 5+ Years
Employment Type: Full-Time
Certification: Databricks Professional-level certification – Mandatory
Job Overview
We are looking for an experienced Databricks Data Engineer with strong hands-on expertise in building scalable data engineering solutions using Databricks, Python, PySpark, Spark SQL, and SQL.
The ideal candidate will be responsible for designing and developing data pipelines, implementing Lakehouse architectures, optimizing large-scale data processing workloads, and ensuring data quality, governance, and performance across the data platform.
Candidates must hold a relevant Databricks Professional-level certification.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Databricks.
- Develop ETL/ELT pipelines using Python, PySpark, Spark SQL, and SQL.
- Build and maintain Lakehouse solutions using Databricks and Delta Lake.
- Implement Bronze, Silver, and Gold layers using Medallion Architecture.
- Develop and optimize Databricks notebooks, jobs, and workflows.
- Process and transform large volumes of structured and semi-structured data.
- Optimize Spark workloads for performance, scalability, and cost efficiency.
- Implement data quality checks, validation, error handling, and monitoring.
- Work with Unity Catalog for data governance, access control, and metadata management.
- Integrate Databricks with cloud storage, databases, APIs, and other enterprise data sources.
- Troubleshoot pipeline failures and resolve performance issues.
- Collaborate with Data Architects, Analysts, Data Scientists, and business stakeholders.
- Follow engineering best practices for version control, CI/CD, testing, documentation, and production deployments.
Mandatory Skills
- Strong hands-on experience with Databricks.
- Strong experience with Python and PySpark.
- Strong knowledge of Spark SQL and SQL.
- Experience developing production-grade ETL/ELT data pipelines.
- Experience with Delta Lake and Lakehouse Architecture.
- Understanding of Medallion Architecture.
- Experience working with large-scale data processing using Apache Spark.
- Experience with Databricks Workflows/Jobs.
- Knowledge of data modeling, data quality, and performance optimization.
- Experience with at least one cloud platform: Azure, AWS, or GCP.
- Good understanding of Git/version control and CI/CD practices.
Mandatory Certification
Candidate must hold a relevant Databricks Professional-level certification.
For a Data Engineer requirement, you can make this more specific:
Databricks Certified Data Engineer Professional – Mandatory
Candidates holding only an Associate-level certification should not be considered.
Preferred Skills
- Unity Catalog
- Databricks Asset Bundles
- Delta Live Tables / Lakeflow Declarative Pipelines
- Databricks SQL
- CI/CD for Databricks workloads
- Terraform
- Azure Data Factory / AWS data services / GCP data services
- Data governance and security
- Streaming technologies such as Kafka or Spark Structured Streaming
- Migration from legacy/on-premise data platforms to Databricks
Key Skills
Databricks, PySpark, Python, Apache Spark, Spark SQL, SQL, Delta Lake, Lakehouse, Medallion Architecture, Unity Catalog, ETL, ELT, Data Engineering, Databricks Workflows, CI/CD, Cloud Data Engineering
