NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Data Engineer in India
1 day agoBe an early applicant
Apply with autofill
Apply with autofill
EXL·1 day ago
1 day agoBe an early applicant

Senior Databricks Engineer

Pune, IndiaHybridSenior · 5-8 yearsData Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at EXL

How you compare FREE

?
Your scoreYour score: not yet known
→
82
Top 10%Top 10%: 82 out of 100

Top 10% of NextRaise users matched against Data Engineer roles in India.

Must-have skills for this role

  • databricks
  • unity catalog
  • delta lake
  • pyspark

PDF or DOCX · no account needed

Apply faster with autofill FREEThe NextRaise extension autofills your application in one click.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
  • Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage.
  • Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance.
  • Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
  • Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management.
  • Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM).
  • Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
  • Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations.
  • Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing.
  • Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).
  • Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
  • Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.

What they're looking for

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
  • Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage.
  • Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
  • Implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction.
  • Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks.
  • Integrate Databricks with cloud-native services and maintain CI/CD pipelines for Databricks notebooks and jobs.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Databricks Platform Engineering ● Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments. ● Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage. ● Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance. ● Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager). ● Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management. Delta Lake & Lakehouse Architecture ● Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM). ● Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization. ● Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations. ● Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing. ● Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).
Data Pipeline Development (PySpark / SQL) ● Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake. ● Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables. ● Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning. ● Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity. ● Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines. MLflow & AI/ML Workloads ● Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks. ● Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances. ● Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features. ● Enable GenAI workloads — LLM fine-tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI). ● Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints. Cloud Integration & DevOps ● Integrate Databricks with cloud-native services: Azure Data Lake Storage (ADLS). ● Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI. ● Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure-as-code (IaC) deployment of Databricks resources. ● Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte). ● Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools. Governance, Security & Optimization ● Implement row-level security, column masking, and dynamic data views using Unity Catalog policies. ● Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations. ● Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement. ● Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog. ● Document architecture decisions, runbooks, and operational guides for Databricks workloads.

Company

EXL
Pune, India

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from EXL's careers site·first seen 21 Sept 2026·last verified 21 Sept 2026·How we source jobs

Similar jobs

  • Principal Data Engineer (AWS, Databricks, Ai, ML Flow, Data Architecture, Apache Airflow) at MastercardPune, India–match not yet calculated
  • Data Engineer - R01571442 at brillio-2Bengaluru, India–match not yet calculated
  • Lead Data Engineer - R01571409 at brillio-2Bengaluru, India–match not yet calculated
  • Associate Data Engineer, DACI Analytics at Lowe's IndiaBengaluru, India–match not yet calculated
  • Java Spark Big Data Engineer - Assistant Vice President at CitiPune, India–match not yet calculated

Browse more jobs

  • Data Engineer jobs in India
  • Data Analyst jobs in India
  • Data Scientist jobs in India
  • Business Intelligence Analyst jobs in India
  • Data Engineer jobs in United States
  • Data Engineer jobs in United Kingdom