Back to board
prudentglobaltech·11 days ago

Data Architect

Hyderabad, IndiaFull-timeMid · 2-5 years

About this role

About the Role
We're looking for a hands-on Data Architect who combines deep architectural thinking with real, hands-on engineering ability. This is not a "design-only" architect role—you'll be expected to code, build, and guide pipelines yourself while also owning the bigger picture: architecture, standards, stakeholder alignment, and end-to-end delivery.
You'll bridge business needs with scalable, production-grade data platforms across Snowflake, Data Engineering, Data Quality, AI/ML enablement, Agentic AI solutions, and downstream analytics consumption.

Key Responsibilities
Architecture & Design
    • Design and own end-to-end enterprise data architecture with Snowflake as the core data platform, leveraging Medallion (Bronze/Silver/Gold) architecture, Lakehouse patterns, and modern cloud data architectures.
    • Define data modelling standards, ingestion strategies, governance, and storage/compute optimization across Snowflake, Databricks, Microsoft Fabric, and Azure.
    • Architect scalable, secure, and high-performance solutions for BI, Analytics, AI/ML, Agentic AI, APIs, and enterprise applications.
    • Establish enterprise standards for scalability, security, metadata management, lineage, observability, cost optimization, and governance.
Hands-On Engineering
    • Personally build and develop production-grade data pipelines using SQL, Python, Snowpark, and PySpark.
    • Design, develop, and optimize Snowflake-native ELT pipelines using Snowpipe, Dynamic Tables, Streams & Tasks, and Snowpark.
    • Modernize legacy ETL pipelines into scalable cloud-native architectures following industry best practices.
    • Build reusable APIs and data services for enterprise applications and downstream systems.
    • Implement enterprise Data Quality frameworks with validation, reconciliation, anomaly detection, monitoring, and alerting.
    • Optimize Snowflake warehouses, Spark jobs, Delta tables, and end-to-end pipeline performance.
Agentic AI & AI Enablement
    • Design enterprise data architectures that enable AI Agents, Copilots, Retrieval-Augmented Generation (RAG), and LLM-powered applications.
    • Build pipelines supporting vector search, embeddings, semantic search, and enterprise knowledge repositories.
    • Integrate Snowflake Cortex AI, Cortex Search, Cortex Analyst, Azure AI Foundry, Microsoft AI Foundry, LangChain, LangGraph, Semantic Kernel, CrewAI, AutoGen, or similar AI orchestration frameworks.
    • Design secure AI-ready data platforms with governance, lineage, RBAC, and metadata management.
    • Develop ML-ready datasets and feature engineering pipelines supporting Machine Learning and Generative AI workloads.
Team Leadership & Stakeholder Management
    • Mentor engineering teams on architecture, coding standards, performance optimization, and engineering best practices.
    • Partner with business and product stakeholders to translate business requirements into scalable technical solutions.
    • Own the complete Software Development Life Cycle (SDLC), including architecture, development, testing, deployment, monitoring, and production support.
    • Drive architecture reviews, technical governance, and engineering excellence across the organization.
Platform & Ecosystem
    • Work extensively across Snowflake, Microsoft Azure, Microsoft Fabric, Databricks, and modern cloud-native ecosystems.
    • Build secure, governed, scalable, and AI-ready enterprise data platforms.
    • Implement CI/CD, Infrastructure as Code (IaC), monitoring, logging, and observability across data platforms.

Required Skills & Experience
    • Proven experience in Data Engineering and Data Architecture with strong hands-on development expertise.
    • Expert-level SQL, Python, Snowpark, PySpark, and Apache Spark.
    • Strong expertise in Snowflake, including Snowsight, Snowpark, Cortex AI, Cortex Search, Cortex Analyst, Snowpipe, Dynamic Tables, Streams & Tasks, CLI, Horizon Catalog, Tags, Data Sharing, and Performance Optimization.
    • Strong expertise in Databricks, including Delta Lake, Unity Catalog, Delta Live Tables (DLT), Spark Optimization, Workflows, and MLflow.
    • Strong understanding of Microsoft Fabric and Azure Data Platform.
    • Deep expertise in Medallion Architecture, Lakehouse Architecture, Data Mesh, Data Vault, and modern ELT/ETL frameworks.
    • Experience designing enterprise Data Quality frameworks using Great Expectations, Soda, Deequ, or custom frameworks.
    • Experience developing REST APIs and enterprise data service layers.
    • Strong understanding of AI/ML lifecycle, Feature Engineering, MLOps, LLM integration, and Agentic AI architectures.
    • Experience implementing enterprise governance, metadata management, data lineage, RBAC, masking policies, and security frameworks.
    • Experience with Git, Azure DevOps, GitHub Actions, CI/CD, automated testing, and release management.
    • Excellent stakeholder communication, solution architecture, and technical leadership skills.

Technology Stack
Cloud & Data Platforms
    • Snowflake (Snowpark, Cortex AI, Cortex Search, Cortex Analyst, Snowsight, Snowpipe, Dynamic Tables, Streams & Tasks, Horizon Catalog, Native Apps, Data Sharing)
    • BigQuery
    • Microsoft Azure
    • Microsoft Fabric
    • Databricks
    • Azure Data Lake Storage Gen2 (ADLS Gen2)
    • Azure Synapse Analytics
Data Engineering & Processing
    • SQL
    • Python
    • Snowpark
    • PySpark
    • Apache Spark
    • Delta Lake
    • Delta Live Tables (DLT)
    • Azure Data Factory (ADF)
    • Microsoft Fabric Data Factory
    • Apache Airflow
Data Architecture & Modeling
    • Medallion Architecture (Bronze/Silver/Gold)
    • Lakehouse Architecture
    • Data Mesh
    • Data Vault
    • Star Schema
    • Dimensional Modeling
Data Quality & Governance
    • Snowflake Governance (Tags, Masking Policies, Row Access Policies, Horizon Catalog)
    • Great Expectations
    • Soda
    • Unity Catalog
    • Microsoft Purview
    • Data Lineage
    • Metadata Management
Agentic AI & AI/ML
    • Snowflake Cortex AI
    • Cortex Search
    • Cortex Analyst
    • Azure AI Foundry
    • Microsoft AI Foundry
    • Azure OpenAI
    • OpenAI APIs
    • LangChain
    • LangGraph
    • Semantic Kernel
    • CrewAI
    • AutoGen
    • Model Context Protocol (MCP)
    • Retrieval-Augmented Generation (RAG)
    • Vector Databases (Azure AI Search, Pinecone, Weaviate, Milvus, ChromaDB)
    • MLflow
APIs & Integration
    • REST APIs
    • FastAPI
    • GraphQL
    • Azure Functions
    • Event Grid
    • Azure Service Bus
DevOps & CI/CD
    • Git
    • GitHub
    • GitHub Actions
    • Azure DevOps
    • Terraform
    • Docker
    • Kubernetes
BI & Analytics
    • Power BI
    • Tableau
    • Microsoft Fabric Real-Time Intelligence
    • Semantic Models
    • DAX

Nice to Have
    • SnowPro Core and SnowPro Advanced Certifications
    • Databricks Certified Data Engineer Professional
    • Microsoft Fabric Analytics Engineer (DP-600)
    • Azure Data Engineer Associate (DP-203)
    • Azure AI Engineer Associate (AI-102)
    • Experience with Kafka, Azure Event Hubs, Apache Flink, or Spark Structured Streaming.
    • Experience designing and implementing enterprise AI Agents, Multi-Agent Systems, MCP Servers, RAG applications, and Knowledge Graphs.
    • Experience with Vector Databases such as Azure AI Search, Pinecone, Weaviate, Milvus, or ChromaDB.
    • Experience working with enterprise data governance, FinOps, and cloud cost optimization across Snowflake and Azure.