TRCD-26-05366
Lead / Principal Data Engineer (Lakehouse + GenAI Integration)
Lead/Principal Data Engineer needed to architect and scale enterprise, production-grade lakehouse data platforms (Medallion/Bronze-Silver-Gold) across cloud stacks (Microsoft Fabric/Azure/Databricks/Snowflake/AWS). The role focuses on building high-throughput batch and streaming pipelines with PySpark/SQL/Python, modern orchestration and transformation (Airflow/ADF/Fabric Pipelines, dbt, Kafka/Spark Streaming), and embedding GenAI capabilities (RAG, vector stores, agentic workflows, LLM APIs) into operational data pipelines, while owning governance, standards, and production troubleshooting.
Position summary
- Location
- India
- Workplace
- Hybrid
- Employment
- Contract
- Experience
- Minimum 7 years and Maximum 12 years
Role overview
Why This Role Matters.
Lead/Principal Data Engineer needed to architect and scale enterprise, production-grade lakehouse data platforms (Medallion/Bronze-Silver-Gold) across cloud stacks (Microsoft Fabric/Azure/Databricks/Snowflake/AWS). The role focuses on building high-throughput batch and streaming pipelines with PySpark/SQL/Python, modern orchestration and transformation (Airflow/ADF/Fabric Pipelines, dbt, Kafka/Spark Streaming), and embedding GenAI capabilities (RAG, vector stores, agentic workflows, LLM APIs) into operational data pipelines, while owning governance, standards, and production troubleshooting.
Your Impact
Deliver Enterprise Value
Help organisations solve complex business problems through modern technology, consulting expertise and measurable outcomes.
Collaboration
Work Across Teams
Collaborate with consultants, architects, engineers and client stakeholders throughout the project lifecycle.
Growth
Learn Continuously
Gain exposure to enterprise technologies, certifications, mentoring and real-world project experience.
Career Path
Grow With Ubique
Build a long-term consulting career with opportunities to take on greater responsibility and leadership over time.
Responsibilities
What You'll Be Doing.
Every role at Ubique contributes directly to solving meaningful business challenges for our clients.
Architect end-to-end multi-cloud data platforms using Lakehouse (Bronze-Silver-Gold / Medallion Architecture) and Star/Snowflake dimensional modeling standards.
Lead data platform migrations, consolidations, and cost-optimization strategies (e.g., Fabric, Azure, Databricks, Snowflake).
Core Data Engineering & Orchestration
Design, build, and optimize robust ETL/ELT pipelines for incremental loads, high-throughput processing, and data quality assurance using PySpark, SQL, and Python.
Implement streaming workflows (Kafka / Spark Streaming) and modern transformation layers (dbt, Airflow).
Tune performance for high-availability reporting, latency reduction (~40%), and large-scale data processing.
GenAI & AI Pipeline Integration
Embed Generative AI models, vector stores (e.g., ChromaDB, Azure AI Search), and agentic workflows (e.g., LangChain, Azure AI Foundry) into core data pipelines for automated triage and natural-language-to-SQL analytics.
Governance & Leadership
Establish reusable engineering frameworks, enterprise data modeling standards, and data governance/lineage controls.
Lead production issue resolution, root-cause analysis, and performance troubleshooting.
Technology stack
Tools & Technologies.
The platforms and technologies you'll use to build modern, enterprise-grade solutions.
Microsoft Fabric
Azure Analytics
Azure Data Factory
Fabric Pipelines
Databricks
Snowflake
AWS
S3
Glue
Lambda
Athena
Lakehouse
Medallion Architecture
Bronze Silver Gold
Star Schema
Snowflake Schema
Dimensional Modeling
PySpark
Spark
SQL
Requirements
Skills & Experience.
We value curiosity, collaboration and continuous learning. If you don't meet every requirement but believe you can make an impact, we'd still love to hear from you.
Essential
Required Qualifications
7+ years in data engineering/data architecture with production platform delivery
Lakehouse / Medallion architecture (Bronze-Silver-Gold)
Dimensional modeling (star/snowflake schemas)
PySpark
SQL
Python
ETL/ELT pipeline design and optimization (incremental loads, high-throughput, data quality)
Streaming pipelines (Apache Kafka and/or Spark Streaming)
Orchestration (Airflow or Azure Data Factory or Fabric Pipelines)
dbt (modern transformation layer)
Cloud data platforms experience: Microsoft Fabric and/or Azure Analytics and/or Databricks and/or Snowflake and/or AWS (S3, Glue, Lambda, Athena)
GenAI integration into data workflows: RAG, vector stores, LLM APIs (OpenAI/Gemini), frameworks/tools (LangChain, Azure AI Search, Azure AI Foundry)
Production issue resolution, root-cause analysis, performance troubleshooting
Data governance/lineage controls and enterprise standards
Preferred
Nice to Have
Microsoft Fabric (explicit deep experience preferred)
Cost optimization and platform migration/consolidation experience (Fabric/Azure/Databricks/Snowflake)
Performance tuning for reporting/latency reduction and high availability
BI semantic modeling
DAX
Row-Level Security (RLS)
What you'll gain
More Than Just A Job.
We're committed to helping every team member grow professionally, personally and technically while working on meaningful projects.
Global Exposure
Collaborate with international clients and multicultural teams on enterprise programmes.
Continuous Learning
Expand your expertise through mentoring, certifications and hands-on project experience.
Career Growth
Take ownership, develop leadership skills and grow your consulting career over time.
Flexible Working
Hybrid and remote collaboration designed around trust and delivering exceptional outcomes.
People First
Join a supportive culture where collaboration, respect and long-term relationships come first.
Enterprise Projects
Work on meaningful technology initiatives for leading organisations across industries.
Apply
Apply for Lead / Principal Data Engineer (Lakehouse + GenAI Integration)
One page, about two minutes. We only ask for what we actually need to have a first conversation.