Ubique Systems

TRCD-26-05366

Lead / Principal Data Engineer (Lakehouse + GenAI Integration)

Lead/Principal Data Engineer needed to architect and scale enterprise, production-grade lakehouse data platforms (Medallion/Bronze-Silver-Gold) across cloud stacks (Microsoft Fabric/Azure/Databricks/Snowflake/AWS). The role focuses on building high-throughput batch and streaming pipelines with PySpark/SQL/Python, modern orchestration and transformation (Airflow/ADF/Fabric Pipelines, dbt, Kafka/Spark Streaming), and embedding GenAI capabilities (RAG, vector stores, agentic workflows, LLM APIs) into operational data pipelines, while owning governance, standards, and production troubleshooting.

Position summary

Location
India
Workplace
Hybrid
Employment
Contract
Experience
Minimum 7 years and Maximum 12 years
Apply now

Role overview

Why This Role Matters.

Lead/Principal Data Engineer needed to architect and scale enterprise, production-grade lakehouse data platforms (Medallion/Bronze-Silver-Gold) across cloud stacks (Microsoft Fabric/Azure/Databricks/Snowflake/AWS). The role focuses on building high-throughput batch and streaming pipelines with PySpark/SQL/Python, modern orchestration and transformation (Airflow/ADF/Fabric Pipelines, dbt, Kafka/Spark Streaming), and embedding GenAI capabilities (RAG, vector stores, agentic workflows, LLM APIs) into operational data pipelines, while owning governance, standards, and production troubleshooting.

Your Impact

Deliver Enterprise Value

Help organisations solve complex business problems through modern technology, consulting expertise and measurable outcomes.

Collaboration

Work Across Teams

Collaborate with consultants, architects, engineers and client stakeholders throughout the project lifecycle.

Growth

Learn Continuously

Gain exposure to enterprise technologies, certifications, mentoring and real-world project experience.

Career Path

Grow With Ubique

Build a long-term consulting career with opportunities to take on greater responsibility and leadership over time.

Responsibilities

What You'll Be Doing.

Every role at Ubique contributes directly to solving meaningful business challenges for our clients.

01

Architecture & Platform Design

02

Architect end-to-end multi-cloud data platforms using Lakehouse (Bronze-Silver-Gold / Medallion Architecture) and Star/Snowflake dimensional modeling standards.

03

Lead data platform migrations, consolidations, and cost-optimization strategies (e.g., Fabric, Azure, Databricks, Snowflake).

04

Core Data Engineering & Orchestration

05

Design, build, and optimize robust ETL/ELT pipelines for incremental loads, high-throughput processing, and data quality assurance using PySpark, SQL, and Python.

06

Implement streaming workflows (Kafka / Spark Streaming) and modern transformation layers (dbt, Airflow).

07

Tune performance for high-availability reporting, latency reduction (~40%), and large-scale data processing.

08

GenAI & AI Pipeline Integration

09

Embed Generative AI models, vector stores (e.g., ChromaDB, Azure AI Search), and agentic workflows (e.g., LangChain, Azure AI Foundry) into core data pipelines for automated triage and natural-language-to-SQL analytics.

10

Governance & Leadership

11

Establish reusable engineering frameworks, enterprise data modeling standards, and data governance/lineage controls.

12

Lead production issue resolution, root-cause analysis, and performance troubleshooting.

Technology stack

Tools & Technologies.

The platforms and technologies you'll use to build modern, enterprise-grade solutions.

01

Microsoft Fabric

02

Azure Analytics

03

Azure Data Factory

04

Fabric Pipelines

05

Databricks

06

Snowflake

07

AWS

08

S3

09

Glue

10

Lambda

11

Athena

12

Lakehouse

13

Medallion Architecture

14

Bronze Silver Gold

15

Star Schema

16

Snowflake Schema

17

Dimensional Modeling

18

PySpark

19

Spark

20

SQL

Requirements

Skills & Experience.

We value curiosity, collaboration and continuous learning. If you don't meet every requirement but believe you can make an impact, we'd still love to hear from you.

Essential

Required Qualifications

7+ years in data engineering/data architecture with production platform delivery

Lakehouse / Medallion architecture (Bronze-Silver-Gold)

Dimensional modeling (star/snowflake schemas)

PySpark

SQL

Python

ETL/ELT pipeline design and optimization (incremental loads, high-throughput, data quality)

Streaming pipelines (Apache Kafka and/or Spark Streaming)

Orchestration (Airflow or Azure Data Factory or Fabric Pipelines)

dbt (modern transformation layer)

Cloud data platforms experience: Microsoft Fabric and/or Azure Analytics and/or Databricks and/or Snowflake and/or AWS (S3, Glue, Lambda, Athena)

GenAI integration into data workflows: RAG, vector stores, LLM APIs (OpenAI/Gemini), frameworks/tools (LangChain, Azure AI Search, Azure AI Foundry)

Production issue resolution, root-cause analysis, performance troubleshooting

Data governance/lineage controls and enterprise standards

Preferred

Nice to Have

Microsoft Fabric (explicit deep experience preferred)

Cost optimization and platform migration/consolidation experience (Fabric/Azure/Databricks/Snowflake)

Performance tuning for reporting/latency reduction and high availability

BI semantic modeling

DAX

Row-Level Security (RLS)

What you'll gain

More Than Just A Job.

We're committed to helping every team member grow professionally, personally and technically while working on meaningful projects.

Global Exposure

Collaborate with international clients and multicultural teams on enterprise programmes.

Continuous Learning

Expand your expertise through mentoring, certifications and hands-on project experience.

Career Growth

Take ownership, develop leadership skills and grow your consulting career over time.

Flexible Working

Hybrid and remote collaboration designed around trust and delivering exceptional outcomes.

People First

Join a supportive culture where collaboration, respect and long-term relationships come first.

Enterprise Projects

Work on meaningful technology initiatives for leading organisations across industries.

Apply

Apply for Lead / Principal Data Engineer (Lakehouse + GenAI Integration)

One page, about two minutes. We only ask for what we actually need to have a first conversation.

Your CV

Optional. A line or two is plenty.