Role & responsibilities
- Design, develop, and implement scalable Enterprise Data Warehouse (EDW) and Data Lakehouse solutions on Microsoft Azure using Azure Databricks and Delta Lake.
- Build, optimize, and maintain enterprise-grade ETL/ELT pipelines to ingest, transform, and process data from multiple sources, including relational databases, APIs, flat files, and other structured/unstructured data sources.
- Develop high-performance data transformation solutions using Python, PySpark, and Spark SQL within Azure Databricks.
- Lead the modernization and migration of legacy data warehouse environments, SQL stored procedures, Unix scripts, and traditional ETL workflows to cloud-native Azure platforms.
- Develop and orchestrate data pipelines using Azure services such as Azure Data Factory (ADF), Azure Data Lake Storage Gen2 (ADLS Gen2), Azure SQL/Synapse, and Azure Key Vault.
- Monitor, troubleshoot, and optimize Spark jobs, SQL queries, and ETL processes to improve performance, scalability, and cost efficiency.
- Design and implement robust data models and warehouse architectures following industry best practices.
- Collaborate with architects, business stakeholders, data scientists, and cross-functional teams to deliver scalable, secure, and reliable data solutions.
- Lead technical discussions, mentor data engineering teams, conduct code reviews, and promote engineering best practices including CI/CD, automated testing, and coding standards.
- Ensure data quality, governance, security, and compliance standards are incorporated into all data engineering solutions.
Preferred candidate profile
- 10-12 years of experience in Data Engineering, Business Intelligence, or Enterprise Data Warehousing.
- Strong experience designing and implementing enterprise-scale Data Warehouse and Data Lakehouse solutions.
- Hands-on expertise in Azure Databricks with at least 4 years of experience developing production-grade applications using PySpark, Spark SQL, and Delta Lake.
- Strong programming skills in Python, including core Python, Pandas, object-oriented programming (OOP), and API integration.
- Proven experience building scalable ETL/ELT frameworks and orchestrating complex enterprise data integration pipelines.
- Extensive experience with Microsoft Azure data services, including:
- Azure Data Factory (ADF)
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Azure SQL Database / Azure Synapse Analytics
- Azure Key Vault
- Expert-level SQL skills with experience in query optimization, performance tuning, window functions, and migration of legacy SQL workloads.
- Strong understanding of data modeling methodologies, including Kimball, Inmon, Data Vault, Star Schema, Snowflake Schema, and Slowly Changing Dimensions (SCD Types 1, 2, and 3).
- Experience migrating legacy on-premises data warehouse platforms such as SQL Server, Oracle, Teradata, or Unix-based ETL solutions to cloud platforms.
- Experience with performance tuning of Spark clusters, ETL workflows, and large-scale data processing environments.
- Familiarity with CI/CD pipelines and DevOps practices using Azure DevOps, Git, or GitHub Actions.
- Good understanding of Data Governance, Data Quality frameworks, security models, RBAC, and Databricks Unity Catalog.
- Strong leadership, stakeholder management, mentoring, and communication skills with experience leading technical teams and driving enterprise data initiatives.
Preferred Certifications
- Microsoft Certified: Azure Data Engineer Associate (DP-203)
- Databricks Certified Data Engineer Professional
