Key Responsibilities :
- Design, develop, and deploy data engineering solutions using Azure Databricks and Azure cloud services.
- Build, optimize, and maintain scalable ETL/ELT pipelines using PySpark, SQL, and Azure Data Factory.
- Develop and maintain high-performance data pipelines to support increasing data volume and business requirements.
- Integrate multiple data sources and develop new data source integrations for enterprise data platforms.
- Perform data transformation, cleansing, validation, and optimization for analytical workloads.
- Collaborate with business analysts, data architects, and cross-functional teams to deliver end-to-end data solutions.
- Monitor, troubleshoot, and optimize data processing jobs for performance and reliability.
- Implement best practices for data governance, security, and quality.
- Participate in code reviews, documentation, and CI/CD deployment processes.
Required Skills :
- Strong hands-on experience with Azure Databricks, PySpark, SQL, and Azure Data Factory (ADF).
- Experience in designing, developing, and deploying data engineering solutions on Microsoft Azure.
- Expertise in building scalable and optimized ETL/ELT pipelines.
- Strong knowledge of Apache Spark architecture and performance tuning techniques.
- Experience working with Delta Lake, Spark SQL, and data optimization strategies.
- Hands-on experience with Azure Data Lake Storage (ADLS Gen2), Blob Storage, and other Azure storage services.
- Experience integrating structured, semi-structured, and unstructured data from multiple data sources.
- Good understanding of data warehousing concepts and dimensional data modeling.
- Experience with Git, Azure DevOps, and CI/CD pipelines.
- Knowledge of data security, governance, and monitoring best practices.
- Strong analytical, debugging, and problem-solving skills.
- Excellent communication and collaboration skills.
(ref:hirist.tech)