Job Description :
Key Responsibilities :
- Design, develop, and maintain scalable data pipelines using Databricks, PySpark, and Apache Spark.
- Build robust ETL/ELT workflows for batch and real-time data processing.
- Develop and optimize data transformation pipelines on AWS.
- Work with AWS services including S3, Glue, IAM, Lambda, EMR, and Redshift.
- Design and implement Data Lake/Lakehouse architectures using Delta Lake.
- Collaborate with business stakeholders, data architects, and engineering teams to deliver scalable data solutions.
- Optimize Spark workloads for performance, scalability, and cost efficiency.
- Implement workflow orchestration using Apache Airflow or similar tools.
- Ensure data quality, governance, security, and operational excellence across data platforms.
- Support production environments, troubleshoot issues, and continuously improve system reliability.
- Follow best practices for version control, CI/CD, and data engineering standards.
Required Skills & Experience :- 9 - 13 years of experience in Data Engineering.
- Strong hands-on expertise in :
1. Databricks
2. PySpark
3. Apache Spark
4. SQL
- Experience designing and developing ETL/ELT pipelines and data transformation workflows.
- Strong experience with AWS services :
1. Amazon S3
2. AWS Glue
3. IAM
4. AWS Lambda
5. Amazon EMR
6. Amazon Redshift
- Good understanding of Data Lake, Lakehouse Architecture, and Delta Lake.
- Experience with workflow orchestration tools such as Apache Airflow.
- Experience with Git, CI/CD pipelines, and production support.
- Strong performance tuning, troubleshooting, and optimization skills.
- Excellent analytical, communication, and stakeholder management skills.
Preferred Skills :- Experience with Kafka and Spark Streaming.
- Knowledge of Snowflake, Apache Iceberg, and dbt.
- Experience with Terraform and Infrastructure as Code (IaC).
- AWS or Databricks certifications are an added advantage.
- Exposure to cloud-native data platforms and modern analytics architectures.
Key Competencies :- Databricks & Apache Spark
- PySpark & SQL
- AWS Data Engineering
- ETL/ELT Development
- Data Lake & Lakehouse Architecture
- Delta Lake & Airflow
- Performance Optimization
- CI/CD & Git
- Data Governance & Production Support
- Cross-functional Collaboration
(ref:hirist.tech)