Key Responsibilities:
- Design, develop, and optimize scalable data pipelines using Python.
- Work extensively with Pandas and PySpark for data processing and transformation.
- Build and maintain efficient data engineering workflows.
- Develop and manage Delta Tables and Parquet datasets for optimized data storage.
- Improve data pipeline performance, scalability, and reliability.
- Collaborate with cross-functional teams to deliver high-quality data solutions.
- Ensure code quality through testing and best engineering practices.
Must-Have Skills:
- Strong proficiency in Python.
- Hands-on experience with Pandas and PySpark.
- Experience in Data Engineering and Workflow Optimization.
- Working knowledge of Delta Tables.
- Experience with Parquet file format.
- Strong problem-solving and debugging skills.
Good-to-Have Skills:
- Experience with Databricks.
- Hands-on knowledge of Apache Spark.
- Experience with DBT (Data Build Tool).
- Experience with Apache Airflow.
- Advanced Pandas performance optimization techniques.
- Knowledge of PyTest and DBT Testing Frameworks.
Wissen Technology - Python Data Engineer • Bangalore