We are looking for an experienced Big Data + Python Engineer to design, develop, and optimize large-scale data processing solutions. The ideal candidate will possess strong expertise in Python, PySpark, distributed data processing, and cloud-based data platforms, with a proven ability to build scalable ETL pipelines and data engineering solutions.
Key Responsibilities :
- Design, develop, and maintain scalable data pipelines and big data processing frameworks using Python and PySpark.
- Build and optimize batch and real-time data processing workflows for large-scale datasets.
- Develop ETL/ELT solutions to ingest, transform, validate, and process structured and unstructured data.
- Utilize Spark technologies including Spark SQL, DataFrames, and MLlib to support analytics and machine learning workloads.
- Perform data cleansing, transformation, feature engineering, and data quality validation.
- Design and optimize data storage solutions using SQL and NoSQL databases.
- Implement workflow orchestration and automation using tools such as Apache Airflow, Jenkins, or GitHub Actions.
- Collaborate with data scientists, analysts, and business stakeholders to deliver data-driven solutions.
- Monitor and improve data pipeline performance, scalability, reliability, and cost efficiency.
- Troubleshoot production issues and perform root cause analysis for data processing failures.
- Implement coding standards, testing frameworks, and CI/CD best practices for data engineering projects.
- Contribute to cloud migration, modernization, and data platform optimization initiatives.
Required Skills & Experience :
- 7-10 years of experience in Data Engineering, Big Data Development, or related roles.
- Strong proficiency in Python with a focus on clean, maintainable, and scalable code development.
- Hands-on experience with PySpark for batch and streaming data processing.
- Strong expertise in :
1. Spark SQL
2. Spark DataFrames
3. Spark MLlib
4. Distributed Data Processing
- Experience with data manipulation and feature engineering using :
1. Pandas
2. NumPy
3. Data visualization libraries
- Strong knowledge of SQL and database optimization techniques.
- Experience with data storage technologies such as Hive, Cassandra, or other SQL/NoSQL databases.
- Experience with workflow orchestration and automation tools including Apache Airflow, Jenkins, or GitHub Actions.
- Familiarity with CI/CD pipelines and Agile development methodologies.
- Strong analytical, debugging, and problem-solving skills.
Synechron - Big Data Engineer - Python • Bangalore