Role Overview :
As a PySpark Developer, you will serve as a critical architect within our data engineering team, responsible for designing, developing, and maintaining high-performance data pipelines that process massive datasets.
You will collaborate closely with data scientists, business analysts, and infrastructure engineers to transform raw data into actionable insights that drive strategic decision-making.
By optimizing complex ETL workflows and ensuring the scalability of our data platforms, you will directly influence the efficiency of our analytics ecosystem and provide the foundation for data-driven business outcomes across the organization.
This role offers the opportunity to work in a fast-paced environment across our offices in Pune, Chennai, Bangalore, Mumbai, Hyderabad, or Gurgaon, contributing to projects that impact millions of end-users.
Key Responsibilities :
- Architect and implement robust data processing pipelines using PySpark to ensure seamless data ingestion and transformation for downstream analytical applications.
- Optimize existing Spark jobs and SQL queries to improve processing speed and reduce infrastructure costs, directly enhancing the performance of our reporting dashboards.
- Partner with cross-functional stakeholders to translate complex business requirements into scalable technical solutions that support long-term data accessibility.
- Implement rigorous data quality checks and validation frameworks to maintain the integrity and reliability of data assets for executive leadership.
- Mentor junior team members and conduct thorough code reviews to uphold high engineering standards and foster a culture of technical excellence.
Required Skillset :
- Demonstrated expertise in building and deploying large-scale data processing solutions using PySpark and Python, with a deep understanding of distributed computing principles.
- Proven ability to design efficient data models and write complex SQL queries to extract meaningful insights from structured and semi-structured datasets.
- Strong proficiency in managing cloud-based data environments and orchestrating workflows using tools like Airflow or similar scheduling frameworks.
- Exceptional communication skills, with the ability to articulate complex technical concepts to non-technical stakeholders and influence project direction.
- A collaborative mindset that thrives in a hybrid work environment, showing adaptability to evolving project needs and cross-regional team dynamics.
- A Bachelor's or Master's degree in Computer Science, Engineering, or a related quantitative field, supported by 5 - 7 years of hands-on experience in data engineering.
PySpark Developer - ETL Tools • Gurgaon