Job Description :
Urgent Opportunity for Lead - Data Engineer
Role Overview :
This pivotal role involves leading the design, development, and optimization of robust and scalable data pipelines and data architectures.
The Lead Data Engineer will be instrumental in transforming raw data into actionable insights, working closely with data scientists, product managers, and business analysts to understand their data needs and deliver high-impact solutions.
This position directly influences our ability to make data-driven decisions, enhance product features, and drive significant business outcomes by ensuring data quality, accessibility, and performance across our platforms.
Key Responsibilities :
- Design, develop, and maintain highly scalable and efficient data pipelines and ETL processes using Python, PySpark, and SQL, ensuring timely and accurate data delivery for analytical and operational needs.
- Lead the architectural planning and implementation of cloud-native data warehousing and data lake solutions on AWS, optimizing for performance, cost-efficiency, and security to support advanced analytics and machine learning initiatives.
- Mentor and guide a team of data engineers, fostering best practices in data engineering, code quality, and system reliability to elevate the team's technical capabilities and project delivery.
- Collaborate cross-functionally with data scientists, product owners, and business stakeholders to translate complex data requirements into robust technical specifications and deliver data solutions that drive strategic business value.
- Implement and enforce stringent data governance, quality, and security standards across all data assets, ensuring compliance and maintaining the integrity and trustworthiness of our data.
- Optimize complex SQL queries, data models, and data processing jobs to enhance system performance, reduce latency, and improve resource utilization across our data ecosystem.
Required Skillset :
- Demonstrated expertise in designing, building, and optimizing large-scale data pipelines and ETL processes using Python, PySpark, and advanced SQL.
- Profound proficiency in cloud-based data warehousing and data lake solutions, specifically leveraging AWS services such as S3, Redshift, Glue, EMR, Lambda, and Athena.
- Strong understanding of data modeling, data architecture principles, and database design across both relational and NoSQL paradigms.
- Proven ability to lead technical initiatives, mentor junior engineers, and drive best practices in a fast-paced, agile environment.
- Exceptional problem-solving skills with a proactive approach to identifying and resolving complex data challenges and architectural bottlenecks.
- Excellent communication and interpersonal skills, capable of articulating complex technical concepts clearly to diverse audiences, including senior leadership and non-technical stakeholders.
- A strong commitment to continuous learning, staying abreast of emerging data technologies, and applying innovative solutions to business problems
(ref:hirist.tech)