Job Description :
Highly skilled Data Engineer with strong experience in ETL development, SQL, database management, Qlik (Qlik Compose and Qlik Replicate), Data Lakes, PySpark, and AWS Glue. The ideal candidate will design and build scalable data pipelines, enable efficient data replication and transformation, support analytics and reporting needs, and ensure high-quality data integration across enterprise systems.
Key Responsibilities :
- Analyze multiple data sources and create data mapping documents, data lineage, and transformation logic to ensure clear documentation of data flows for analytics, reporting, and governance.
- Design, develop, and maintain efficient ETL pipelines using AWS Glue, PySpark, and SQL, ensuring seamless ingestion and transformation of structured and semi-structured data.
- Utilize Qlik Compose to automate data warehouse design, schema generation, and ETL processes, accelerating delivery of curated, analytics-ready data models.
- Develop and maintain applications to integrate data from various sources into Data Lakes (e.g., AWS S3) and other platforms, ensuring consistency and reliability.
- Work extensively with relational databases such as Oracle, SQL Server, and PostgreSQL, writing and optimizing complex SQL queries, stored procedures, and data models.
- Perform query tuning and database performance optimization to enhance system efficiency.
- Use PySpark for large-scale distributed data processing and transformations, enabling efficient handling of big data workloads.
- Perform root cause analysis for data discrepancies, missing data, and performance bottlenecks, and implement effective solutions.
- Optimize the efficiency, scalability, cost-effectiveness, and quality of development processes, including AWS Glue jobs, Spark workloads, SQL queries, and Qlik-based data pipelines.
- Collaborate with cross-functional teams including Data Analysts, BI Developers, Architects, and DevOps to ensure efficient development and enhancement of applications.
- Follow best practices, reusable frameworks, and coding standards while considering system-wide performance and impact.
Required Skills & Qualifications
- Strong experience in ETL development, advanced SQL, Qlik Compose and Qlik Replicate, PySpark, and AWS Glue.
- Hands-on experience with relational databases such as Oracle, SQL Server, PostgreSQL, or similar systems.
- Experience working with Data Lake architectures
- Strong understanding of data modelling, data warehousing concepts, and distributed data processing.
Preferred Skills
- Experience with AWS services such as S3, Athena, Redshift, and Lambda.
- Familiarity with workflow orchestration tools such as Airflow, CI/CD pipelines, and data governance or lineage tools.
- Experience with Qlik Compose / Replicate
- Exposure to Agile development practices.
Key Competencies
- Strong analytical and problem-solving skills
- Attention to detail and focus on data quality
- Ability to work collaboratively in a team environment
- Effective communication and stakeholder management skills
(ref:hirist.tech)