Role : GCP Data Engineer
Role Overview :
We are seeking a talented and experienced GCP Data Engineer to join our dynamic team. In this role, you will be responsible for designing, building, and maintaining scalable and reliable data pipelines on Google Cloud Platform (GCP).
You will collaborate closely with data scientists, analysts, and other engineers to deliver high-quality data solutions that drive business insights and improve decision-making. Your work will directly impact our ability to leverage data effectively, enabling us to better serve our customers and achieve our business goals.
Key Responsibilities :
- Design and implement robust data pipelines using GCP services such as BigQuery, Dataflow, Cloud Storage, and Dataproc to ingest, process, and transform large datasets for analytical and reporting purposes.
- Develop and maintain data models and schemas in BigQuery to ensure data quality, consistency, and accessibility for end-users.
- Automate data workflows and orchestrate data pipelines using Apache Airflow to improve efficiency and reduce manual effort.
- Optimize data processing performance and scalability by leveraging Apache Spark and PySpark to handle large-scale data transformations.
- Monitor data pipeline performance and troubleshoot issues to ensure data reliability and availability for critical business applications.
- Collaborate with data scientists and analysts to understand their data requirements and provide them with the necessary data infrastructure and tools to perform their analyses.
Required Skillset :
- Demonstrated ability to design, build, and maintain data pipelines on Google Cloud Platform (GCP) using services like BigQuery, Dataflow, and Cloud Storage.
- Proven expertise in data modeling, schema design, and data warehousing concepts to ensure data quality and accessibility.
- Strong proficiency in programming languages such as Python and SQL to develop data transformations and perform data analysis.
- Hands-on experience with Apache Spark and PySpark to process large-scale datasets efficiently.
- Experience with workflow orchestration tools such as Apache Airflow to automate data pipelines.
- Excellent problem-solving and communication skills to collaborate effectively with cross-functional teams.
- A Bachelor's or Master's degree in Computer Science, Engineering, or a related field is preferred.
- Ability to adapt to a fast-paced and evolving environment, with a willingness to learn new technologies and techniques.
(ref:hirist.tech)