Job Summary :
We are looking for a Senior/Principal Data Architect with deep expertise in both Google Cloud Platform (GCP) and Databricks to lead the architecture and evolution of enterprise-scale data platforms. This is a senior individual contributor role responsible for defining architecture standards, governance, data strategy, and platform modernization across analytics, machine learning, and operational workloads.
Mandatory Skills :
GCP :
- BigQuery
- Dataproc
- Cloud Composer (Airflow)
- Pub/Sub
- Dataflow
- Google Cloud Storage (GCS)
- Dataplex
- Vertex AI
- GCP IAM
- VPC Service Controls
- Multi-region and High Availability (HA) Architecture
- GCP Cost Optimization
Databricks :
- Unity Catalog
- Delta Lake
- Delta Live Tables (DLT)
- MLflow
- Photon Engine
- SQL Warehouses
- Databricks Asset Bundles (DAB)
- Databricks Workflows
- Workspace Architecture
- Cluster Policies
- Job Compute
- Delta Lake Optimization
Data Architecture & Engineering :
- Data Lakehouse Architecture
- Data Modeling (Kimball, Data Vault 2.0, Domain-Driven Design)
- PySpark
- SQL
- Apache Spark 3.x
- Real-Time & Streaming Architecture
- Apache Kafka
- Data Governance
- Data Cataloging
- Metadata Management
- Data Quality Frameworks
- CI/CD
- Terraform
- Infrastructure as Code (IaC)
Required Experience :
- 15 to 22 years of overall experience.
- Minimum 8+ years in Data Engineering or Data Architecture.
- At least 3 years in a dedicated Data Architect role.
- Expert-level experience with GCP Data Services.
- Expert-level experience with Databricks Lakehouse Platform.
- Strong hands-on expertise in PySpark and SQL.
- Experience designing enterprise-scale analytics, machine learning, and operational data platforms.
- Strong experience with real-time streaming architectures using Kafka, Pub/Sub, or Spark Structured Streaming.
Preferred Certifications :
- Google Cloud Professional Data Engineer
- Google Cloud Professional Cloud Architect
- Databricks Certified Data Engineer Professional
- Databricks Certified Associate Developer for Apache Spark
- dbt Certified Developer
- AWS Solutions Architect
Job Responsibilities :
- Platform Architecture & Design : Define enterprise data platform architecture across GCP and Databricks. Design and govern Lakehouse architecture with Bronze, Silver, and Gold layers. Architect batch, micro-batch, and real-time data ingestion frameworks.
- Data Modeling & Governance : Design logical and physical data models. Establish governance standards for cataloging, lineage, access control, and data classification.
- Engineering Standards : Define standards for PySpark, SQL development, testing, and pipeline design. Implement CI/CD practices, automated testing, and schema management.
- Leadership & Stakeholder Management : Partner with Data Engineering, Data Science, Analytics, and Product teams. Lead architecture reviews and provide technical direction.
(ref:hirist.tech)