Experience : 68 Years
Location : Chennai / Pune (Hybrid - 3 Days Work from Office)
Employment Type : Direct Contract with Haparz
About the Role :
We are looking for a Senior Data Engineer to architect and build the foundational Data Engine for an AI Agent Development Platform (ADP). The role focuses on creating secure, agent-ready data systems that power autonomous AI agents for Lease Abstraction, CAM Reconciliation, and other Commercial Real Estate use cases.
The ideal candidate should have strong expertise in building scalable ETL/ELT pipelines, designing AI-native data architectures, and enabling secure data access for AI agents through governed interfaces.
Key Responsibilities :
1. Build the Data Engine :
- Design and develop robust ETL/ELT pipelines using Python and SQL.
- Transform complex relational datasets into secure, structured, and agent-consumable data views.
- Develop data materialization strategies and optimize data processing workflows.
2. Architect AI-Native Data Platforms :
- Design and manage multi-layered data environments using :
1. PostgreSQL for transactional workloads
2. Vector databases for embeddings and semantic retrieval
3. Graph databases for relationship mapping and knowledge representation
- Build scalable and high-performance data models supporting AI-driven applications.
3. Develop Secure Agent-Data Interfaces :
- Build and maintain Model Context Protocol (MCP) adapters and secure SDK endpoints.
- Enable AI agents to access data through governed interfaces without direct database connectivity.
- Implement schema mapping and data governance controls.
4. Build Knowledge Graph & RAG Infrastructure :
- Develop knowledge graphs to model relationships between leases, entities, invoices, and CAM charges.
- Configure Retrieval-Augmented Generation (RAG) pipelines, document ingestion frameworks, and metadata management processes.
- Support semantic search and contextual retrieval capabilities.
5. Implement Data Security & Multi-Tenant Isolation :
- Design logical data isolation mechanisms for multi-tenant environments.
- Implement data masking, access controls, and zero-trust security patterns.
- Ensure solutions align with SOC 2 and GDPR compliance standards.
6. Cloud & Platform Engineering :
- Manage and optimize AWS data services, including RDS, S3, and event-driven architectures.
- Monitor performance, scalability, reliability, and data availability.
- Collaborate with AI Engineers, Platform Architects, and Product teams to deliver production-grade AI data solutions.
Required Experience :
- 6-8 years of experience in Data Engineering.
- Strong expertise in Python, SQL, ETL/ELT pipeline development, and PostgreSQL.
- Hands-on experience with AWS data services including RDS and S3.
- Experience with Vector Databases and Graph Databases for AI-native applications.
- Exposure to Redshift and modern AI data architectures.
- Understanding of secure, scalable data platform design.
Preferred Experience :
- Experience with dbt and data transformation frameworks.
- Exposure to Model Context Protocol (MCP), Knowledge Graphs, and RAG infrastructure.
- Familiarity with Temporal workflow orchestration.
- Understanding of LangChain and LangGraph ecosystems.
- Experience working with Commercial Real Estate data domains is highly preferred.
Why Join :
- Build foundational data infrastructure for next-generation AI agents.
- Work on Agentic AI, Knowledge Graphs, RAG, and Vector Databases.
- Opportunity to architect enterprise-grade AI data platforms from the ground up.
- Collaborative environment with exposure to cutting-edge AI technologies.
(ref:hirist.tech)