Job Description :
We are looking for a highly skilled Staff Software Engineer Distributed Systems to lead the design and development of large-scale, mission-critical backend platforms.
The ideal candidate will have deep expertise in building distributed systems, event-driven architectures, microservices, and cloud-native applications capable of handling high transaction volumes with low latency and high availability.
As a Staff Engineer, you will provide technical leadership across multiple engineering teams, drive architectural decisions, establish engineering best practices, and mentor senior developers while remaining hands-on with system design and development.
Key Responsibilities :
- Design and develop highly scalable, fault-tolerant distributed systems capable of handling millions of transactions and requests.
- Architect low-latency, high-throughput backend services using modern microservices principles.
- Lead architecture reviews and define technology roadmaps aligned with business objectives.
- Design event-driven systems using messaging and streaming platforms.
- Evaluate and recommend technologies, frameworks, and architectural patterns for scalability and maintainability.
- Ensure systems meet requirements for availability, reliability, security, and performance.
- Develop robust backend services using Java and Spring Boot frameworks.
- Build and maintain RESTful APIs and asynchronous communication services.
- Implement domain-driven design (DDD) and distributed system design principles.
- Create reusable frameworks, libraries, and platform components.
- Optimize application performance through profiling, tuning, and architectural improvements.
- Design and implement event streaming solutions using Kafka.
- Develop caching strategies using Redis to improve application responsiveness.
- Design scalable NoSQL data models using Cassandra.
- Implement search and analytics capabilities using Elasticsearch.
- Handle distributed transactions, eventual consistency, and fault tolerance mechanisms.
- Design data replication, partitioning, and sharding strategies.
- Build and deploy cloud-native applications on Kubernetes.
- Design containerized solutions with scalability and resilience in mind.
- Implement observability practices including logging, monitoring, and tracing.
- Improve deployment automation and operational excellence.
- Drive platform reliability and disaster recovery planning.
- Provide technical leadership and guidance to engineering teams.
- Mentor senior engineers and help elevate engineering standards.
- Conduct design reviews, code reviews, and technical discussions.
- Collaborate with Product Managers, Architects, DevOps, and Engineering Leadership.
- Participate in hiring and technical evaluations of engineering talent.
- Identify system bottlenecks and drive performance optimization initiatives.
- Implement resiliency patterns such as circuit breakers, retries, bulkheads, and rate limiting.
- Improve system observability through metrics, monitoring, and alerting.
- Lead root cause analysis and incident resolution activities.
- Define and track system SLAs, SLOs, and reliability metrics.
Required Technical Skills :
1. Programming & Frameworks :
- Strong expertise in Java (Java 8/11/17+)
- Extensive experience with Spring Boot
- Experience with Spring Ecosystem :
i. Spring MVC
ii. Spring Data
iii. Spring Security
iv. Spring Cloud
v. Spring WebFlux
2. Backend Architecture :
- Microservices Architecture
- Event-Driven Architecture
- Domain-Driven Design (DDD)
- Distributed System Design
- REST APIs
- API Gateway Patterns
- Service Discovery
- Circuit Breaker Patterns
3. Messaging & Streaming :
- Apache Kafka
- Kafka Streams
- Event Processing
- Message Queues
- Pub/Sub Architectures
4. Databases & Storage :
- Cassandra
- Redis
- Elasticsearch
- SQL Databases (PostgreSQL/MySQL preferred)
- Data Modeling
- Query Optimization
5. Cloud & Containerization :
- Kubernetes
- Docker
- Cloud Platforms (AWS/Azure/GCP)
- Helm
- Container Orchestration
- Service Mesh Concepts
6. Observability & Reliability :
- Prometheus
- Grafana
- ELK Stack
- OpenTelemetry
- Distributed Tracing
- Monitoring & Alerting
7. DevOps & CI/CD :
- Git
- Jenkins/GitHub Actions/GitLab CI
- Infrastructure as Code concepts
- Automated Testing Frameworks
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
- 915 years of experience in software engineering with significant experience in distributed systems.
- Proven experience building and operating large-scale, high-availability backend platforms.
- Strong understanding of distributed computing concepts including CAP theorem, consistency models, consensus, and fault tolerance.
- Experience designing systems that handle high concurrency and massive scale.
- Strong problem-solving, analytical, and debugging skills.
- Excellent communication and stakeholder management abilities.
- Experience working with internet-scale applications.
- Exposure to cloud-native architectures and platform engineering.
- Knowledge of service mesh technologies such as Istio.
- Experience with stream processing and real-time analytics platforms.
- Understanding of security best practices in distributed systems.
- Contributions to open-source projects or technical communities.
- Experience leading engineering teams or serving as a technical architect.
(ref:hirist.tech)