We're looking for a curious and innovative Site Reliability Engineer to join our team at Visa. Here, you'll be part of a diverse group of problem-solvers who ensure billions of transactions flow seamlessly across the world's largest payment network.
What You'll Do
• Lead Investigations: Use your detective skills to solve complex technical puzzles and prevent future incidents.
• Drive Innovation: Contribute to our evolution from traditional application to cloud-native solutions.
• Collaborate Globally: Work with talented engineers across the world to build, support, and deploy application services.
Why You'll Love It
• Growth Opportunities: Regular learning sessions, mentorship programs, and exposure to cutting-edge technology.
• Work-Life Integration: Hybrid work model (2-3 days in office) with flexible scheduling.
• Inclusive Culture: Join a team that actively promotes diverse perspectives and collaborative problem-solving.
Your Experience & Skills
We encourage you to apply even if you don't meet every requirement. We value potential and enthusiasm over perfection.
Core Skills (Some combination of):
• Programming Languages: Basic knowledge & experience with C++, Java, Python, JavaScript, HTML, CSS.
• Application Technologies: Basic understanding on modern application technologies including Tomcat, Apache, Spring Boot, SQS, JBoss, IBM MQ, IBM DataPower, Hazelcast, Flink, Connect Direct, and SSL. Skilled in designing.
• Systems and Networking: Basic knowledge of Linux/Unix systems and networking. Basic knowledge on cloud platforms such as AWS, Azure, and GCP for scalable and resilient infrastructure solutions.
• Containerization and DevOps: Basic knowledge in containerization technologies like Kubernetes and Docker.
• Monitoring and Logging: Basic knowledge in monitoring and logging systems including Prometheus, Grafana, Datadog, and the ELK stack.
Technical Areas You'll Grow In:
• Observability & Performance: Master advanced monitoring, tracing, and performance optimization techniques.
• Automation & Intelligence: Build smart alerting systems and automated remediation workflows.
What Makes You Thrive:
• You're energized by solving complex problems.
• You believe in automation over manual processes.
• You enjoy mentoring others and sharing knowledge.
• You're comfortable with ambiguity and rapid change.
• You value building reliable systems over quick fixes.
Job Responsibilities
• Build in-depth expertise on the 24*7 systems of Visa that support our merchants, Developers, Support team.
• Support & publish the issues within SLAs for established issue priorities, fix issues/bugs.
• Collaborate with the cross functional team ,development and product team to improve the overall development process. Support and track the production support activities to enhance efficiency and productivity.
• Participate in post-release monitoring and validation. Collaborate with the DEV team to ensure no 'release issues' occur in PROD.
• Manage the team who does the application and infrastructure performance plans/models for a highly scalable, low-latency, highly-available, and high-throughput payment processing system.
• Contribute to capacity planning and disaster recovery exercises.
• Support in triaging and troubleshooting of performance degradation incidents in the production environment with multiple stakeholders in regular intervals.
Required Skills
• Basic Coding Knowledge in programming languages like Java/Python and scripting languages.
• Should have experience in defining production baselines regularly and reliability of system.
• Ability to work independently and with managing a small set of engineers.
• Professional work experience in highly scalable web services.
• Exposure to containerized micro-services architecture and stacks is an added plus.
Preferred Skills
• Understanding of Disaster Recovery methodologies.
• Experience working with Agile teams.
• Knowledge of monitoring tools like Splunk/Keynote/Graphana is added plus
• A Bachelor's degree in Computer Science or Engineering; a Master's degree is a plus.
• Experience working in fast-paced 24*7 environments.
• Excellent oral and written communication skills.
• Knowledge of GenAI, Chatgpt,LLM is a plus
The Team Culture
You'll join a collaborative team that:
• Celebrates diverse perspectives and approaches to problem-solving.
• Values teaching and learning from each other.
• Promotes work-life balance and sustainable on-call rotations.
• Encourages innovation and experimentation.
• Champions personal growth and career development.
Impact & Growth
In this role, you'll:
• Shape the reliability standards for global payment systems.
• Mentor and be mentored by talented engineers.
• Drive automation and observability initiatives.
This is a hybrid position. Expectation of days in office will be confirmed by your hiring manager.
Qualifications
Basic Qualifications
• 8+ years of relevant work experience and a Bachelors degree, OR 11+ years of relevant work experience
Preferred Qualifications
• 13 to 16 Yr years of work experience with a Bachelor's Degree or more than 12 years of work experience with an Advanced Degree (e.g., Masters, MBA, JD, MD).
• People manager skills is an added advantage,
• Bachelor's or Master's degree in Computer Science or related field, or equivalent experience.
• We value hands-on experience and continuous learning over specific degrees.
Skills Required
Spring Boot, C++, Apache, Java, Python