Site Reliability Engineer II

mastercard

Pune NM Years Exp Posted 4h ago

Job Description

As part of the Business Operations team, you will:
• Work independently on elements of projects/processes within the Site Reliability Engineering area by applying intermediate/practical knowledge and area best practices to meet organizational standards of quality and excellence.
• Support the implementation and maintenance of high-availability systems to ensure operational stability.
• Assist in evaluating operational needs and developing technical solutions under guidance.
• Contribute to automation and scripting projects to streamline routine operational tasks.
• Troubleshoot and resolve basic to moderate system issues, escalating more complex problems as needed.
• Document operational procedures and shares knowledge with team members.
• Participate in quality checks and reviews to ensure system stability and reliability.
• Utilize experience and a comprehensive understanding of area processes and tools to make minor adjustments or enhancements to resolve identifiable issues. May manage smaller project/initiatives as an experienced individual contributor with specialized knowledge within the Site Reliability Engineering area.


Role qualifications:
The ideal candidate will apply the following skills independently in routine and moderately complex situations, requiring occasional guidance typically only in unfamiliar or highly complex scenarios. They will demonstrate growing consistency and reliability in applying the skills.

• Observability - Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement.
• Programming and Scripting - Ability to write and maintain code and scripts to automate tasks, build operational tools, and support monitoring, deployment, and incident response using languages such as Python, Go, Bash, or similar.
• Systems and Network Administration - Ability to configure, operate, and troubleshoot Linux/Unix systems and network components, applying knowledge of networking concepts, protocols, security, and system reliability.
• Cloud Computing and Infrastructure - Ability to design, deploy, and manage applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP), ensuring scalability, security, availability, and operational efficiency.
• Reliability and Scalability - Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand
• DevOps Practices - Ability to apply DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery and operations.
• Troubleshooting - Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations.
• Capacity Planning and Performance Optimization - Ability to monitor resource utilization, forecast future capacity needs, and optimize system performance to support growth, scalability, and efficient infrastructure usage.
• IT Service Management - Ability to apply IT service management principles to incident, problem, and change management, ensuring reliable service delivery, effective incident response, and continuous service improvement aligned to business needs.
• Proactive Monitoring and Improvement (SRE Applications) - The ability to use application reliability signals to anticipate issues, identify risks, and drive preventative improvements that enhance application performance and availability.

Similar Openings for You