Senior Cloud Site Reliability Engineer AP
workforcenow
Job Description
Site Reliability Engineering & Operational Excellence
- Drive and implement Site Reliability Engineering (SRE) best practices across cloud platforms and services.
- Define, maintain, and improve:
- Service Level Indicators (SLIs)
- Service Level Objectives (SLOs)
- Service Level Agreements (SLAs)
- Error Budgets
- Improve service reliability, resiliency, scalability, and operational efficiency.
- Establish operational standards, reliability governance, and production readiness practices.
- Conduct Root Cause Analysis (RCA), postmortems, and reliability improvement initiatives.
- Participate in on-call rotations, incident management, and major incident resolution activities.
- Continuously improving operational processes to reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR)
Observability, Monitoring & Telemetry
- Design, implement, and maintain enterprise observability and telemetry platforms.
- Build operational dashboards, reliability scorecards, and service health monitoring solutions.
- Configure proactive alerting, anomaly detection, and incident correlation mechanisms.
- Implement centralized monitoring and telemetry using:
- Grafana
- Prometheus
- Azure Monitor
- Log Analytics
- ELK Stack / ElasticSearch
- Power BI dashboards
- Develop actionable operational metrics and telemetry reporting for engineering and leadership teams.
- Enhance visibility into infrastructure, application, Kubernetes, and platform health.
Automation & Auto-Healing Engineering
- Drive automation-first operational practices across infrastructure and platform services.
- Develop Infrastructure-as-Code (IaC) solutions using:
- Terraform
- ARM/Bicep
- Ansible
- Build operational automation scripts using:
- Python
- Bash
- PowerShell
- Develop self-healing and auto-remediation capabilities for recurring operational incidents.
- Automate infrastructure provisioning, monitoring, scaling, backup, recovery, and deployment workflows.
- Reduce manual operational effort and improve engineering productivity through intelligent automation.
Collaboration & Engineering Partnership
- Collaborate closely with:
- Cloud Engineering teams
- Product Engineering teams
- DevEx teams
- Security teams
- DBA teams
- Operations teams
- Support engineering teams in improving production readiness and operational maturity.
- Contribute to continuous improvement initiatives, reliability reviews, and operational excellence programs.
Bring your passion and together we will shine. It would also be great if you had the following:
- Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience/certification).
- Knowledge of Python, scripting, or Infrastructure-as-Code tools (e.g., Terraform, Ansible, ARM/Bicep).
- Experience managing cloud platforms (e.g., Azure, AKS, Pivotal Cloud Foundry, or equivalent).
- Strong understanding of Kubernetes and containerization concepts.
- Experience with application packaging, deployment automation, and release management.
- Solid knowledge of relational databases (MS-SQL) and exposure to NoSQL technologies (e.g., Redis, ElasticSearch, MongoDB).
- Experience with CI/CD tools (Azure DevOps, Jenkins, GitHub Actions, or similar).
- Familiarity with monitoring and logging tools (Grafana, ELK stack, Prometheus, PowerBI, etc.).
- Proficiency with Git and modern branching/merging workflows.
- Strong Linux administration and troubleshooting skills.
- Excellent problem-solving, communication, and teamwork skills.
Work Environment and Physical Demands
- Duties are performed in a typical office environment while at a desk or computer table.
- Duties require the ability to use a computer, communicate over the telephone, and read printed material, in a quiet and professional setting.
- Duties may require being on call periodically and working outside normal working hours (evenings and weekends).