SRE & Deployments Engineer
recrew
Job Description
- Build and operate on-call rotations, alerting pipelines, and incident response workflows across product pods; drive post-incident reviews and structural fixes to reduce repeat pages
- Define, instrument, and enforce SLOs for compliance-critical services; own error budgets and escalation paths
- Own the end-to-end deployment pipeline: reproducible packaging, artifact signing, SBOM generation, and install/upgrade automation across cloud and restricted-access environments
- Build remote observability and diagnostics tooling for deployments where direct environment access is limited or forbidden
- Build and maintain access-governance infrastructure: audited break-glass flows, session recording, and deploy-only paths that eliminate standing production access
- Operate and migrate core shared services with full runbook coverage, capacity planning, and DR validation
- Maintain and improve IaC, environment parity, and deployment automation across cloud and on-prem/hybrid environments
Must Have Criteria
- 3–5 years in a DevOps or SRE role with real uptime accountability — carried a pager and made structural changes to reduce alert volume or MTTR
- Hands-on Linux administration and container operations in production environments
- Infrastructure-as-Code experience managing cloud resources at production scale
- Proficiency with at least one major cloud provider at an infrastructure operations level
- Scripting fluency in at least one language for automation, tooling, and diagnostic workflows
- Demonstrated experience with observability stacks — metrics, logging, and alerting pipelines in production