DevOps Specialist Engineer

clarusadvisers

Bengaluru 6 Years Exp Posted 10h ago

Job Description

  • Design, build, operate, and continuously improve large-scale cloud-native production systems.
  • Define and own SLIs, SLOs, SLAs, error budgets, and reliability objectives.
  • Lead production incident response, on-call operations, root-cause analysis, and reliability improvements.
  • Build and manage CI/CD pipelines, Kubernetes platforms, infrastructure automation, and multi-environment deployments.
  • Implement Infrastructure as Code using Terraform and deployment automation using tools such as ArgoCD.
  • Develop production-grade observability across metrics, logging, and distributed tracing.
  • Work with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, Azure Monitor, and AWS CloudWatch.
  • Operate AI/ML and GenAI workloads in production, addressing reliability, performance, model drift, output variance, and train/serve skew.
  • Support MLOps/LLMOps platforms and AI control-plane capabilities such as model gateways and guardrails.
  • Implement Kubernetes/Docker-based solutions and cloud-native networking across multiple environments.
  • Conduct load and performance testing using tools such as LoadRunner, k6, or JMeter.
  • Implement chaos engineering, capacity planning, autoscaling, and resilience testing.
  • Drive cloud and AI FinOps, including GPU, inference, and token-cost attribution and optimization.
  • Implement security controls including RBAC, least privilege, secrets management, deployment approvals, and segregation of duties.
    • Collaborate with software engineering, security, data/AI, and product teams to improve platform reliability and operational excellence.

Similar Openings for You