Sr Technical Consultant
blueyonder
Job Description
- Review and act on incidents, service requests, infrastructure requests and provisioning failures logged by implementation teams and platform users.
- Own L2/L3 cloud-platform issues from initial triage through recovery, validation, communication, root-cause analysis and closure.
- Troubleshoot Kubernetes pods, deployments, replicas, services, events, health checks, node pools, scheduling and resource constraints.
- Diagnose container image-pull, startup, shutdown, registry-authentication, rollout and workload-reconciliation failures.
- Investigate CPU, memory, GPU, quota, region, placement and capacity issues affecting platform workloads.
- Troubleshoot workload and event-driven autoscaling using Kubernetes metrics and technologies such as KEDA.
- Support Azure resource provisioning, provider operations, resource lifecycle workflows and reconciliation between desired and actual state.
- Diagnose ACR, private endpoint, firewall, allowlist, CIDR, DNS, TLS and runtime-connectivity problems.
- Trace logs, metrics and distributed telemetry from the workload through collectors, exporters and observability ingestion pipelines.
- Support Elastic/Logstash/Kibana ingestion, index mappings, access, dashboards, alerts and environment filters.
- Investigate operational issues involving MongoDB/Atlas, SQL, Redis, NFS and related managed data or storage services.
- Support certificates, service principals, API keys, image-pull credentials and infrastructure credential rotation.
- Support regional releases, environment configuration, disaster-recovery workflows and post-deployment validation.
- Develop automation to improve platform reliability, reduce manual provisioning and shorten incident recovery time.
- Maintain runbooks, dashboards, alerts, known-error records and diagnostic procedures for use by platform support teams.
- Participate in capacity planning, incident reviews, change reviews, release readiness and an agreed production on-call rotation.