Senior_Principal_Agentic_Engineer
lilly
Job Description
1) Agentic resolution platform — intake to action
- Architect the agentic resolution platform end to end: intake, classification, action execution, verification, and human-in-the-loop fallback.
- Build the agent runtime and orchestration layer: agent state and memory, tool integration, multi-agent coordination patterns, and confidence-thresholded handoffs to humans.
- Define agent-decision observability, audit-ready posture, and the data contracts that let the platform consume durable fixes from upstream engineering teams.
2) LLM application patterns, knowledge, and evaluation rigor
- Design production LLM patterns: prompt engineering, retrieval-augmented generation (RAG), structured outputs, multi-model routing, and hybrid retrieval over the knowledge corpus.
- Own the knowledge-base strategy as a compounding deflection lever — every resolved incident becomes training data and structured retrieval input for future automation.
- Establish evaluation and guardrail frameworks for non-deterministic systems: automated evals, quality scoring, drift detection, and feedback loops that compound agent quality over time.
3) Cloud-native platform — build and operate
- Architect and operate Kubernetes (EKS or equivalent) at scale for container and serverless workloads supporting agentic and LLM inference traffic patterns.
- Write production platform services and internal tooling (Python or Go) that automate provisioning, deployment, and operational workflows — not just infrastructure configuration.
- Define and maintain infrastructure as code (Terraform) integrated with a major cloud's AI stack (AWS Bedrock/SageMaker, Azure AI Foundry, or Vertex AI), with secrets management and audit-ready posture for regulated environments.
4) CI/CD, observability, and developer experience
- Build and maintain CI/CD pipelines tuned for agentic and AI workloads: model and agent versioning, canary rollouts, evaluation gates, and rollback.
- Own the observability stack (Prometheus/Grafana/OpenTelemetry plus enterprise tooling) and instrument platform health, agent-decision telemetry, and model-inference metrics.
- Establish SLOs, SLIs, and reliability standards for both platform and agentic system health; design self-service patterns and golden-path templates that accelerate delivery.
5) Cross-team partnership, security, and talent development
- Partner with the Reliability Engineering team on which production patterns become agent-assisted automations, and with senior architects on the patterns that scale.
- Own security posture: network policies, pod security, secrets rotation, vulnerability scanning, and access controls in a regulated pharmaceutical environment.
- Set the engineering bar through code quality standards, architectural reviews, and role modeling; mentor senior engineers in agentic AI, LLM application patterns, and platform engineering. Influence engineering leaders to adopt automation-friendly patterns at the source, not just downstream.