Job description
Operationalizes AI: deployment, monitoring, cost, and governance for models and LLM applications running in production.
What you'll do
- Build CI/CD pipelines for model and LLM application deployment
- Stand up model monitoring, versioning, and rollback systems
- Manage cost, latency, and reliability of production AI workloads
- Implement governance controls: access management, audit logging, usage tracking
- Partner with the security team on AI-specific compliance requirements
- Support multi-cloud AI infrastructure across AWS, Azure, and GCP
Skills
What you bring
- 4+ years in DevOps or platform engineering, including 1–2+ years specific to ML/AI systems
- Experience with Kubernetes, Docker, and Terraform or an equivalent infrastructure-as-code tool
- Familiarity with MLOps tooling such as MLflow, Kubeflow, SageMaker, or Vertex AI
- Strong scripting ability in Python or Go, with CI/CD pipeline experience
- Working understanding of model versioning, monitoring, and cost optimization practices
Nice to have
- Experience with LLM-specific observability tools such as LangSmith or Arize
- AWS, Azure, or GCP cloud certification