AWS Platform Engineer
Here is a comprehensive Job Description tailored specifically for a Senior AWS DevOps / Platform Engineer transitioning into MLOps, structured to attract candidates with strong consulting or client-delivery backgrounds.
Senior AWS Platform / DevOps Engineer (MLOps Focus)
About Indicium AI
Indicium AI was Anthropic’s first European launch partner, a Preferred Anthropic Partner, and sits among a handful of organizations globally trusted to deploy Claude at an enterprise scale. Named the 2026 Databricks Consulting Partner of the Year and backed by Databricks Ventures, we build and deploy production AI systems for complex enterprises, making AI a genuine competitive advantage.
About the Role
We are seeking a Senior AWS Platform / DevOps Engineer to design, automate, and scale our cloud infrastructure with a core focus on enabling our next-generation AI/ML applications.
In this role, you will leverage your expertise in AWS, Kubernetes (EKS), Infrastructure as Code (Terraform), and CI/CD automation to build resilient platform foundation, while expanding your hands-on scope into MLOps, deploying and orchestrating AI workloads, vector stores, and model inference pipelines (e.g., Amazon Bedrock, SageMaker).
This is an ideal opportunity for a high-performing DevOps or Platform Engineer (especially those with a cloud consulting or professional services background) who wants to build production-grade AI infrastructure at scale.
Key Responsibilities
-
Cloud & Platform Automation: Design, build, and maintain modular Infrastructure as Code (IaC) using Terraform and AWS best practices across multi-account environments.
-
Kubernetes Orchestration: Manage production AWS EKS clusters, Helm charts, and ingress controllers; configure dynamic autoscaling (Karpenter/HPA) to handle specialized CPU/GPU compute workloads.
-
GitOps & Continuous Delivery: Establish automated GitOps delivery pipelines (ArgoCD or Flux) to ensure zero-downtime releases for microservices and AI application services.
-
MLOps Integration: Build and optimize cloud infrastructure supporting Machine Learning workflows, including model serving, vector databases (e.g., OpenSearch, pgvector), and API integrations with services like Amazon Bedrock and SageMaker.
-
Security & Reliability: Implement fine-grained IAM policies, network security (VPCs, security groups), secrets management, and observability stacks (Prometheus, Grafana, CloudWatch).
-
Developer Enablement: Partner with software developers and data/AI teams to streamline internal developer workflows, reduce friction, and standardize deployment patterns.
Key Qualifications
-
Core DevOps / IaC Expertise: 4+ years of hands-on experience in DevOps, Platform, or Infrastructure Engineering using AWS, Terraform, and Kubernetes (EKS) in production environments.
-
Containers & Orchestration: Deep understanding of Docker, Helm, Kubernetes architecture, and continuous delivery tools (ArgoCD, Flux, GitHub Actions, or GitLab CI).
-
Interest & Exposure to MLOps / AI: Practical exposure or strong motivation to build infrastructure for AI/ML workloads (e.g., model serving, SageMaker, Amazon Bedrock, vector databases, GPU node groups).
-
Consulting or High-Growth Background: Experience working in cloud consulting, system integration, or fast-paced delivery environments where adapting to new tools and client requirements is second nature.
-
Scripting & Automation: Proficiency in scripting/programming with Python, Bash, or Go.
-
Strong Communication: Ability to articulate architectural trade-offs, collaborate across software and data teams, and document infrastructure patterns.
Preferred Qualifications
-
AWS Certified Solutions Architect or DevOps Engineer - Professional.
-
Hands-on experience with GPU node scaling via Karpenter on EKS.
-
Familiarity with MLOps frameworks like MLflow, Kubeflow, or Ray.
-
Experience implementing cost optimization (FinOps) strategies for heavy compute/AWS workloads.
Why Indicium AI
-
Deep Anthropic Partnership: Work at the bleeding edge as a Preferred Anthropic Partner, with direct access to partner teams, official training, and early access to unreleased capabilities.
-
High Autonomy Culture: Sharp, high-agency teams with no middle-management layers or corporate theater.
-
Frontier-Level L&D: Generous learning budget, dedicated research time, and unencumbered access to state-of-the-art AI tools.
-
Top-Tier Benefits: Competitive pay with performance bonuses, comprehensive health coverage, 401k Employer Match, generous PTO, flexible holidays, parental leave, and a paid company shutdown the last week of December.
The anticipated base salary range for this role is $135,000 - $190,000. In addition to base pay, this position may be eligible for an annual discretionary bonus. An individual's final salary offer will be determined based on a variety of factors, including geographic location, experience, specialized skills, and qualifications. This compensation range is subject to updates or modifications at the company’s discretion
Apply for this job
*
indicates a required field
