Back to jobs
New

MLOps Engineer

Abu Dhabi, UAE

About AI71:

AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.

The Role:

As an MLOps Engineer you set the ML infrastructure and reliability strategy across AI71's platform, including how LLMs and other deep learning models are deployed, fine-tuned, and served at scale. You own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive multi-quarter ML infrastructure strategy. You are a force multiplier.

What You'll Do:

  • Define ML infrastructure architecture across the platform: model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP)
  • Set direction for ML system reliability: monitoring, latency / throughput / availability targets, and incident response across research and production environments.
  • Mentor senior MLOps engineers; raise the operational bar across multiple teams.
  • Drive cross-team initiatives that improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).
  • Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.
  • Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.

What You'll Bring:

  • 10+ years of MLOps, ML infrastructure, or machine learning engineering with history of architectural ownership.
  • Proven track record architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.
  • Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong Python proficiency
  • Mentorship record — engineers you have grown now operate independently at higher levels.
  • Deep comfort architecting ML systems that run in both managed SaaS and on-premises / disconnected air-gapped environments.
  • Kubernetes at architectural depth — GPU scheduling, multi-tenancy, operators, and the failure modes of distributed workloads on shared clusters.
  • Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams. 

Strong Preference:

  • Ownership of production reliability at platform level: SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.
  • Architecture-level experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.
  • Deep GPU systems knowledge: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.
  • Model optimization strategy at portfolio level: quantization (FP8, AWQ, GPTQ), speculative decoding, with measurable cost or latency outcomes across multiple systems.
  • Experience in regulated or security-constrained environments — compliance-driven architecture, model governance, lineage, audit, and secrets management.
  • On-prem / air-gap ML delivery architecture experience at scale.
  • Track record maturing MLOps practice in a growing organization: standards, platform abstractions, and paved paths that outlived your involvement.
  • Bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.

Nice to Have: 

  • Conference speaking, technical writing, or industry thought leadership.
  • Open-source contributions to inference, serving, or ML infrastructure projects, particularly maintainer-level involvement.
  • C/C++ or CUDA kernel experience for performance-critical paths.
  • Arabic language skills.

Why AI71:

  • Mission-Driven Work: Work on cutting-edge AI applications with a talented and passionate team, solving real-world challenges in critical sectors.
  • Unparalleled Opportunity: This is a chance to innovate and solve real-world challenges using AI at a company with unique access to world-leading models and resources.
  • Career Growth: We offer competitive compensation, benefits, and significant career growth opportunities as a foundational member of the team.
  • World-Class Environment: Enjoy a flexible working environment and the latest tools & technologies needed to do your best work.

 

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...
Select...
Select...

As part of our commitment to fostering a diverse and inclusive workplace, we invite applicants to voluntarily provide gender and ethnicity information. This data is for internal reporting only, kept confidential, and has no impact on hiring decisions. Sharing is completely optional — your application will be considered equally whether or not you provide this information.