Back to jobs
New

Site Reliability Team Leader

Tel Aviv

At Optimove, we believe people are capable of more than a single job description. You’re not hired just to fill a position- you’re empowered to shape it, grow it, and make it your own.
We call this being Positionless.
And Positionless isn’t just our culture. It’s our product.
Optimove is the creator of Positionless Marketing, an AI-powered platform that gives every marketer the power to analyze, create, launch, and optimize independently. The result is faster execution, deeper personalization, and 88% greater campaign efficiency.
Recognized as a Visionary in Gartner’s Magic Quadrant, we partner with leading brands like Sephora, Staples, and Entain. Today, more than 500 Optimovers across NYC, London, Tel Aviv, Scotland, Brazil, Estonia, and beyond are building the future of marketing together, in an environment that actively encourages ownership and growth, with two out of every three managers promoted from within.
If you’re looking for a place where you can do more, be more, come grow with us.

Are you passionate about building reliable, scalable, and highly available production systems? Do you enjoy solving complex engineering challenges through automation, observability, and software engineering? Optimove is looking for an SRE Team Lead to join our global SRE organization. In this role, you'll lead a team of SREs while partnering with engineering teams globally to improve the reliability,
scalability, and operational excellence of our cloud platform. As an SRE Team Lead, you'll combine hands on technical leadership with people management,
driving engineering focused reliability initiatives while growing and mentoring a team of SREs across automation, deployment processes, observability, and operational excellence in
production. This is an Israel based role leading a global team of SRE engineers based in Israel and Ukraine.

Responsibilities

  • Team Leadership: Manage, mentor, and grow a global team of SRE engineers across Israel and Ukraine, running regular 1:1s, setting goals, and supporting career development. Own hiring,
    onboarding, and performance management for the team.
  • Distributed Team Operations: Keep the team working as one unit across sites, with shared standards, consistent handoffs, and clear ownership so that reliability work is not fragmented by location or timezone.
  • Technical Direction: Set the technical roadmap for the team's reliability, automation, and observability initiatives, and stay hands on enough to guide design decisions and unblock complex problems.
  • Reliability Engineering: Guide the design and implementation of solutions that improve the reliability, availability, and scalability of our production platform, and ensure the team proactively identifies and eliminates operational risks.
  • Automation and Platform Engineering: Prioritize and oversee the build of internal tools and automation that eliminate manual operational work, improve engineering productivity, and streamline production workflows.
  • Observability: Drive the team's roadmap for monitoring, alerting, dashboards, and production visibility, reducing alert fatigue and strengthening operational insight across services.
  • Production Rollouts: Oversee the team's work on deployment processes using modern release strategies such as Canary, Blue/Green, and Feature Flags, ensuring safe and reliable releases.
  • Production Reliability and Incident Response: Act as an escalation point for critical production incidents, guide root cause analysis, and ensure long term preventive improvements are implemented and tracked.
  • On-Call and Operational Excellence: Own the team's on-call rotation and coverage across sites,participate as needed, and drive continuous improvements that reduce operational toil and prevent future incidents.
  • Cloud and Infrastructure: Oversee the team's work on our Kubernetes based cloud platform,CI/CD pipelines, and production infrastructure running on GCP and AWS.
  • Cross Team Partnership: Represent the SRE team in planning and decision making with Software Engineering, DevOps, DBA, and Product leadership, and align the team's priorities with broader engineering goals.


Requirements

  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Infrastructure Engineering, including some experience leading or mentoring other engineers.
  • Hands on experience operating Kubernetes in production environments.
  • Strong experience working with public cloud platforms (GCP or AWS).
  • Strong programming and scripting skills (Python preferred, Go or Bash are a plus).
  • Experience designing and building automation and internal engineering tools.
  • Experience working with CI/CD pipelines and modern deployment methodologies.
  • Hands on experience with observability platforms such as Datadog, Prometheus, or Grafana.
  • Strong understanding of Linux, networking, distributed systems, and cloud native architectures.
  • Excellent troubleshooting, debugging, and root cause analysis skills.
  • Strong communication skills and proven experience working with remote colleagues across sites and timezones, with a genuine interest in growing into people leadership.
  • Fluent English, written and spoken, as the team works across multiple countries

    Advantages
  • Prior formal people management or team lead experience.
  • Experience leading or coordinating engineers who are not co-located.
  • Experience with Infrastructure as Code (Terraform, Ansible, etc.).
  • Experience with messaging and distributed technologies such as Kafka, Pub/Sub, or Redis.
  • Experience supporting modern deployment strategies such as Canary, Blue/Green, or Feature Flags.
  • Familiarity with OpenTelemetry and modern observability tooling.
  • Understanding of Site Reliability Engineering principles, including SLIs, SLOs, and error budgets.
  • Experience working in large scale SaaS production environments.
  • Relevant cloud or Kubernetes certifications (GCP, AWS, CKA, CKAD).

    Why Join Us?

    At Optimove, SRE is an engineering discipline, not just an operations function. You'll work on large-scale cloud infrastructure serving enterprise customers while building automation, improving observability, and enabling engineering teams to deliver software safely and reliably. You'll join a collaborative global SRE organization working with modern cloud-native technologies, Kubernetes, distributed systems, and large-scale SaaS workloads. You'll have the opportunity to lead impactful reliability initiatives, influence engineering best practices, and solve complex production challenges that directly improve the experience of both our customers and engineering teams.

Optimove is an equal opportunity employer. We consider all qualified applicants fairly, without regard to race, ethnicity, gender, age, religion, disability, sexual orientation, or any other characteristic protected by applicable law. If you require any adjustments during the recruitment process, please let us know.


By submitting this application, you agree that Optimove will process your personal data in accordance with applicable data protection laws, including GDPR. For details, see our Privacy Policy.

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf