Back to jobs
New

Cloud Operations Engineer

India

Are you looking for an opportunity to help solve one of today's biggest business challenges? AI is changing the pace of business, and organizations everywhere are struggling to help their workforce, partners, and customers keep up. At Litmos, we're building the Learning Acceleration Platform that helps organizations build human capability faster—and we're looking for people who are passionate about making a meaningful impact for customers while growing alongside a collaborative, people-first team.

Litmos is the Learning Acceleration Platform that helps organizations build capability faster, adapt at the speed business changes, and scale learning to anyone, anywhere. Combining an intuitive platform, AI-powered capabilities, trusted content, expert services, and a broad ecosystem of integrations, Litmos helps organizations accelerate workforce productivity, improve customer adoption and retention, enable high-performing partners, and reduce organizational risk through continuous learning.

Organizations such as Hewlett Packard Enterprise, Graco, Sabre, and Russell Mineral Equipment trust Litmos to accelerate learning across their workforce, partners, and customers. Today, more than 11K customers with 30 million learners across 150 countries and 37 languages use Litmos to build the capabilities their organizations need to succeed. Backed by Francisco Partners, one of the world's leading technology investment firms, we're investing in the future of learning—and the people who are building it. Learn more at www.litmos.com.

 

We are seeking a Cloud Operations Engineer to join our cloud engineering team and support the operational excellence of our Azure-based SaaS platform. This mid-level role is focused on the day-to-day operational tasks that keep our infrastructure stable, performant, and compliant. You will monitor cloud infrastructure, manage patching schedules, triage and respond to alerts, maintain comprehensive incident documentation, and track critical SLA metrics. Working closely with the Cloud Engineering Manager and broader cloud team, you will develop deeper expertise in cloud operations while contributing directly to platform reliability and uptime.

Key Responsibilities:

Infrastructure and Systems Monitoring (IST)

  • Monitor Azure infrastructure and on-premises systems for health, performance, and security using established observability and monitoring platforms
  • Proactively identify and investigate anomalies, performance degradation, and potential issues before they impact users
  • Maintain dashboards and reporting tools that provide visibility into infrastructure utilization, performance metrics, and trending data
  • Work with the team to establish and refine SLOs (Service Level Objectives) and ensure monitoring aligns with agreed-upon targets

Patch Management and Scheduling

  • Schedule and coordinate security patches and infrastructure updates across Azure resources and supporting systems
  • Communicate patch windows to stakeholders and coordinate maintenance activities to minimize business impact
  • Execute patch deployments following established procedures and validate successful implementation
  • Maintain patch schedules and compliance documentation to ensure security and regulatory requirements are met

Alert Triage and Incident Response

  • Triage incoming alerts, assess severity and impact, and determine appropriate response actions
  • Respond to operational incidents including diagnosing root causes, implementing fixes, and restoring service when issues occur
  • Escalate complex issues to the Cloud Engineering Manager or senior engineers with clear context and preliminary findings
  • Execute runbooks and documented procedures for common operational scenarios
  • Participate in on-call rotation and respond to critical alerts outside normal business hours

Incident Logging and Documentation

  • Document all operational incidents with clear timelines, root causes, and resolutions in incident management system
  • Maintain comprehensive runbooks and operational procedures for common infrastructure tasks and incident responses
  • Create and update monitoring alerts, thresholds, and escalation policies based on operational learnings
  • Participate in post-incident reviews (blameless postmortems) to identify improvements and prevent recurrence

SLA Tracking and Reporting

  • Monitor and report on Service Level Agreement (SLA) metrics including uptime, availability, and response times
  • Track Mean Time To Recovery (MTTR), Mean Time Between Failures (MTBF), and other key reliability metrics
  • Generate regular reports and dashboards for internal stakeholders and customer-facing SLA obligations
  • Identify trends and patterns in operational metrics to inform reliability improvements and capacity planning

Collaboration and Continuous Improvement

  • Collaborate with the Cloud Engineering Manager and cloud engineering team on operational improvements and automation opportunities
  • Identify opportunities to automate manual operational tasks and provide input on Infrastructure-as-Code improvements
  • Support CI/CD pipeline operations and assist with troubleshooting deployment-related issues
  • Actively participate in team knowledge-sharing sessions and contribute to operational documentation

Required Qualifications

  • 3+ years of experience in cloud operations, infrastructure support, or systems administration roles
  • Hands-on experience operating Microsoft Azure infrastructure and services in production environments
  • Strong experience with monitoring, alerting, and observability tools (e.g., Azure Monitor, Application Insights, or similar platforms)
  • Solid understanding of infrastructure fundamentals including networking, virtual machines, storage, and databases
  • Experience with scripting in PowerShell, Python, or similar automation languages
  • Familiarity with incident management practices, runbooks, and operational procedures
  • Experience tracking SLAs, uptime metrics, and producing operational reports
  • Strong problem-solving skills and ability to troubleshoot complex infrastructure issues
  • Excellent communication skills and ability to work effectively in a collaborative team environment
  • Ability to work on-call schedules and respond to critical incidents outside normal business hours

Preferred Qualifications

  • Azure certifications such as AZ-104 (Azure Administrator) or AZ-900 (Azure Fundamentals)
  • Experience with Infrastructure-as-Code tools (Bicep, ARM templates, or Terraform)
  • Familiarity with SRE (Site Reliability Engineering) principles and practices
  • Experience with Azure Kubernetes Service (AKS) or container orchestration platforms
  • Background in supporting SaaS or multi-tenant cloud platforms
  • Experience with CI/CD pipelines and deployment automation
  • Exposure to AWS or experience with multi-cloud environments
  • Experience with incident management tools and practices

Key Competencies

  • Operational Excellence: Commitment to maintaining reliable, secure, and performant infrastructure with attention to detail
  • Proactive Problem-Solving: Ability to identify issues before they impact users and implement preventative measures
  • Continuous Learning: Eagerness to expand technical knowledge and stay current with cloud technologies and operational practices
  • Collaboration: Strong ability to work effectively with team members, provide clear context during escalations, and support peers
  • Attention to Detail: Precision in incident documentation, SLA tracking, and operational procedures
  • Adaptability: Ability to handle multiple priorities, respond to urgent incidents, and thrive in a dynamic environment
  • Communication: Clear and concise communication of technical issues, status updates, and escalations to team members and stakeholders

Salary Range: ₹22,00,000 to ₹28,00,000 plus 10% bonus

 

 

As a learning company we believe in the potential of everyone; if you don't have experience in all the details mentioned in this job post, then we still encourage you to apply and we'll get back to you as soon as we can.  

We are an equal opportunity workplace employer. We are committed to the values of Equal Employment Opportunity and provide accessibility accommodations to applicants with physical and/or mental disabilities.

Applicants will receive consideration for employment without regard to their age, race, religion, national origin, ethnicity, age, gender (including pregnancy, childbirth, et al), sexual orientation, gender identity or expression, protected veteran status, or disability.

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf