Back to jobs
New

Senior Operational Engineer

US

.

 

About Nscale

Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work.

 

The role

We are looking for a Senior Operational Engineer to lead cross-service improvements that make Nscale’s engineering services more reliable, secure, controlled, and cost-aware.

You will design and deliver shared operational capabilities—such as service-readiness standards, canary checks, runbooks, reporting, alerting workflows, and automation—and partner with service teams to address systemic operational risks. This is a hands-on technical role with significant influence across engineering.

What you’ll do

  • Take ownership of the design and delivery of shared operational capabilities, including readiness checks, canaries, runbooks, service-health reporting, and operational automation.

  • Establish practical standards for service ownership, on-call readiness, alerts, dashboards, recovery procedures, and operational evidence.

  • Partner with service teams to identify and resolve recurring operational issues and cross-team blockers.

  • Design and improve integrations and workflows across Grafana, PagerDuty, Jira, Backstage, public-cloud platforms, and reporting systems.

  • Turn incident findings, change failures, and near misses into lasting engineering improvements.

  • Act as an incident commander and lead technical analysis for incidents; support corrective-action planning and delivery.

  • Build safe, observable, auditable automations that reduce manual work and improve the speed and quality of operational execution.

  • Automate operational reporting, including SLA performance, health indicators, cost trends, patching status, and action closure.

  • Mentor engineers, review operational designs, and raise the technical standard of the wider team.

  • Contribute to continuity testing and focused failure-and-recovery experiments.

What you’ll bring

  • 6–10 years’ experience in operations engineering, SRE, cloud infrastructure, platform engineering, or a similar production-focused engineering role.
  • Experience designing and operating shared capabilities for monitoring, alerting, on-call, runbooks, service readiness, and operational reporting.
  • Hands-on experience building integrations and automations across Grafana, PagerDuty, Jira, Backstage, and public cloud.
  • Experience using incident, change, service-health, continuity, patching, or cost data to drive operational improvement.
  • Ability to lead cross-service work, set practical standards, and support teams through incidents.
 

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...
Select...

Nscale uses AI-powered tools to assist in reviewing and prioritising applications against the requirements of this role. All final hiring decisions are made by humans. To learn more about how AI is used and your rights, click "Learn more" below.

Learn more