Staff Software Engineer - Fleet Management
.
About Nscale
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
About the role
Nscale is hiring a Staff Software Engineer to build Fleet Manager — the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale.
This role sits at the intersection of distributed systems, infrastructure automation, and physical hardware. You'll own domain-level architecture within Fleet Manager: Python-based systems that manage the entire operational lifecycle of our compute infrastructure, from initial device enrollment through multi-day burn-in testing to ongoing health monitoring and automated remediation. The problems are challenging and the stakes are high — the software you design and build determines how quickly and how reliably Nscale scales its GPU fleet to meet demand.
This is an opportunity to shape a foundational platform early, setting the patterns and standards that engineers across Fleet Manager build on.
What you'll work on
- Device provisioning and enrollment: automation that takes bare-metal GPU nodes and network switches from first power-on to production-ready — BMC configuration, DHCP reservations, and provisioning state machines.
- Burn-in and validation: multi-day testing workflows that qualify hardware before it enters, and re-enters, the fleet.
- Workflow orchestration: durable, event-driven state machines that span multiple days, survive crashes, resume from checkpoints, support human-in-the-loop approval gates, and let thousands of concurrent idempotent workflows run without stepping on each other.
- GPU health monitoring and self-healing: detection, diagnosis, and automated remediation workflows that keep nodes healthy in production.
- Network configuration: switch lifecycle automation and network state management across the fleet.
- Integrations: keeping Fleet Manager consistent with datacenter inventory tooling (DCIM, NetBox), bare-metal provisioning systems (MAAS, Ironic, IPMI), credential stores, and monitoring infrastructure.
- Observability: structured logging, metrics, distributed tracing, and tooling that lets operators troubleshoot effectively.
Responsibilities
- Domain-level technical direction. Own the architecture for a major Fleet Manager domain — such as provisioning, validation, or remediation — influencing engineers across the team and adjacent squads.
- Design and build production-grade automation. Implement device provisioning, burn-in testing, network configuration, and hardware health validation workflows in Python.
- Engineer for reliability and auditability. Treat idempotency, resumability, checkpointing, retries, replay, and failure handling as first-class design concerns.
- Integrate broadly. Connect Fleet Manager with datacenter infrastructure management systems, cloud orchestration platforms, and bare-metal provisioning tools.
- Create leverage through standards. Establish shared patterns, libraries, conventions, and operational runbooks that other engineers build on.
- Own production outcomes. Operate what you build with strong observability, alerting, incident response, and day-2 operational discipline.
- Mentor and influence. Raise the bar through design reviews, implementation guidance, and operational best practices.
- Use AI to accelerate delivery while maintaining architectural coherence.
Requirements
- Extensive experience designing, building, and operating distributed systems in production, ideally in infrastructure automation, workflow tooling, or platform engineering.
- Strong proficiency in Python — Fleet Manager is built entirely in Python.
- Strong understanding of event-driven and workflow architecture, including reliable delivery, idempotency, retries, replay, and failure handling.
- Track record of delivering automation systems from ambiguous requirements to production, with hands-on day-2 operations experience (monitoring, incident response, performance optimization).
- Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority.
- You are driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement.
- You use AI tools like Claude or Cursor as a core part of your development workflow to create leverage, increase quality, and accelerate delivery.
- Excellent communication skills to build consensus with stakeholders, both internally and externally, in a fast-paced, high-agency environment.
Preferred
- Experience with workflow orchestration tools like Temporal, Airflow, Prefect, or similar
- Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems
- Bare-metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation
- Experience building hardware lifecycle automation: provisioning, validation, testing, or remediation workflows
- GPU infrastructure experience: health monitoring, burn-in testing, or cluster management
- HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE)
- Deep knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP
- Open-source contributions in infrastructure automation or cloud-native tooling
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Salary Range
$220,000 - $320,000 USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
Apply for this job
*
indicates a required field
