Back to jobs
New

Design Reliability Engineer — Power and Energy

Houston; US

About Nscale

At Nscale, we are building the infrastructure for the AI revolution. Nscale is developing cutting-edge, sovereign generative AI solutions powered by a new generation of high-performance, sustainable data centers and GPUs built specifically for AI workloads.

The rapid growth of artificial intelligence is driving unprecedented global demand for compute and GPU infrastructure. Positioned at the heart of this transformation, Nscale is delivering the platforms that will enable the next decade of innovation.

Working closely with the world's most advanced AI technology providers — including Microsoft and NVIDIA — Nscale integrates next-generation compute hardware and GPU clusters across both partner and self-delivered facilities. This is a unique opportunity to join Nscale's journey, play a pivotal role in delivering transformative AI infrastructure, and help shape a company designed to scale rapidly across North America, Europe, and beyond.

About the Role

We are hiring a Design Reliability Engineer to own the analysis that proves our systems will hit their availability targets, and the discipline that makes those numbers real.

Nscale's Power and Energy group develops behind-the-meter generation colocated with our data center campuses: fleets of hundreds of reciprocating engines, battery energy storage, and medium- and high-voltage distribution operating as islanded microgrids, serving loads with contractual availability commitments measured at the GPU. At this scale and redundancy, availability is not a brochure number. It is an engineered quantity: allocated from committed SLAs down to systems and equipment, modeled with rigor, defended against common-mode risks, and eventually measured live against the models that predicted it. This role owns that quantity, end to end, across power and mechanical scopes together rather than in silos.

This is a hands-on technical leadership role built for someone who runs with minimal oversight. You will own the availability model of record for each campus, the FMEA program across major equipment, and the reliability foundations of our digital twin: the living, data-fed model that will carry RAM analysis from design studies into day-to-day operations. You will direct the RAM consultants and reliability data work supporting the portfolio, challenge methodology on its merits, and put analysis in front of executives that they can bet commercial commitments on.

You will sit within the Power and Energy technology organization, working alongside the electrical, controls, mechanical, and pipeline engineering managers whose designs your models test, and with the commissioning and operations teams whose data will eventually close the loop. You will be the owner's technical authority on reliability, directing the work of consultants rather than simply accepting it, and resolving the questions where textbook RAM practice does not account for a system this large, this redundant, or this consequential.

Location: Houston, TX (remote considered within the US, with regular site and vendor travel)

What You'll Be Doing

RAM modeling and availability assurance

  • Own the availability model of record for each generation campus: system-level RAM models spanning generation, electrical distribution, fuel supply, and cooling, built and maintained as living engineering assets.
  • Allocate committed SLA targets down through the system: availability budgets by subsystem, redundancy requirements, MTTR assumptions, and sparing and maintenance strategies that make the top-level number achievable.
  • Apply the right method for each question — Monte Carlo simulation (BlockSim or equivalent) where it earns its complexity, closed-form analytical models where they are faster and more auditable — and defend the choice on its merits.
  • Hunt common-mode and dependent failures relentlessly: shared fuel supply, shared cooling, control system dependencies, and site-wide events that redundancy counts conceal.
  • Report availability the way commitments are written: both single-path and contracted-capacity views, with sensitivities that show which assumptions carry the number.

FMEA and reliability engineering

  • Own the FMEA/FMECA program for major equipment and systems: engines and generators, BESS, switchgear and transformers, fuel gas systems, and cooling infrastructure, establishing baselines and keeping them current as designs mature.
  • Curate the failure rate and repair data underpinning every model: industry sources such as IEEE 493 and OREDA, OEM data challenged rather than transcribed, and field data as it accumulates.
  • Turn analysis into design influence: redundancy configuration, single-point-of-failure treatment, equipment selection input, and testability and maintainability requirements fed into the engineering teams while designs can still change.
  • Bring reliability analysis into design reviews, HAZOPs, and vendor evaluations as a routine discipline, not a report delivered after decisions are made.

Digital twin and reliability data systems

  • Own the reliability core of Nscale's digital twin: the model architecture, data structures, and analytical methods that let availability be recomputed on the fly as system state, configuration, and failure data change.
  • Define the operational data requirements — event capture, failure coding, downtime attribution, and run-hour tracking — so the twin is fed by trustworthy data from day one of operations.
  • Work with our software and data teams to move RAM analysis from static studies into instrumented, queryable tooling, and champion analytical approaches that are automatable and auditable rather than locked in desktop tools.
  • Establish the feedback loop: measured availability versus model prediction, model recalibration, and reliability growth tracking from commissioning onward.

Delivery, commissioning, and operations support

  • Support commissioning and startup with reliability input: burn-in and reliability run design, failure tracking during startup, and acceptance criteria grounded in the availability model.
  • Lead and support root cause analyses of significant failures and availability events, and drive corrective actions back into designs, models, and standards.
  • Develop Nscale's reliability engineering standards, methods, and reference models, so each campus builds on the last instead of starting over.

Team and vendor leadership

  • Direct RAM and reliability consultants: set the basis they work to, review and challenge their methodology and deliverables, and integrate their output into one coherent picture across power and mechanical scopes.
  • Coordinate with data center reliability and operations teams so that availability is engineered and measured across the full path to the GPU, not just to the fence line.
  • Grow the internal reliability engineering capability over time, building a small team as the portfolio scales.
  • Communicate complex reliability analysis clearly to project leadership and executives, with honest assessments of confidence, sensitivity, and risk.

About You

  • 5+ years of reliability engineering experience on power generation, process, or mission-critical facilities, with significant time owning RAM analysis for large, redundant systems.
  • Deep RAM modeling capability across both discrete-event simulation (BlockSim, Raptor, or equivalent) and closed-form analytical methods, with the judgment to know which the question deserves.
  • Proven FMEA/FMECA leadership on major rotating, electrical, or process equipment, and fluency with the standard failure data sources (IEEE 493, OREDA, IEEE 3006 series or similar) and their limitations.
  • Experience allocating availability targets from commercial commitments down to systems, and defending the resulting analysis to executives, customers, or insurers.
  • A sharp eye for common-mode and dependent failure mechanisms, and a track record of finding the risks that redundancy arithmetic hides.
  • Working software capability: comfort scripting analyses (Python or similar) and structuring reliability data for automation, not just operating desktop tools.
  • Experience feeding reliability analysis into live design processes on major capital projects, and the standing to influence engineering decisions with it.
  • Demonstrated ability to run a scope with minimal oversight: establishing the basis, setting the pace, and escalating with solutions rather than problems.
  • Exposure to digital twin, telemetry-driven reliability, or operational availability measurement programs is a strong advantage, as is CRE certification or equivalent.
  • Bachelor's degree in Mechanical, Electrical, Chemical, or Reliability Engineering or a related field; PE license or CRE preferred.
  • Willing and able to travel to project sites and vendor facilities regularly (typically 15–25%, higher during commissioning campaigns).

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and our people are driving it.

  • Highly competitive US compensation package (base + bonus + equity), with performance reviews every 12 months
  • Comprehensive medical, dental, and vision coverage
  • 401(k) retirement plan with company match
  • Generous PTO plus US federal holidays
  • A career-defining opportunity to be an early member of one of the fastest growing AI infrastructure companies in the world
  • A human-first approach: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy needed to get the job done.

Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$130,000 - $200,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...

Nscale uses AI-powered tools to assist in reviewing and prioritising applications against the requirements of this role. All final hiring decisions are made by humans. To learn more about how AI is used and your rights, click "Learn more" below.

Learn more