Back to jobs
New

Senior Platform Support Engineer (Remote)

Malaysia

 

About Radian Arc.
Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.

What impact you will have 

Mission: Provide advanced technical support for customers running workloads on the GPU cloud platform, ensuring reliable operation of edge- and large-scale GPU clusters and infrastructure services.

The Senior Cloud Support Engineer acts as a technical escalation point for complex incidents, helping diagnose and resolve issues across compute, networking, storage, and orchestration layers. This role works closely with engineering and operations teams to improve platform reliability, reduce incident frequency, and enhance the overall customer experience.

 

What you’ll do  

Customer Support & Incident Management

  • Provide advanced technical support for customers operating workloads on bare-metal and virtualized GPU infrastructure.
  • Diagnose and resolve complex customer issues affecting GPU clusters, compute nodes, networking, and storage.
  • Investigate incidents across multiple layers of the stack including firmware, drivers, operating systems, and platform services.
  • Perform root cause analysis (RCA) for major incidents and contribute to long-term remediation efforts.
  • Serve as a technical escalation point for complex or high-priority support cases.

Infrastructure Troubleshooting

  • Troubleshoot issues affecting:
    • GPU compute nodes
    • Kubernetes clusters
    • Networking infrastructure
    • Local NVMe, hyperconverged and distributed storage systems
  • Analyze logs, telemetry, and monitoring signals to identify underlying causes of platform instability.
  • Investigate issues related to GPU drivers, firmware, networking, and system performance.

Security Monitoring & Incident Triage

  • Monitor and investigate security alerts generated by the platform security stack.
  • Analyze and triage alerts generated by
    • Wazuh
    • TheHive
    • Cortex
  • Validate alerts, determine impact, and escalate potential security incidents to the security engineering team.
  • Assist in collecting system telemetry, logs, and forensic data required for incident investigations.
  • Improve alert runbooks and operational procedures to reduce false positives and improve response time.

GPU HPC Workload Support

  • Provide advanced support for large-scale GPU workloads running distributed training and inference jobs.
  • Diagnose failures affecting multi-GPU and multi-node workloads.
  • Investigate performance issues impacting distributed workloads, including:
    • GPU utilization
    • Communication latency
    • Storage bottlenecks
    • networking congestion
  • Support scheduling systems used for GPU workloads, including troubleshooting:
    • Job queue failures
    • Scheduling constraints
    • Cluster resource fragmentation

High-Performance Networking Troubleshooting

  • Diagnose issues affecting high-performance networking fabrics used by distributed workloads.
  • Support environments using:
    • RDMA
    • RoCE networking
  • Investigate performance issues affecting GPU-to-GPU communication and distributed training pipelines.

Data Center Coordination

  • Coordinate with customer’s data center technicians and infrastructure teams to perform remote diagnostics and hardware interventions when required.
  • Assist in validating on-premise installations and deployments of GPU infrastructure, ensuring hardware, networking, and platform components are correctly installed and operational.
  • Support hardware troubleshooting and identify faulty components across GPU nodes, networking equipment, and storage systems.
  • Coordinate and track hardware replacements and RMA processes with vendors and data center staff.
  • Validate hardware health after replacements, including GPU nodes, NICs, DPUs, storage devices, and power components.
  • Work closely with deployment and infrastructure teams to verify service readiness after installations, expansions, or hardware maintenance activities.

Operational Excellence

  • Participate in on-call 24/7 rotations to ensure production platform availability.
  • Respond to monitoring alerts and resolve operational incidents in accordance with defined SLAs.
  • Improve operational runbooks, troubleshooting guides, and support documentation.
  • Contribute to improving incident response processes and operational tooling.

Cross-Team Collaboration

  • Work closely with platform engineering, infrastructure engineering, and networking teams to resolve systemic issues.
  • Provide feedback to engineering teams on recurring operational problems affecting customers.
  • Help translate customer issues into actionable improvements for the platform.

Automation & Tooling

  • Develop automation scripts and tools to streamline support workflows.
  • Improve observability dashboards and alerts to enable faster issue detection and resolution.
  • Contribute to automation initiatives that reduce manual intervention in operational processes.

Knowledge Sharing & Mentorship

  • Mentor medior support engineers and share troubleshooting expertise across the team.
  • Lead the creation of knowledge base articles, troubleshooting guides, and operational documentation.
  • Contribute to training initiatives that improve the team’s technical capabilities.

Technical Stack

Operating Systems

  • Linux (Ubuntu)

GPU Infrastructure

  • NVIDIA GPU platforms
  • CUDA drivers
  • GPU monitoring tools (nvidia-smi)

Platform Infrastructure

  • Kubernetes
  • Container runtimes
  • KubeVirt
  • Distributed compute environments

Networking

  • NVIDIA Cumulus
  • OOB, north-south, and east-west fabric topologies
  • TCP/IP
  • VLAN, VXLAN, OVS/OVN
  • Routing fundamentals (BGP, VRFs)
  • DNS / DHCP
  • High-performance networking (RDMA/NVLink/NCCL)

Observability

  • Grafana
  • Zabbix

Security Monitoring

  • Wazuh
  • TheHive
  • Cortex

Automation

  • Python
  • Bash
  • Ansible, Terraform

Collaboration & Documentation

  • Jira
  • Confluence
  • Zendesk
  • PagerDuty
  • Slack

What you'll need

Core Experience

  • 5+ years of experience in cloud support, infrastructure operations, or systems administration.
  • Experience supporting large-scale infrastructure environments or GPU clusters.

Systems Expertise

  • Strong Linux systems administration skills.
  • Experience troubleshooting issues across compute, networking, and storage layers.
  • Familiarity with Kubernetes platforms and containerized workloads.

GPU Infrastructure

  • Experience working with GPU hardware platforms or HPC environments.
  • Familiarity with GPU monitoring tools and debugging GPU-related issues.

Networking

  • Solid understanding of networking fundamentals including L2/L3 concepts, routing, and load balancing.
  • Ability to diagnose connectivity issues affecting distributed workloads.

Operational Mindset

  • Strong troubleshooting and incident response skills.
  • Experience participating in on-call rotations and handling production incidents.
  • Ability to perform root cause analysis and drive operational improvements.

Communication & Collaboration

  • Excellent written and verbal communication skills.
  • Ability to explain complex technical concepts to both technical and non-technical stakeholders.
  • Proven ability to collaborate effectively with cross-functional engineering teams.

 

Location & work modality: Malaysia or comparable time zone

Start: August 2026

Type of Contract: Contractor

Average 40 hours per week, 9x5 business hour support with after hour on-call response/resolution for category 1 incidents

 

What we offer
• Attractive compensation package reflecting your expertise and experience.
• A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach.
• You'll be part of a fast-growing scale-up with a mission to make a positive impact, offering an exciting career evolution.

Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.

Our inclusive responsibility
Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.

 

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...
Select...
Select...
Select...

Select all that apply:
● Linux (Ubuntu)
● Kubernetes
● KubeVirt
● NVIDIA GPU platforms
● CUDA drivers
● nvidia-smi
● Grafana
● Zabbix
● Wazuh
● TheHive
● Cortex
● Ansible
● Terraform
● Jira / Confluence / Zendesk /
PagerDuty

Select...
Select...
Select...
Select...

Answer options:
● Yes, strong direct experience
● Yes, moderate experience
● Limited experience
● No