Senior Platform Support Engineer (Remote)
Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks.
Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.
What impact you will have?
As a Senior Platform Support Engineer, you’ll be one of the main technical escalation points for Radian Arc’s production AI and GPU infrastructure. You’ll take on the incidents that need deeper investigation across Linux, Kubernetes, NVIDIA GPU infrastructure, networking and storage. The expectation is that you can get into the detail, work across multiple layers and own the technical investigation through to recovery or a clear engineering/vendor escalation. This is not a coordination-only support role. We need someone who has already operated and troubleshot GPU and Kubernetes environments in production and can help raise the technical capability of the wider support team.
You’ll work closely with Engineering, Data Center Operations, Service Delivery and our technology partners.
What you’ll do
- Own complex and high-priority technical incidents.
- Troubleshoot across Linux, Kubernetes, GPU compute, networking, storage and supported platform services.
- Use logs, monitoring data, hardware diagnostics and recent changes to work out where the problem sits.
- Drive recovery actions and workarounds where appropriate.
- Validate service recovery and contribute to root-cause analysis for major or recurring incidents.
- Assess customer and SLA impact during major incidents and provide the SDM with clear technical status, recovery actions and realistic next steps.
- Diagnose GPU node failures, degraded GPU health, driver issues and infrastructure-related GPU performance problems.
- Use NVIDIA diagnostic and telemetry tools to investigate GPU, driver, firmware and hardware issues.
- Troubleshoot production Kubernetes issues across nodes, pods, scheduling, networking and the container runtime.
- Diagnose problems that span Kubernetes, Linux and the underlying infrastructure rather than treating each layer separately.
- Troubleshoot network issues across L2/L3, routing, VLAN/VXLAN and high-performance networking.
- Investigate storage availability, connectivity and performance issues.
- Escalate platform defects and systemic problems to Engineering with enough evidence to make the escalation useful.
- Work with Data Center Operations where hardware needs to be replaced or checked onsite.
- Engage OEMs and technology partners when specialist support is required.
- Stay technically engaged through the escalation rather than simply handing the issue off.
- Build or improve Bash/Python tooling for diagnostics and repetitive support tasks.
- Work with Engineering and Observability teams to improve alerts, telemetry and supportability.
- Turn repeatable L2 investigations into better runbooks, tooling or L1 procedures.
- Mentor Platform Support Engineers and help improve their troubleshooting capability.
- Contribute to readiness reviews for new platforms, infrastructure and customer environments.
What you’ll need
- Showcase 5+ years of hands-on experience in infrastructure engineering/support, cloud/platform operations, systems engineering or production support.
- Advanced Linux administration and troubleshooting skills.
- Strong production Kubernetes experience, including cluster, node, pod, scheduling, networking and container-runtime troubleshooting.
- Hands-on experience operating and troubleshooting NVIDIA GPU infrastructure in production.
- Experience diagnosing GPU hardware, drivers and system-level GPU issues using NVIDIA or equivalent tooling.
- Strong networking knowledge, including TCP/IP, VLANs, routing and structured L2/L3 troubleshooting.
- Experience troubleshooting storage availability, connectivity and performance.
- Strong experience using observability platforms and working with logs, metrics and telemetry during complex incidents.
- Strong Bash and/or Python scripting skills.
- Proven experience owning significant production incidents through diagnosis, recovery and escalation.
- Ability to troubleshoot across multiple infrastructure layers and identify the actual source of a problem.
- Experience working directly with Engineering teams, infrastructure specialists and/or hardware vendors.
- Clear written and spoken English.
Nice to Have:
- Hands-on experience with NVIDIA HGX/DGX-class systems, including H100, H200, B200/B300 or comparable platforms.
- NVIDIA DCGM and other GPU diagnostic or telemetry tooling.
- RDMA/RoCE and high-performance Ethernet.
- NVIDIA Spectrum-X or Cumulus networking.
- NCCL and multi-GPU/multi-node communication troubleshooting.
- Large-scale bare-metal GPU environments.
- Distributed or high-performance storage such as DDN, WEKA, VAST or similar.
- BMC, BIOS and firmware troubleshooting.
- Infrastructure automation using Ansible, Terraform or similar.
- HPC, distributed AI training or large-scale inference environments.
- Experience with private cloud and virtualisation technologies such as CloudStack, KVM/KubeVirt, OVS/OVN, VyOS and Citrix NetScaler/WAF.
Location: Malaysia or a comparable APAC time zone preferred
Employment type:
- Contractor or Employee of Record
- Average 40 hours per week
- Paid as a monthly rate
What we offer
• Attractive compensation package reflecting your expertise and experience.
• A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach.
• You'll be part of a fast-growing scale-up with a mission to make a positive impact, offering an exciting career evolution.
Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.
Our inclusive responsibility
Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.
Apply for this job
*
indicates a required field

