Platform Support Engineer (Remote-Malaysia)
Platform Support Engineer (Malaysia-Remote)
About Radian Arc.
Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.
What impact you will have
As a Platform Support Engineer, you’ll support Radian Arc’s production AI and GPU infrastructure across Linux, Kubernetes, GPU compute, networking and storage. This is a hands-on technical role. You’ll investigate alerts and incidents, work out what is actually failing, resolve issues where you can, and escalate when deeper engineering or physical intervention is needed.
We are not looking for someone whose main experience is monitoring queues and passing tickets on. You should be comfortable getting onto systems, looking at logs and metrics, running diagnostics and troubleshooting problems yourself.
You’ll work closely with the Senior Platform Support Engineers, Engineering, Data Center Operations and Service Delivery teams.
What you’ll do
- Support production GPU infrastructure and platform services.
- Troubleshoot Linux systems, Kubernetes, GPU nodes, networking and storage.
- Investigate Kubernetes node, pod, service and cluster issues using logs, metrics and standard tooling.
- Check GPU health, driver state, device visibility, utilization and hardware status using NVIDIA and platform tools.
- Troubleshoot common network issues including DNS, routing, VLANs and connectivity.
- Investigate common storage, filesystem and mount issues.
- Respond to alerts and support incidents and assess the likely customer and SLA impact.
- Take recovery or remediation actions where the issue is within your scope.
- Verify that services are healthy again after an incident or maintenance activity.
- Provide the SDM with clear technical status and recovery updates during customer-impacting incidents.
- Escalate complex issues with useful diagnostics, logs and a clear summary of what has already been checked.
- Escalate suspected platform defects or systemic issues to Engineering with the appropriate diagnostics and technical evidence.
- Work with Data Center Operations when physical hardware intervention is needed.
- Keep technical records and handovers clear and up to date.
- Improve runbooks and troubleshooting guides based on what you see in production.
- Help identify recurring incidents, monitoring gaps and noisy alerts.
- Write or adapt simple Bash or Python scripts where this makes support easier or more repeatable.
What you’ll need
- 3+ years of hands-on experience in infrastructure support, systems administration, cloud/platform operations or production support.
- Strong Linux troubleshooting skills.
- Hands-on Kubernetes experience, including troubleshooting nodes, pods, services and common cluster issues.
- Experience supporting GPU infrastructure or GPU-enabled environments, including basic NVIDIA GPU and driver diagnostics.
- Good networking fundamentals, including TCP/IP, DNS, VLANs, routing and connectivity troubleshooting.
- Working knowledge of storage and filesystem concepts.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Zabbix or similar.
- Ability to work with logs, metrics and system telemetry and use them to narrow down a problem.
- Basic Bash and/or Python scripting skills.
- Experience working in a structured production support environment.
- Good judgement on when to keep troubleshooting and when to escalate.
- Clear written and spoken English.
Useful experience
- NVIDIA HGX/DGX or similar GPU server platforms.
- NVIDIA tools such as nvidia-smi, DCGM or similar.
- RDMA/RoCE or other high-performance networking.
- Large-scale bare-metal infrastructure.
- Distributed or high-performance storage.
- BMC/IPMI/Redfish and server hardware diagnostics.
- Ansible or other infrastructure automation tools.
- Experience with private cloud or virtualization platforms such as CloudStack, KVM or KubeVirt.
- HPC or AI infrastructure environments.
- Supporting business-critical or 24/7 services.
What we offer
• Attractive compensation package reflecting your expertise and experience.
• A great work environment characterized by friendliness, international diversity, flexibility, and a hybrid-friendly approach.
• You'll be part of a fast-growing scale-up with a mission to make a positive impact, offering an exciting career evolution.
Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.
Our inclusive responsibility
Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.
Apply for this job
*
indicates a required field

