layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
Nscale
Houston; San Francisco; Seattle
Source: Nscale careers · View original posting
From Nscale's posting. “We” and “our” refer to the employer.
Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.
At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work.
If you join our team, you'll be contributing to building the technology that powers the future.
Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets — tickets, alerts, hardware faults, and customer issues — across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world.
Experience.
3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services).
Communication.
Clear written notes, concise updates, and reliable follow-through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately.
GPU and hardware troubleshooting.
Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers — reseating components, swap testing, working via BMC/out-of-band management — through to preparing RMA evidence. A strong technical base here is required, not a learning goal.
Linux.
Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate.
Networking.
Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role.
Ticketing and ITSM discipline.
Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation.
Observability foundations.
Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review.
Scripting and automation basics.
Comfortable reading and writing simple Bash or Python, and using Git for version control.
Platform and DC fundamentals.
Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background.
Growth mindset.
Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior.
Adaptability.
Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed.
Nice to Have
High-performance fabrics and GPU-HPC: exposure to RDMA/InfiniBand, link-level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL-based troubleshooting, or NVLink concepts.
High-performance storage: exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale, including basic storage–network troubleshooting.
OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet-scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar).
Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role.
Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault.
Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time.
At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.
Highly competitive package (base + equity) with reviews every 12 months. 🚀
Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. ✨
Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
We treat you as humans first. 🫶🏽 Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work.
Equal Opportunities Statement
We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there’s anything we can do to accommodate your specific situation, please let us know.
The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Compensation
$100,000 — $140,000 USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:
Here.
Nscale does not accept unsolicited candidate submissions from recruitment agencies.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: