layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
US, CA, Santa Clara
Source: NVIDIA careers · View original posting
From NVIDIA's posting. “We” and “our” refer to the employer.
NVIDIA is seeking an engineering director to lead the software teams behind our bare-metal datacenters and cloud compute infrastructure. The organization supports hundreds of megawatts of datacenter capacity already online, with more capacity coming. You will build the software that makes this growing physical infrastructure available as reliable, secure, and efficient compute services for NVIDIA engineering.
Our team provides NVIDIA’s continuous integration (CI) infrastructure: the environment where hardware, firmware, drivers, networking, and system software are integrated and tested on their path to production. You will enable engineering teams to bring up preproduction systems, reproduce failures, validate changes, and move platforms toward production readiness. The fleet spans multiple hardware generations and maturity levels: x86 and Arm servers, GPUs, DPUs, high-speed networking, storage, and rack-scale, liquid-cooled systems.
Platforms such as Grace Blackwell and Vera Rubin illustrate the breadth of compute and interconnect technology involved. This role combines hands-on systems judgment with leadership of the teams making that diversity manageable at scale.
15+ overall years of experience in software engineering, distributed systems, or cloud infrastructure, including 7+ years leading engineering teams and experience managing engineering managers.
Direct engineering ownership of the underlying services of a public or private compute cloud, such as AWS EC2, Google Compute Engine, Azure Compute, OCI Compute, or a comparable infrastructure-as-a-service platform.
Strong technical depth in distributed systems, Linux, virtualization, containers, and the networking and storage services that support large compute fleets.
Experience building software and automation for bare-metal infrastructure, including server provisioning, hardware inventory, firmware or operating-system lifecycle management, and recovery across heterogeneous systems.
Experience running production infrastructure with demanding availability requirements, including failure isolation, incident response, observability, and reliable deployment and recovery mechanisms.
The ability to guide architecture, evaluate implementation choices, and debug complex interactions among hardware, firmware, drivers, operating systems, and distributed services with senior engineers.
A record of sustained ownership as platforms evolve, with measurable improvements in reliability, scalability, or engineering delivery. Strong communication, collaboration, and people-development skills.
A degree in computer science, computer engineering, or a related discipline, or equivalent experience.
Experience building or operating OpenStack infrastructure, particularly Nova, Neutron, Cinder, or Ironic, or contributing to related open-source projects.
Engineering leadership spanning compute control planes and site reliability, including multi-tenant isolation, scheduling, placement, and capacity allocation.
Experience with GPU clusters, rack-scale computing, NVLink, InfiniBand or high-speed Ethernet, and AI or high-performance computing workloads.
Experience supporting preproduction hardware, new product introduction, or hardware-in-the-loop CI, including repeatable validation environments and systematic regression isolation.
Familiarity with liquid-cooled, high-density datacenters and how power, thermal constraints, cooling, and component health affect fleet availability and scheduling and delivery of compute platforms across multiple regions, datacenters, or hybrid-cloud environments, including fleet upgrades and workload migration with minimal service disruption.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.
You will also be eligible for equity and benefits .
Applications for this job will be accepted at least until October 13, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted:
Marvell · Santa Clara, CA; US-CA - Santa Clara; Irvine, CA; Austin, TX; Burlington, VT; Boise, ID; Hudson Valley, NY
NVIDIA · US, WA, Redmond; US, MA, Westford; US, NC, Durham; US, CA, Santa Clara
Apple · Austin, Texas; Beaverton, Oregon; Santa Clara, California
ServiceNow · Santa Clara, California, United States
ServiceNow · Santa Clara, California, United States
Tenstorrentuniversity · Austin, Texas, United States; Santa Clara, California, United States