layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
NVIDIA
US, CA, Santa Clara; US, TX, Austin; US, OR, Hillsboro
Source: NVIDIA careers · View original posting
From NVIDIA's posting. “We” and “our” refer to the employer.
We are now looking for a Senior Hardware Architect for our Tegra System-on-Chips (SoC) focused on Reliability, Availability, and Serviceability (RAS). Do you want to be part of the Artificial Intelligence (AI) revolution and help define resilient computing platforms for datacenters, autonomous vehicles, edge systems, and other high-reliability applications?
We are looking for an exceptional SoC architect to help define, drive, and deliver RAS hardware architecture across advanced CPUs and SoCs, from early architectural concepts through design implementation, verification, validation, and production readiness.
This position offers the opportunity to have real impact in a dynamic, technology-focused company developing state-of-the-art processor and system architectures at the forefront of machine learning, autonomous vehicles, high-performance computing, and edge computing.
You will work with world-class systems architects, RAS experts, design teams, verification teams, validation teams, firmware teams, and software partners to define end-to-end hardware RAS features that improve system resiliency, observability, debuggability, error containment, recovery, and serviceability.
Space and radiation-aware design are important areas of interest for this role, including understanding how radiation effects can influence SoC reliability, but the primary focus is broad SoC RAS architecture and driving features successfully through the product development flow.
MS or PhD degree in computer engineering, electrical engineering, or equivalent experience.
At least 8+ years of SoC architecture, design, verification, reliability, silicon validation, or related hardware development experience.
Strong understanding of Reliability, Availability, and Serviceability (RAS) in the SoC context, including fault detection, correction, containment, telemetry, recovery, degradation modes, debug visibility, and serviceability mechanisms.
Experience defining and driving hardware architecture features through the full development lifecycle, including architecture definition, design implementation, verification planning, validation, debug, and production readiness.
Strong understanding of overall SoC architecture and the ability to reason across micro-architecture, full-chip integration, firmware interfaces, software-visible behavior, platform flows, and customer use cases.
Meaningful industry expertise in one or more SoC architecture areas such as RAS, safety, debug, clocks, resets, interconnects, memory controllers, IO technologies, platform integration, firmware-visible error handling, or diagnostic infrastructure.
Hands-on experience with design verification, silicon validation, fault injection, coverage analysis, resiliency modeling, diagnostic development, or reliability validation methodology.
Familiarity with radiation effects, space operation, or other high-reliability deployment environments is strongly valued, including understanding how hardware architecture can mitigate single-event effects and related reliability risks.
Excellent analytical, written, and verbal interpersonal skills with the ability to work effectively across architecture, design, verification, firmware, software, validation, and customer-facing teams.
Demonstrated history of architecting and delivering complex SoC RAS features across design, verification, validation, and production phases.
Deep familiarity with architectural resiliency techniques such as ECC, parity, redundancy, replay, checkpoint/restart, scrubbing, isolation, containment, telemetry, error logging, recovery flows, and graceful degradation.
Experience with cross-functional debug of hardware failures in simulation, emulation, post-silicon validation, production, customer deployments, or other high-reliability systems.
Familiarity with Design for Debug, Design for Test, Design for Reliability, silicon observability, fault-injection methodology, and coverage-driven validation flows.
Experience with radiation effects analysis, soft-error-rate analysis, radiation test campaigns, accelerated stress testing, heavy-ion or proton testing, or space qualification methodology.
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you are creative, autonomous, and motivated to define and deliver resilient SoC architectures that improve reliability, debuggability, and serviceability across demanding applications, we want to hear from you!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.
You will also be eligible for equity and benefits .
Applications for this job will be accepted at least until September 6, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: