layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
NVIDIA
US, CA, Remote; US, TX, Remote; US, NY, Remote; US, WA, Remote
Source: NVIDIA careers · View original posting
From NVIDIA's posting. “We” and “our” refer to the employer.
NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners.
We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions.
You will be responsible for the architecture, technical plan, and production results of a major platform area like ingestion and orchestration, data quality and reconciliation, or data serving and consumption. You will clarify requirements with customers, define technical objectives, guide design and development among engineers and partner teams, and stay actively engaged in coding, debugging, and production tasks.
Successful candidates have already led complex technical work across team boundaries and delivered improvements that other groups adopted. Our primary implementation environment is Python, SQL, Databricks, and Spark.
BS or MS in Computer Science, Engineering, or a related field (or equivalent experience), and at least 12+ years of equivalent experience
A sustained record of building and operating production software, data platforms, databases, or distributed systems. This includes owning a major component or complex project from requirements and architecture through release and ongoing operation.
Proven ability to outline a component’s technical plan, establish objectives for engineers, assign design and implementation tasks, and guide delivery within your team and nearby teams with little supervision.
Extensive practical experience in one or more of these areas: distributed processing using Spark or a similar system; relational, distributed, or analytical databases; production ETL, change-data capture, streaming, or event handling; or backend and cloud platforms managing large data volumes. You must grasp the interfaces and failure modes of nearby layers thoroughly to inform solid architectural choices.
Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL. Experience designing reusable abstractions, reviewing substantial changes, and personally implementing and debugging critical code paths.
Strong SQL and data-modeling skills, with practical depth in query execution, incremental processing, schema evolution, consistency, and analytical consumption. Ability to reason about idempotency, replay, late-arriving data, partial failure, and correctness across system boundaries.
Experience leading complex investigations involving multiple components and teams. Ability to use logs, metrics, traces, query plans, profiles, and controlled experiments to establish root cause, coordinate resolution, and prevent recurrence.
Demonstrated architectural judgment: evaluating alternatives, anticipating future requirements, and balancing reliability, performance, cost, security, compatibility, and maintainability. Experience leading significant migrations or architectural changes while preserving production service.
Experience establishing production quality and operational practices that other engineers adopt, including testing, CI/CD, monitoring, alerting, rollback, incident response, and secure deployment.
Proven success in influencing technical decisions without official authority, advising engineers outside your immediate project, and advancing workflow improvements across closely related teams. Ability to simplify complex issues, offer a course of action, and communicate decisions and delivery risks clearly.
Proven expertise in building, refining, and running Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, or Unity Catalog workloads along with shared platform features.
Experience designing and operating Kafka or comparable streaming systems, including partitioning, consumer behavior, offset management, backpressure, replay, and schema compatibility.
Experience scaling, migrating, or tuning relational, distributed, time-series, object-storage, or information retrieval systems, including Elasticsearch or OpenSearch.
Background operating compute or GPU clusters, or working with Kubernetes, Slurm, cloud infrastructure, and fleet telemetry across AWS, Azure, GCP, or other providers.
Experience building production agentic systems or agent harnesses, including tool integration, context management, evaluation, permissions, observability, and failure recovery. Evidence of measurable improvements in engineering productivity or operational outcomes.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD.
You will also be eligible for equity and benefits .
Applications for this job will be accepted at least until October 3, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: