layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
AMD
Santa Clara, California
Source: AMD careers · View original posting
From AMD's posting. “We” and “our” refer to the employer.
As a core member of the team, you will play a pivotal role in optimizing and developing deep learning frameworks for AMD GPUs. Your expertise will be critical in enhancing GPU kernels, deep learning models, and training/inference performance across multi-GPU and multi-node systems. You will engage with both internal GPU library teams and open-source maintainers to ensure seamless integration of optimizations, utilizing cutting-edge compiler technologies and advanced engineering principles to drive continuous improvement.
Seeking an Industry Leading Expert C++ developer with advanced technical and analytical skills in Linux environments. The ideal candidate will excel in providing technical leadership, guiding teams, and driving projects/initiatives independently. You will define goals, scope, and own development efforts while collaborating effectively within a high-performing team.
Deep expertise in designing and optimizing
GPU kernels for deep learning on AMD GPUs using HIP, CUDA, and assembly (ASM). Strong knowledge of AMD architectures (GCN, RDNA) and low-level programming to maximize performance for AI operations, leveraging tools like Compute Kernel (CK), CUTLASS, and Triton for multi-GPU and multi-platform performance.
Proven ability and experience to integrate GPU-accelerated compute into ML frameworks (e.g., PyTorch
, TensorFlow), with a focus on throughput, scalability, and efficient execution for training and inference workloads.
Software Engineering
Advanced proficiency in Python and C++ with deep experience in performance tuning, debugging, and robust test design, ensuring reliable, maintainable, high-performance codebases.
Broad and indepth experience with large-scale, heterogen eous co mpute environments; adep t at optim izing
AI workloads for performance, efficiency, and reso urce utiliz ation across clusters.
Thorough and detailed understanding of compiler internals, LLVM, and
ROCm
, with the ability to drive system-level optimizations from source to machine code.
Master’s and/ PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
As a core member of the team, you will play a pivotal role in optimizing and developing deep learning frameworks for AMD GPUs. Your expertise will be critical in enhancing GPU kernels, deep learning models, and training/inference performance across multi-GPU and multi-node systems. You will engage with both internal GPU library teams and open-source maintainers to ensure seamless integration of optimizations, utilizing cutting-edge compiler technologies and advanced engineering principles to drive continuous improvement.
Seeking an Industry Leading Expert C++ developer with advanced technical and analytical skills in Linux environments. The ideal candidate will excel in providing technical leadership, guiding teams, and driving projects/initiatives independently. You will define goals, scope, and own development efforts while collaborating effectively within a high-performing team.
Deep expertise in designing and optimizing
GPU kernels for deep learning on AMD GPUs using HIP, CUDA, and assembly (ASM). Strong knowledge of AMD architectures (GCN, RDNA) and low-level programming to maximize performance for AI operations, leveraging tools like Compute Kernel (CK), CUTLASS, and Triton for multi-GPU and multi-platform performance.
Proven ability and experience to integrate GPU-accelerated compute into ML frameworks (e.g., PyTorch
, TensorFlow), with a focus on throughput, scalability, and efficient execution for training and inference workloads.
Software Engineering
Advanced proficiency in Python and C++ with deep experience in performance tuning, debugging, and robust test design, ensuring reliable, maintainable, high-performance codebases.
Broad and indepth experience with large-scale, heterogen eous co mpute environments; adep t at optim izing
AI workloads for performance, efficiency, and reso urce utiliz ation across clusters.
Thorough and detailed understanding of compiler internals, LLVM, and
ROCm
, with the ability to drive system-level optimizations from source to machine code.
Master’s and/ PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: