layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
AMD
San Jose, California
Source: AMD careers · View original posting
From AMD's posting. “We” and “our” refer to the employer.
ADVANCE YOUR CAREER. ADVANCE THE WORLD.
At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.
AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.
We are looking for a strong, Principal or Fellow level software engineer to join our AI Infrastructure team. You will work on improving the performance, efficiency, and reliability of AI workloads across both model training and inference.
Our team supports a broad range of machine learning systems, including large language models, diffusion models, and recommendation models. You will collaborate closely with customers and internal engineering teams to understand performance bottlenecks, optimize workloads, and ensure that models run reliably at scale.
This role is a strong fit for an engineer who enjoys working across the AI software and hardware stack, solving technically challenging performance problems, and partnering directly with customers to make them successful.
You will help customers achieve meaningful improvements in model performance and system reliability. You will identify difficult bottlenecks, develop reusable solutions, and help shape the infrastructure and product capabilities needed to run demanding AI workloads efficiently at scale.
Profile and optimize AI model training and inference workloads.
Improve model throughput, latency, memory efficiency, scalability, and reliability.
Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
Optimize workloads involving large language models, diffusion models, recommendation systems, and other modern machine learning architectures.
Develop performance tooling, benchmarks, automation, and observability systems.
Investigate and resolve complex production issues affecting AI workloads.
Collaborate with customers to understand their technical requirements, reproduce issues, and recommend effective solutions.
Translate customer feedback into product and infrastructure improvements.
Work closely with machine learning engineers, systems engineers, hardware teams, and product teams.
Document performance findings, technical recommendations, and best practices.
We are looking for a strong, Principal or Fellow level software engineer to join our AI Infrastructure team. You will work on improving the performance, efficiency, and reliability of AI workloads across both model training and inference.
Our team supports a broad range of machine learning systems, including large language models, diffusion models, and recommendation models. You will collaborate closely with customers and internal engineering teams to understand performance bottlenecks, optimize workloads, and ensure that models run reliably at scale.
This role is a strong fit for an engineer who enjoys working across the AI software and hardware stack, solving technically challenging performance problems, and partnering directly with customers to make them successful.
You will help customers achieve meaningful improvements in model performance and system reliability. You will identify difficult bottlenecks, develop reusable solutions, and help shape the infrastructure and product capabilities needed to run demanding AI workloads efficiently at scale.
Profile and optimize AI model training and inference workloads.
Improve model throughput, latency, memory efficiency, scalability, and reliability.
Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
Optimize workloads involving large language models, diffusion models, recommendation systems, and other modern machine learning architectures.
Develop performance tooling, benchmarks, automation, and observability systems.
Investigate and resolve complex production issues affecting AI workloads.
Collaborate with customers to understand their technical requirements, reproduce issues, and recommend effective solutions.
Translate customer feedback into product and infrastructure improvements.
Work closely with machine learning engineers, systems engineers, hardware teams, and product teams.
Document performance findings, technical recommendations, and best practices.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: