layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
AMD
Santa Clara, California
Source: AMD careers · View original posting
From AMD's posting. “We” and “our” refer to the employer.
ADVANCE YOUR CAREER. ADVANCE THE WORLD.
At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.
Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.
We are hiring a AI Research Scientist, Reinforcement Learning (LLM) and Post-Training
, specializing in reinforcement learning to advance post-training and interactive learning for large generative models applied to demanding engineering and hardware-adjacent tasks (code, optimization, tool use, and long-horizon decision making).
You will invent and analyze RL algorithms—policy optimization, preference-based methods, exploration, credit assignment, and reward modeling—run rigorous empirical studies, and partner with infra and product teams to land methods that improve measurable task success without sacrificing stability or safety.
You publish and ship. You are fluent in both RL theory and the practical path from ablation to production-scale training. You care about reward misspecification, variance reduction, and evaluation that reflects real constraints—not only toy environments.
We are hiring a AI Research Scientist, Reinforcement Learning (LLM) and Post-Training
, specializing in reinforcement learning to advance post-training and interactive learning for large generative models applied to demanding engineering and hardware-adjacent tasks (code, optimization, tool use, and long-horizon decision making).
You will invent and analyze RL algorithms—policy optimization, preference-based methods, exploration, credit assignment, and reward modeling—run rigorous empirical studies, and partner with infra and product teams to land methods that improve measurable task success without sacrificing stability or safety.
You publish and ship. You are fluent in both RL theory and the practical path from ablation to production-scale training. You care about reward misspecification, variance reduction, and evaluation that reflects real constraints—not only toy environments.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: