layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
Seattle, WA
Source: Amazon careers · View original posting
From Amazon's posting. “We” and “our” refer to the employer.
Senior Software Development Manager, FAIM Evaluations.
Amazon Advertising is building toward a future where an advertiser specifies a few marketing parameters (budget, success definition, which products to promote) and a set of AI agents handles the rest. The Full Funnel Agentic Intelligence and Models (FAIM) organization owns that bet: the Ads Nova agent, the Ads Nova model it reasons with, and the agent infrastructure, learning environments, and evaluations that connect the two. We are looking for a Senior Software Development Manager to found and lead the FAIM Evaluations team.
You will report directly to the Vice President of Full Funnel Agentic Intelligence and Models and own how the entire organization answers one question: is this actually good at advertising?
This is a ground-up build. Evaluation today lives inside individual model and agent teams, measured task by task. You will create the standalone engineering team that turns it into a shared, rigorous system spanning the Ads Nova model, the Ads Nova agent, and the internal agent that serves our Sales, Services, and Operations teams. No inherited harness, no pattern to follow, and a direct line to the VP who sponsors the work.
An advertising benchmark: a representative set of real advertising tasks, organized by domain and difficulty, from single-step questions through multi-step analysis to long-horizon strategic work, each with structured criteria for what a correct end-to-end response looks like
Evaluation infrastructure that scores models and agents deterministically against that benchmark, compares Ads Nova to frontier models on the tasks that matter to advertisers, and gives every science and product team in FAIM the same yardstick
Own the technical vision and roadmap for FAIM evaluations end to end: task taxonomy, rubric design, environment construction, scoring, benchmark versioning, and the separation between what we evaluate on and what we train on Build for the whole org, not one product: your team evaluates the Ads Nova model, the Ads Nova agent, and the internal agent, and you participate in the planning and reviews for all three
Working fluency in how modern models are trained and improved (supervised fine-tuning, reinforcement learning from rubric or verifier signal) and what makes an eval useful as a training asset rather than only a scorecard
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location.
Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, WA, SEATTLE - 220,100.00 - 297,700.00 USD annually
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted:
OCT Consulting, LLC · Kittery, ME, US
Anduril · Quincy, Massachusetts, United States
CHAOS Industries · El Segundo, California, United States
West Monroe · Chicago; New York
eVisit · Mesa, AZ, US
Curaleaf · Phoenix, AZ