layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
Bellevue, WA
Source: Amazon careers · View original posting
From Amazon's posting. “We” and “our” refer to the employer.
Amazon Devices (Lab126) builds products and services that delight millions of customers globally. The Edge AI ML Platform and Infrastructure team is building the platform that enables Amazon teams to train, optimize, evaluate, and deploy generative AI models on devices and in the cloud.
Today, optimizing a large model for a new hardware target requires experts to connect model onboarding, distributed training, compression, evaluation, compilation, and deployment systems by hand. We are turning that work into a repeatable, self-service workflow. Our platform supports large language, vision, audio, multimodal, and mixture-of-experts models, and gives scientists and engineers the tools to move new optimization techniques from research code into reliable production workflows.
We are looking for a Software Development Manager to build and lead the ML infrastructure team behind this platform. You will own distributed training on multi-node GPU clusters, compute capacity and utilization, CI/CD, observability, and operational reliability for GPU-intensive workloads.
You will hire and develop a team of software and ML infrastructure engineers, set its technical direction and roadmap, and deliver platform capabilities that scientists and product teams depend on to ship models with hundreds of billions of parameters.
This role combines people leadership with deep technical judgment. You will grow engineers and managers-in-the-making, drive architecture decisions with your senior engineers, turn ambiguous science and product needs into a prioritized plan, and hold a high bar for delivery and operational excellence.
Key job responsibilities
Build, lead, and grow a team of software and ML infrastructure engineers: recruit and hire, set clear goals, coach for growth, and manage performance across the team.
Own the roadmap for ML infrastructure—distributed training, GPU capacity, workflow orchestration, CI/CD, and observability—balancing near-term deliveries with long-term platform health.
Drive the architecture of distributed training capabilities (data, tensor, pipeline, and model parallelism) for large language and multimodal models, partnering with senior engineers and applied scientists.
Establish operational excellence for production platform services, including metrics, alarms, runbooks, on-call processes, and root-cause correction of recurring issues, while owning GPU fleet efficiency, capacity planning, and cost optimization.
Partner with applied science, compiler, runtime, hardware, security, and product teams to align requirements, manage dependencies, and deliver cross-team programs.
The Edge AI ML Platform and Infrastructure team brings together software engineers, ML infrastructure engineers, and GPU performance specialists. We build reusable model training, optimization, and deployment capabilities for Amazon product teams, working closely with applied scientists across Edge AI. Our customers need to adapt rapidly changing model architectures to constrained hardware and production workloads without rebuilding the toolchain for every model.
The team owns the platform foundations that connect model development to deployment. Because our scope runs end to end, we can improve training, compression, evaluation, and deployment as one system. We value clear interfaces, measurable performance, automated quality gates, and direct collaboration between science and engineering.
Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
Experience partnering with product or program management teams
Experience managing a team of high calibre Software Engineers developing complex, world class, scalable software systems that have been successfully delivered to customers
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location.
Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, WA, Bellevue - 184,900.00 - 250,200.00 USD annually
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: