Amazon Web Services (AWS)
SDE II, ML Infra Services, Annapurna Labs
Full Time · In Office · Seattle, Washington (USA)
Posted Sep 19, 2026
Description
Annapurna Labs was a startup company acquired by AWS in 2015, and is now fully integrated. If AWS is an infrastructure company, then think Annapurna Labs as the infrastructure provider of AWS. Our org covers multiple disciplines including silicon engineering, hardware design and verification, software, and operations. AWS Nitro, ENA, EFA, Graviton and F1 EC2 Instances, AWS Neuron, Inferentia and Trainium ML Accelerators, and in storage with scalable NVMe, are some of the products we have delivered, over the last few years.
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trn1 and Inf1 servers that use them. This position is for a Software Engineer that will lead the development of machine learning tools to run, optimize, and analyze machine learning workloads. This candidate must have had experience leading machine learning tool projects, preferably starting from architecture through several generations of delivery to customers. Deep knowledge of profiling and optimization, resource management, scheduling, code generation are needed. The ideal candidate will have worked on new instruction set architectures, which may include CPU, NPU, GPU and other forms of compute.
Key job responsibilities
This engineer will lead the design and implementation of ML infrastructure platform, building systems for capacity management, workload scheduling, and fleet orchestration across ML accelerators. They will work with ML scientists, training infrastructure engineers, hardware teams, and internal customers to ensure the ML Infra service delivers seamless ML Accelerator access with low wait times, high utilization, and zero-config deployment from various environments.
A day in the life
As You Design And Code Solutions To Help Our Team Drive Efficiencies In Software Architecture, You’ll Create Metrics, Implement Automation And Other Improvements, And Resolve The Root Cause Of Software Defects. You’ll Also
Build high-impact solutions to deliver to our large customer base.
Participate in design discussions, code review, and communicate with internal and external stakeholders.
Work cross-functionally to help drive business decisions with your technical input.
Work in a startup-like development environment, where you’re always working on the most important stuff.
About The Team
- High-impact, high-visibility: You'll directly accelerate every Neuron team's ability to ship — your work multiplies the output of 100+ engineers
- Greenfield opportunities: We're actively building new capabilities with significant design ownership for SDEs
- Small, senior team: where every person owns major components and drives architectural decisions
- AI infrastructure: Work at the intersection of Kubernetes, custom silicon, and large-scale ML workloads
Diverse Experiences
We value diverse experiences and non-traditional career paths. If your career is just starting or includes alternative experiences, we encourage you to apply.
Inclusive Team Culture
Our employee-led affinity groups foster inclusion. Events like CORE and AmazeCon inspire us to embrace our uniqueness.
Work/Life Balance
We strive for flexibility as part of our working culture, supporting you both at work and at home.
Mentorship & Career Growth
We offer knowledge-sharing, mentorship, and one-on-one code reviews to help you grow as a professional.
Basic Qualifications
- 3+ years of non-internship professional software development experience
- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Experience programming with at least one software programming language
Preferred Qualifications
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Bachelor's degree in computer science or equivalent
- 2+ years of building large-scale machine-learning infrastructure for online recommendation, ads ranking, personalization or search experience
- Strong proficiency in Go/Java, Python and working knowledge Javascript/TypeScript
- Experience building and operating large-scale distributed systems on Kubernetes
- Experience designing, deploying, and maintaining production services at scale, including on-call ownership
- Experience with machine learning infrastructure — orchestration, scheduling, or resource management at scale
- Proficiency in application and kernel-level performance profiling and optimization
- Experience with integrated software/hardware performance analysis in heterogeneous compute environments
- Proficiency in observability and telemetry — instrumentation, metrics collection, alarming, dashboarding, and monitoring
- Experience debugging complex issues in large-scale distributed systems and driving best practices
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, WA, Seattle - 143,700.00 - 194,400.00 USD annually
Company - Amazon.com Services LLC
Job ID: A10554141
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
SDE II, ML Infra Services, Annapurna Labs
Amazon Web Services (AWS)
Similar Other Jobs
View all Other jobs→BOOST LLC
Consultant - Workforce Analytics & Human Capital
TalentHop
Program Manager, Ting Data Products
JPMorganChase
Data Product Studio Technical Product Manager
VySystems
Data Product Owner
MediSystem Pharmacy
Pharmacy Data Specialist
Robert Half
Senior Manager, Insights & Analytics
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.