JobLarper
JobLarperCompaniesAWS › Sr. Software Development Engineer AI/ML, Inference Model En…

Sr. Software Development Engineer AI/ML, Inference Model Enablement, AWS Neuron

AWS — The world's largest cloud platform

Cupertino, California, USA Senior+ Posted Jul 29, 2026
JavaC++CloudML

About the role

We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware. As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Trainium accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation. Key job responsibilities * Deliver high-performance models using distributed inference libraries * Drive technical excellence in performance optimization and system reliability across the Neuron ecosystem * Mentor team members and provide technical leadership across multiple work streams * Drive architectural decisions that impact the entire Neuron serving stack * Collaborate with customers, product owners, and engineering teams to define technical strategy * Author technical documentation, design proposals, and architectural guidelines A day in the life You'll lead critical technical initiatives while mentoring team members. You'll collaborate with cross-functional teams of applied scientists, system engineers, and product managers to architect and deliver state-of-the-art inference capabilities. Your day might involve: * Leading design reviews and architectural discussions * Debugging complex performance issues across the stack in collaboration with the compiler and runtime teams * Mentoring junior engineers on system design and model optimization across model enablement teams * Driving technical decisions that shape the future of Neuron's inference stack About the team The inference model enablement team releases its models in the vLLM Neuron plugin: https://github.com/vllm-project/vllm-neuron
- 5+ years of programming using a modern programming language such as Java, C++, or C#, including object-oriented design experience - 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience - 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - 5+ years of non-internship professional software development experience - Experience as a mentor, tech lead or leading an engineering team

What the index says about this role

  • First seen by JobLarper — Aug 2, 2026, 6 days ago. Older postings collect hundreds of applicants — a tailored résumé matters more the longer a role has been live.
  • No pay range in our index for this listing. Across 1,685 indexed Software Engineer roles in US that do publish one, the middle half sits between $185k and $285k, median $225k — JobLarper's read of the market, not a figure from AWS.
  • What Software Engineer roles ask for — across 10,230 indexed openings: REST/APIs (46%), Backend (42%), Cloud (38%), Python (38%), Go (25%). This posting names Cloud, ML.
  • AWS is hiring actively — 219 open roles indexed.

Derived from the 27,000 roles JobLarper indexes daily from official company boards — not from the job description above.

⚡ JobLarper watched this role appear on AWS's official board on Jul 29, 2026. Sign up free to get alerted minutes after roles like this go live, and tailor your real résumé to the exact description — nothing invented.

More open roles at AWS

All 219 open roles at AWS →

Similar Software Engineer roles at other companies