Software Development Engineer I – AI/ML Network Infrastructure, Annapurna Labs
Amazon (AI roles) · Cupertino, California, USA
No salary listed. Estimated from 1122 salary-disclosed Engineering roles on this board: $201K–$307K (interquartile range; estimate, not the employer's figure).
Apply at Amazon (AI roles)Find your warm intro on LinkedInAbout the role
We're looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI models to train across massive GPU clusters, developing support for communication libraries and frameworks like NCCL, NVSHMEM, and NIXL.
This is a ground-floor opportunity to work at the intersection of high-performance computing, networking, and machine learning infrastructure - building the systems that power the largest AI workloads in the cloud.
Key job responsibilities
- Write high-performance C/C++ code for network communication libraries running on custom AWS hardware
- Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
- Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
- Design mechanisms to detect functional and performance regressions before they reach production
- Work across many instance types, software stacks, and Linux environments - Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field (recent graduates welcome)
- Strong proficiency in C/C++
- Solid coursework or project experience in: 1/ Operating Systems (Linux internals, kernel concepts, memory management) 2/ Parallel Computer Architecture (multi-threading, SIMD, GPU programming, cache coherence) 3/ Distributed Systems (consensus, message passing, fault tolerance, scalability)
- Familiarity with Linux development environments and toolchains
This is a ground-floor opportunity to work at the intersection of high-performance computing, networking, and machine learning infrastructure - building the systems that power the largest AI workloads in the cloud.
Key job responsibilities
- Write high-performance C/C++ code for network communication libraries running on custom AWS hardware
- Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
- Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
- Design mechanisms to detect functional and performance regressions before they reach production
- Work across many instance types, software stacks, and Linux environments - Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field (recent graduates welcome)
- Strong proficiency in C/C++
- Solid coursework or project experience in: 1/ Operating Systems (Linux internals, kernel concepts, memory management) 2/ Parallel Computer Architecture (multi-threading, SIMD, GPU programming, cache coherence) 3/ Distributed Systems (consensus, message passing, fault tolerance, scalability)
- Familiarity with Linux development environments and toolchains
More like this