Prime Video is looking for a Research Engineer to support scientific research efforts in our Content Reasoning, Enrichment and Localization team. We build and deploy Generative AI models across Prime Video Lines of Business, including Derivative Content, Trust & Safety, and Metadata Localization. A Research Engineer supports Applied Scientists on the team by building experiment frameworks, infrastructure and deployment packages. The role involves directly facilitating science benchmarking, model training, model deployment, and deep dives on any deep learning related technical problem. The engineer may lead or facilitate a joint project between science and engineering teams. Responsibilities include end‑to‑end ownership of model training and deployment, from managing our GPU cluster to monitoring job queues and containerizing trained models for production. The role supports several research teams across lines of business and machine learning methodologies.
Key job responsibilities
Cluster Management
Architect and operate SLURM cluster infrastructure for distributed training and evaluation workloads
Maintain healthy job status and manage queues from different users and teams
Monitor feedback regarding cluster features and functions
Identify and drive infrastructure optimization opportunities
Develop self‑service tooling, automation, and APIs
Model Deployment
Gather requirements from collaborators on research teams
Containerize Machine Learning models
Shepherd models through deployment into beta environments, including debugging builds
Develop data collection, aggregation and monitoring pipelines
Build observability and monitoring systems or dashboards for ML infrastructure
About The Team
This team's mission is to deeply understand all content and empower customers with relevant language options, innovative accessibility assists, and rich title‑information across all their content experiences on Prime Video. We delight our customers by pushing the boundaries of content understanding and enrichment. Through inclusion and innovation, we do the most fulfilling work of our career.
Basic Qualifications
Experience in automating, deploying, and supporting large‑scale infrastructure
Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
Experience with Linux/Unix
Bachelor's degree, or 2+ years of working with Advanced Compute technologies including: Accelerated Compute, High Performance Compute, Visual/Spatial Compute, and/or IoT
3+ years of cloud computing technologies experience
Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience leading the architecture and design of new and current systems
Preferred Qualifications
Experience with distributed systems at scale
Master's degree
Experience with large‑scale machine learning systems such as profiling and debugging and understanding of system performance and scalability
Publications or contributions to open‑source HPC projects
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information.
#J-18808-Ljbffr

Research Engineer, Prime Video - Content Understanding
Prime Video & Amazon MGM Studios · Seattle, WA, USA ·
- Pay:
- 100.000 - 130.000
- Job type:
- Full Time