Qualcomm Technologies, Inc. is seeking a Staff Engineer to lead AI model optimization for LLMs, VLMs, and diffusion models on Qualcomm accelerators. You will transform PyTorch models, drive graph capture with PyTorch and TorchDynamo, and implement kernel fusion using DSLs like Triton.
Collaborate with compiler, performance, and accuracy teams to balance throughput, latency, memory, and quality. You will own transformer optimizations, KVcache behavior, and continuous batching while scaling
#J-18808-Ljbffr

AI Model Optimization Architect for Scalable Inference
Nutanix · San Diego, CA, USA ·
- Job type:
- Full Time