Mediabistro logo
job logo

AI Model Optimization Architect for Scalable Inference

Nutanix · San Diego, CA, USA ·

Job type:
Full Time

Qualcomm Technologies, Inc. is seeking a Staff Engineer to lead AI model optimization for LLMs, VLMs, and diffusion models on Qualcomm accelerators. You will transform PyTorch models, drive graph capture with PyTorch and TorchDynamo, and implement kernel fusion using DSLs like Triton.
Collaborate with compiler, performance, and accuracy teams to balance throughput, latency, memory, and quality. You will own transformer optimizations, KVcache behavior, and continuous batching while scaling

#J-18808-Ljbffr