Mediabistro logo
job logo

LLM Inference Frameworks Engineer High-Performance Serving

Together AI · San Francisco, CA, USA ·

Job type:
Full Time

Together AI is recruiting a highly skilled Inference Frameworks and Optimization Engineer to design and optimize distributed inference engines for multimodal models at scale. You will focus on low-latency, high-throughput inference, GPU/accelerator optimization, and software-hardware co-design to enable efficient deployment of LLMs and vision models.
Join a research-driven team shaping AI infrastructure, collaborating with researchers to craft end-to-end serving pipelines and push the boundaries

#J-18808-Ljbffr