Together AI is recruiting a highly skilled Inference Frameworks and Optimization Engineer to design and optimize distributed inference engines for multimodal models at scale. You will focus on low-latency, high-throughput inference, GPU/accelerator optimization, and software-hardware co-design to enable efficient deployment of LLMs and vision models.
Join a research-driven team shaping AI infrastructure, collaborating with researchers to craft end-to-end serving pipelines and push the boundaries
#J-18808-Ljbffr

LLM Inference Frameworks Engineer High-Performance Serving
Together AI · San Francisco, CA, USA ·
- Job type:
- Full Time