Job Responsibilities
Design, build, and evaluate multi-step agentic AI systems, including autonomous agents capable of planning, tool use, memory management, and multi-agent collaboration.
Research and implement state‑of‑the‑art techniques in agentic architectures, such as ReAct, reflection loops, chain‑of‑thought prompting, and tool‑augmented reasoning.
Develop and maintain agent orchestration frameworks, defining how agents decompose tasks, delegate to sub‑agents, and handle failure and recovery.
Integrate large language models (LLMs) with external tools, APIs, databases, and code execution environments to enable real‑world task completion.
Define and own evaluation frameworks for agentic systems, measuring task success, reliability, latency, cost, and safety across diverse benchmarks and production scenarios.
Collaborate closely with product, engineering, and research teams to translate business requirements into agentic system designs and deliver production‑grade solutions.
Identify and mitigate risks specific to agentic systems, including prompt injection, unintended actions, hallucination in long‑horizon tasks, and unsafe tool use.
Stay current with the rapidly evolving agentic AI landscape, synthesizing academic research and industry developments to inform the team’s technical direction.
Mentor junior ML engineers and scientists, providing technical guidance on agentic design patterns, LLM best practices, and experimentation methodology.
Requirements
8+ years of related industry experience.
Demonstrated experience designing and deploying agentic or multi‑step AI systems (e.g., ReAct, tool‑calling agents, multi‑agent pipelines) in production or research settings.
Strong proficiency in Python and ML frameworks (PyTorch, TensorFlow, or JAX); experience with LLM APIs and orchestration libraries (e.g., LangChain, LlamaIndex, or similar).
Experience integrating LLMs with external tools, APIs, and structured data sources for real‑world task completion.
Solid understanding of prompt engineering techniques including chain‑of‑thought, few‑shot prompting, and structured output generation.
Experience defining and running evaluation frameworks for ML systems, including offline benchmarking and production monitoring.
#J-18808-Ljbffr

Senior Machine Learning Scientist – Agentic Experience
Jobtailor · California, MO, USA ·
- Pay:
- 150.000 - 200.000
- Job type:
- Full Time