Description

Accelerate AI innovation as a Senior Inference Engineer at Thomson Reuters, focusing on optimizing AI workloads in a hybrid work model. Collaborate with cross-functional teams to enhance capacity forecasting for AI workloads and integrate research models into production.<br/><br/>In this pivotal role within Platform Engineering and Enterprise AI Services, youll be responsible for productionizing, optimizing, and scaling AI and LLM workloads powering Thomson Reuters’ AI-driven products. With 5+ years of experience, you’ll leverage strong ML/LLM fundamentals and inference optimization techniques.<br/><br/>Key Responsibilities:
• Optimize LLM and ML model performance for inference
• Deploy and scale workloads on AWS, Azure, and GCP
• Implement routing strategies for AI traffic
• Develop containerized inference pipelines using Kubernetes
• Collaborate with teams to align with cloud-native patterns<br/><br/>Requirements:
• 5+ years in relevant AI/ML roles
• Skill in GPU programming, CUDA is preferred
• Proficiency in Python and C++ systems language
• Experience with Kubernetes and cloud deployments
• Knowledge of distributed systems and microservices<br/><br/>Utilize your expertise in AI inference to enhance performance and reliability at Thomson Reuters.

J-18808-Ljbffr