Description
ABOUT THE ROLE
In this role, you will lead AI engagements end-to-end for enterprise clients – from discovery and architecture design to production deployment on NVIDIA’s ecosystem. You’ll work closely with NVIDIA stakeholders and internal engineering teams to shape GenAI and Agentic AI solutions, drive pre-sales, and turn business challenges into production-grade AI architectures.
RESPONSIBILITIES
- Lead AI engagements end-to-end, from discovery through architecture design and production implementation, to deliver clearly scoped client outcomes
- Design and validate production-grade AI architectures using NVIDIA NeMo, NIM, Triton, Riva, DeepStream, Metropolis, and Omniverse across cloud and on-premise environments
- Translate business challenges into AI use-case definitions and reference architectures, collaborating with clients, NVIDIA stakeholders, and engineering teams
- Define reference architectures for GenAI and Agentic AI solutions, advising on GPU-accelerated infrastructure, Kubernetes orchestration, and MIG/vGPU partitioning
- Benchmark LLM serving stacks such as vLLM, TGI, and Triton+TensorRT-LLM, measuring throughput, TTFT, TPOT, and tokens/sec/GPU
- Tune deployment configurations, including batch size, concurrency, parallelism, and quantization, based on profiling data to hit latency, throughput, and cost targets
- Lead pre-sales activities, including proposals, discovery sessions, and workshops, to drive GenAI proof-of-concept initiatives with measurable value
- Contribute to the NVIDIA alliance go-to-market strategy and mentor teams on NVIDIA ecosystem technologies
REQUIREMENTS
- 6+ years in AI consulting, GenAI/Agentic AI development, or Deep Learning, leading client-facing engagements end-to-end
- Hands-on expertise in Generative AI, Agentic AI, multimodal AI, and LLMs/VLMs across the full lifecycle, from experimentation to deployment
- Practical experience with at least three NVIDIA platforms (NeMo, NIM, Riva, Metropolis, Omniverse, Triton, DeepStream, or TensorRT-LLM)
- Solid command of Python and AI/ML frameworks: PyTorch, TensorFlow, Pandas, NumPy, and Hugging Face
- Experience deploying AI/ML workloads on Kubernetes (Helm, Operators, GPU Operator, MIG/vGPU) and self-hosted stacks like vLLM or Ollama
- Working knowledge of quantization and inference optimization, using GPU profiling tools such as Nsight Systems and DCGM metrics
- Experience designing AI solutions on at least one major cloud (AWS, Azure, or GCP), with a strong grasp of TensorRT, Triton, CUDA, and ONNX
- Confident stakeholder management skills, presenting to executives, and leading pre-sales engagements.
This is the anticipated salary range that SoftServe expects to offer for this position in Canada. The range reflects base salary only. Final compensation will depend on several factors, including relevant experience, skills, education, and qualifications.
The anticipated salary range for this role is 15,958 – 17,554 CAD monthly, in line with our internal compensation framework. Most candidates are offered a salary within this range. This range applies only to regular employment contracts.
Any variable or discretionary compensation components – such as participation in the Variable Pay Plan (VPP) – are considered separately. VPP eligibility depends on the role and SoftServe policy, and if applicable, this information will be communicated to you before the offer is confirmed.
SoftServe is an equal opportunity employer. Qualified applicants will receive consideration regardless of race, color, ancestry, ethnicity, national origin, religion, sex, sexual orientation, gender identity or expression, age, citizenship, disability, health condition, marital or family status, veteran status, or any other characteristic protected by applicable law.





