NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
We are seeking a highly-skilled Senior On-Device Model Inference Optimization Engineer to join our team and lead efforts in improving the performance and efficiency of AI models enabling the next generation of autonomous vehicles technology at NVIDIA!
What you'll be doing:
Develop and implement strategies to optimize AI model inference for on-device deployment.
Employ techniques like pruning, quantization, and knowledge distillation to minimize model size and computational demands.
Optimize performance-critical components using CUDA and C++.
Collaborate with multi-functional teams to align optimization efforts with hardware capabilities and deployment needs.
Benchmark inference performance, identify bottlenecks, and implement solutions.
Research and apply innovative methods for inference optimization.
Adapt models for diverse hardware platforms and operating systems with varying capabilities.
Create tools to validate the accuracy and latency of deployed models at scale with minimal friction.
Recommend and implement model architecture changes to improve the accuracy-latency balance.
What we need to see:
MSc or PhD in Computer Science, Engineering, or a related field, or equivalent experience.
Over 5 years of confirmed experience specializing in model inference and optimization.
10+ overall years of work experience in a relevant area
Expertise in modern machine learning frameworks, particularly PyTorch, ONNX, and TensorRT.
Proven experience in optimizing inference for transformer and convolutional architectures.
Strong programming proficiency in CUDA, Python, and C++.
In-depth knowledge of optimization techniques, including quantization, pruning, distillation, and hardware-aware neural architecture search.
Skilled in building and deploying scalable, cloud-based inference systems.
Passionate about developing efficient, production-ready solutions with a strong focus on code quality and performance.
Meticulous attention to detail, ensuring precision and reliability in safety-critical systems.
Strong collaboration and communication skills for working optimally across multidisciplinary teams.
A proactive, diligent mentality with a drive to tackle complex optimization challenges.
Ways to stand out from the crowd:
Publications or industry experience in optimizing and deploying model inference at scale.
Hands-on expertise in hardware-aware optimizations and accelerators such as GPUs, TPUs, or custom ASICs.
Active contributions to open-source projects focused on inference optimization or machine learning frameworks.
Experience in designing and deploying inference pipelines for real-time or autonomous systems.
You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.
If an employer mentions a salary or salary range on their job, we display it as an "Employer Estimate". If a job has no salary data, Rise displays an estimate if available.
Experienced Senior Account Manager needed to lead NVIDIA’s growth and technology adoption efforts within the USAF and USSF, leveraging deep defense sector expertise.
Lead NVIDIA’s Data Center Quality efforts as a Senior Program Manager, driving customer satisfaction and operational excellence across enterprise partners.
Experienced Staff Software Engineer needed to advance cloud-native applications and database management at NBCUniversal in a fully remote role.
An exciting opportunity to contribute as a .NET Developer at TRINETIX, driving advanced software architecture solutions in a dynamic, global tech company.
Drive the development of highly available Elixir-based distributed systems for AI infrastructure at a cutting-edge startup.
Contribute to innovative research software development at Emory University as a Systems Software Engineer specialized in medical imaging and data integration.
IMO Health is looking for a Software Engineer specializing in AI Applications to build scalable, intelligent solutions advancing healthcare innovation.
Contribute as a Senior Software Engineer at NVIDIA, advancing cloud-native technologies and container orchestration for GPU and DPU accelerated computing.
Contribute to ResMed’s digital health platform as a Senior Software Engineer specializing in API and data model design, driving innovation in a collaborative global environment.
Peraton is hiring a skilled Full Stack Software Developer to build cutting-edge software tools for space domain analysis and asset protection.
Innovative IT firm IGS is hiring a remote AI Developer to build, deploy, and integrate data-driven AI solutions for federal government clients.
Contribute to blockchain mass adoption as a Wallet Engineer at Backpack, building and optimizing wallet features with a hybrid work setup.
Palo Alto Networks is looking for a Principal Software Engineer in Test to advance cutting-edge cloud-based cybersecurity products through automated testing and innovative problem-solving.
Innovate front-end experiences at Swoop by building interactive 2D and 3D visualizations for a cutting-edge infrastructure platform in a hybrid Minneapolis-based role.
Contribute to cutting-edge autonomous vehicle technology by designing and optimizing firmware and drivers for ML accelerator chips at Waymo.
NVIDIA is a publicly traded, multinational technology company headquartered in Santa Clara, California. NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, and ignited the era of modern AI.
493 jobsSubscribe to Rise newsletter