??% Match

DGX Cloud Performance Engineer

NVIDIA • 2 months ago

Location

Hyderabad, Bengaluru, Pune

Job Type

Full-Time

Experience Level

Executive or Staff-level (10+ Years)

Salary Range

Not disclosed

Job Description

The ideal candidate will have a deep understanding of the methodology to conduct end to end performance analysis of critical AI applications running on large scale parallel and distributed systems. Candidates will work closely with the cross functional teams to define DGX Cloud cluster architecture for different CSPs, optimize workloads running on these systems and develop the methodology that will drive the HW-SW codesign cycle to develop world class AI infrastructure at scale and make them more easily consumable by users (via improved scalability, reliability, cleaner abstractions, etc). What you will be doing: Develop benchmarks, end to end customer applications running at scale, instrumented for performance measurements, tracking, sampling, to measure and optimize performance of important applications and services; Construct carefully designed experiments to analyze, study and develop critical insights into performance bottlenecks, dependencies, from an end to end perspective; Develop ideas on how to improve the end to end system performance and usability by driving changes in the HW or SW (or both). Collaborate with AI researchers, developers, and application service providers to understand internal developer and external customer pain points, requirements, project future needs and share best practice. Develop the necessary modeling framework and the TCO (total cost of ownership) analysis to enable efficient exploration and sweep of the architecture and design space Develop the methodology needed to drive the engineering analysis to Inform the architecture, design and roadmap of DGX Cloud What we need to see: Expertise in working with large scale parallel and distributed accelerator-based system systems Expertise optimizing performance and AI workloads on large scale systems Experience with performance modeling and benchmarking at scale Strong background in Computer Architecture, Networking, Storage systems, Accelerators Familiarity with popular AI frameworks (PyTorch, TensorFlow, JAX, Megatron-LM, Tensort-LLM, VLLM) among others Experience with AI/ML models and workloads, in particular LLMs as well as an understanding of DNNs and their use in emerging AI/ML applications and services Bachelors/Masters in Engineering or equivalent experience (preferably, Electrical Engineering, Computer Engineering, or Computer Science) 10 years experience in the above areas Proficiency in Python, C/C++ Expertise with at least one of public CSP infrastructure (GCP, AWS, Azure, OCI, …); Ways to stand out from the crowd: PhD in the relevant areas Very high intellectual curiosity; Confidence to dig in as needed; Not afraid of confronting complexity; Able to pick up new areas quickly; Proficiency in CUDA, XLA Excellent interpersonal skills

About NVIDIA

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry

Connections

Sai Charan

Senior Developer

5+ years

Kalpana Sharma

Team Lead

3+ years

Rahul Patel

Full Stack Developer

4+ years

Priya Singh

Frontend Developer

2+ years

Connect with professionals in your network

Coming Soon

Job Match Score

??%

Based on your resume

Skill Match Analysis

??% skills matched (?? of 32 skills)

💡 This is keyword matching for reference only. Your actual match score uses AI semantic analysis.

Jobs