Senior Software Engineer - NIM Factory Container and Cloud Infrastructure at NVIDIA

Santa Clara, California, United States

NVIDIA Logo
Not SpecifiedCompensation
Senior (5 to 8 years)Experience Level
Full TimeJob Type
UnknownVisa
Technology, Artificial IntelligenceIndustries

Requirements

  • 10+ years building production software with a strong focus on containers and Kubernetes
  • Strong Python skills building production-grade tooling/services
  • Experience with Python SDKs and clients for Kubernetes and cloud services
  • Expert knowledge of Docker/BuildKit, containerd/OCI, image layering, multi-stage builds, and registry workflows
  • Deep experience operating workloads on Kubernetes
  • Strong understanding of LLM inference features, including structured output, KV-cache, and LoRa adapter
  • Hands-on experience building and running GPU workloads in k8s, including NVIDIA device plugin, MIG, CUDA drivers/runtime, and resource isolation
  • Excellent collaboration and communication skills; ability to influence cross-functional design
  • A degree in Computer Science, Computer Engineering, or a related field (BS or MS) or equivalent experience

Responsibilities

  • Design, build, and harden containers for NIM runtimes, inference backends; enable reproducible, multi-arch, CUDA-optimized builds
  • Develop Python tooling and services for build orchestration, CI/CD integrations, Helm/Operator automation, and test harnesses; enforce quality with typing, linting, and unit/integration tests
  • Help design and evolve Kubernetes deployment patterns for NIMs, including GPU scheduling, autoscaling, and multi-cluster rollouts
  • Optimize container performance: layer layout, startup time, build caching, runtime memory/IO, network, and GPU utilization; instrument with metrics and tracing
  • Evolve the base image strategy, dependency management, and artifact/registry topology
  • Collaborate across research, backend, SRE, and product teams to ensure day-0 availability of new models
  • Mentor teammates; set high engineering standards for container quality, security, and operability

Skills

Key technologies and capabilities for this role

PythonKubernetesDockerBuildKitcontainerdOCIHelmOperatorCUDACI/CDGPU

Questions & Answers

Common questions about this position

What compensation and benefits does NVIDIA offer for this role?

NVIDIA offers competitive salaries and a generous benefits package.

Is this position remote or does it require working in an office?

This information is not specified in the job description.

What are the key skills required for this Senior Software Engineer role?

The role requires 10+ years building production software with a strong focus on containers and Kubernetes, strong Python skills for production-grade tooling/services, expert knowledge of Docker/BuildKit and container technologies, deep experience operating workloads on Kubernetes, and hands-on experience with GPU workloads in Kubernetes.

What is the company culture like at NVIDIA for this team?

The role involves collaborating across research, backend, SRE, and product teams, mentoring teammates, and setting high engineering standards for container quality, security, and operability.

What makes a candidate stand out for this position?

Candidates stand out with expertise in Helm chart design systems, Operators, and platform APIs; experience with OpenAI API, Hugging Face API, and inference backends like vLLM, SGLang, TRT-LLM; background in benchmarking and optimizing inference performance; prior experience with multi-tenant or multi-cluster container delivery; and contributions to open-source container, k8s, or GPU ecosystems.

NVIDIA

Designs GPUs and AI computing solutions

About NVIDIA

NVIDIA designs and manufactures graphics processing units (GPUs) and system on a chip units (SoCs) for various markets, including gaming, professional visualization, data centers, and automotive. Their products include GPUs tailored for gaming and professional use, as well as platforms for artificial intelligence (AI) and high-performance computing (HPC) that cater to developers, data scientists, and IT administrators. NVIDIA generates revenue through the sale of hardware, software solutions, and cloud-based services, such as NVIDIA CloudXR and NGC, which enhance experiences in AI, machine learning, and computer vision. What sets NVIDIA apart from competitors is its strong focus on research and development, allowing it to maintain a leadership position in a competitive market. The company's goal is to drive innovation and provide advanced solutions that meet the needs of a diverse clientele, including gamers, researchers, and enterprises.

Santa Clara, CaliforniaHeadquarters
1993Year Founded
$19.5MTotal Funding
IPOCompany Stage
Automotive & Transportation, Enterprise Software, AI & Machine Learning, GamingIndustries
10,001+Employees

Benefits

Company Equity
401(k) Company Match

Risks

Increased competition from AI startups like xAI could challenge NVIDIA's market position.
Serve Robotics' expansion may divert resources from NVIDIA's core GPU and AI businesses.
Integration of VinBrain may pose challenges and distract from NVIDIA's primary operations.

Differentiation

NVIDIA leads in AI and HPC solutions with cutting-edge GPU technology.
The company excels in diverse markets, including gaming, data centers, and autonomous vehicles.
NVIDIA's cloud services, like CloudXR, offer scalable solutions for AI and machine learning.

Upsides

Acquisition of VinBrain enhances NVIDIA's AI capabilities in the healthcare sector.
Investment in Nebius Group boosts NVIDIA's AI infrastructure and cloud platform offerings.
Serve Robotics' expansion, backed by NVIDIA, highlights growth in autonomous delivery services.

Land your dream remote job 3x faster with AI