CoreWeave

Operations Engineer, Fleet Reliability

Livingston, New Jersey, United States

Not SpecifiedCompensation
Mid-level (3 to 4 years), Senior (5 to 8 years)Experience Level
Full TimeJob Type
UnknownVisa
AI Hyperscaler, Cloud Computing, Accelerated Computing, Data CentersIndustries

Position Overview

  • Location Type: Remote
  • Employment Type: Full-time
  • Salary: Not specified

CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting-edge services powering the next wave of AI. Our technology provides enterprises and leading AI labs with the most performant, efficient, and resilient solutions for accelerated computing. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe. CoreWeave was ranked as one of the TIME100 most influential companies of 2024. CoreWeave powers the creation and delivery of the intelligence that drives innovation.

What You'll Do

The Fleet Reliability Operations team is responsible for the day-to-day provisioning, management, and uptime of CoreWeave’s ever-expanding fleet of server nodes. Playing a central role in CoreWeave’s growth strategy, this team is on the front line for configuration, updates, and remote troubleshooting of our highest tier of supercomputing clusters and their networking, delivery platforms, and tools dependencies. You will be in a daily battle with the forces of entropy to maximize the number of nodes CoreWeave can deliver to customers.

We are seeking curious, creative, and persistent problem solvers to join our Fleet Reliability Operations team to help us drive batches of server nodes through our provisioning and validation processes while efficiently and effectively troubleshooting node or cluster problems as they arise. This individual will join a team of committed engineers working to deploy nodes as fast as they can be racked and turned on.

  • Configure and maintain large-scale high-performance supercomputing clusters running state-of-the-art GPUs
  • Troubleshoot hardware and software issues; escalate and coordinate as needed with data center, network, hardware, and platform teams to drive resolution
  • Monitor and analyze system performance and take appropriate remediation actions for cloud health
  • Approach your work with flexibility and optimism anticipating shifting business and technical priorities
  • Create and maintain documentation of team processes, knowledge, and best practices for system management
  • Think critically about your day-to-day work and work collaboratively to improve team processes and efficiency
  • Participate in oncall rotations which include after hours and weekend work

Who You Are

Minimum Qualifications:

  • Strong understanding of Linux system administration and internals
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently

Skills

Fleet Reliability
Operations
Provisioning
Server Management
Uptime
Configuration
Updates
Remote Troubleshooting
Supercomputing Clusters
Networking
Hardware Troubleshooting
Software Troubleshooting
System Performance Monitoring
GPU

CoreWeave

Cloud service for GPU-accelerated workloads

About CoreWeave

CoreWeave provides cloud computing services that focus on GPU-accelerated workloads, which are essential for tasks requiring high computational power. Their services cater to industries such as artificial intelligence, machine learning, visual effects rendering, and data processing. Clients can access powerful computing resources on a pay-as-you-go basis, allowing them to avoid the costs of purchasing expensive hardware. CoreWeave's infrastructure utilizes a bare metal serverless Kubernetes platform, which enhances performance while minimizing operational complexity for users. This setup enables clients to optimize their computing needs with a variety of NVIDIA GPUs, ensuring they can balance performance and cost effectively. The company's goal is to offer flexible and scalable computing solutions that meet the demands of diverse clients, from tech companies to film studios.

New York City, New YorkHeadquarters
2017Year Founded
$1,625.4MTotal Funding
SECONDARYCompany Stage
Enterprise Software, AI & Machine LearningIndustries
501-1,000Employees

Benefits

Health Insurance
Dental Insurance
Vision Insurance
Life Insurance
Disability Insurance
Health Savings Account/Flexible Spending Account
Tuition Reimbursement
Mental Health Support
Family Planning Benefits
Paid Parental Leave
Hybrid Work Options
401(k) Company Match
Unlimited Paid Time Off
Catered lunch each day in our office and data center locations
A casual work environment

Risks

Increased competition from Vultr, backed by AMD, may impact CoreWeave's market share.
The planned IPO in 2025 could expose CoreWeave to market volatility and scrutiny.
Reliance on NVIDIA GPUs poses risks if supply chain issues arise.

Differentiation

CoreWeave specializes in GPU-accelerated workloads, catering to AI and high-performance computing.
Their infrastructure uses a bare metal serverless Kubernetes platform for high performance.
CoreWeave offers a flexible pay-as-you-go pricing model, appealing to cost-conscious clients.

Upsides

CoreWeave's partnership with Dell enhances infrastructure, attracting clients seeking advanced AI solutions.
The $600M data center funding in Virginia supports CoreWeave's expansion in the AI cloud market.
Experienced leadership, like CIO Sandy Venugopal, positions CoreWeave for strategic growth.

Land your dream remote job 3x faster with AI