We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

Staff Software Engineer, ML Infra, Autonomy

Rivian
$206500.00-$258100.00
sick time, 401(k)
United States, California, Palo Alto
Sep 23, 2026
About Rivian

Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract.

As a company, we constantly challenge what's possible, never simply accepting what has always been done. We reframe old problems, seek new solutions and operate comfortably in areas that are unknown. Our backgrounds are diverse, but our team shares a love of the outdoors and a desire to protect it for future generations.


Role Summary

Rivian Autonomy is building an ML Infrastructure team to give hundreds of ML engineers a training platform they can trust at fleet scale. We are seeking a Staff Software Engineer to help design and build the platform that trains and evaluates our autonomous driving models: the control plane that schedules jobs across accelerator clusters, the storage and I/O layer that feeds them, the observability that tells us where every GPU-hour goes, and the training framework layer that lets ML engineers write model code once and run it on any silicon we operate.

Autonomy at Rivian runs a training fleet of thousands of GPUs on Kubernetes over petabyte-scale sensor data, with additional accelerator types arriving in the coming months. The fleet is fully allocated, so the next step change comes from goodput - how much useful training each GPU-hour delivers. Raising it, through scheduling that keeps large gang jobs fed, a storage and I/O layer that keeps up with the accelerators, and per-workload baselines that make every optimization measurable, is the heart of this role.

The control plane and training fleet already exist and serve every ML engineer in Rivian Autonomy; the scheduling, storage/I/O, observability and framework layers on top of them are largely still to be built. This is a role for someone who wants to set technical direction rather than maintain it.

The platform spans four areas - job scheduling and multi-tenant cluster management, training data storage and I/O, observability and workload optimization, and the training framework layer. You will lead one or two of these areas end-to-end and contribute across the rest; we do not expect one person to be an expert in all four. You will partner closely with the model training teams who are the platform's customers, with the Cloud Infrastructure team that owns the clusters underneath, and with the Data Infrastructure team that produces the datasets the platform serves.

As an early member of the team, you will help define its technical direction, operating model and future hiring.


Responsibilities

Job scheduling and multi-tenant cluster management

  • Evolve our control plane into the single entry point for every accelerator cluster we operate, across cloud providers and silicon types: a user asks for N accelerators of a given type, not for a specific cluster.
  • Design scheduling that maximizes fleet goodput rather than queue order: gang scheduling for jobs of hundreds of nodes, multi-factor priority and fair sharing across teams, quota borrowing with enforceable reclaim, and execution-time-aware backfill.
  • Treat availability as an engineered system: continuous node health checking with automatic cordon, drain and replace, and automatic classification of every failed job (user, out-of-memory, communication, hardware, platform).

Training data storage and I/O

  • Design the storage tiering between object storage, shared or node-local caches and memory, and the request patterns that keep thousands of concurrent readers from overwhelming the object store.
  • Kill the small-file problem for good: shard formats and indexing that serve both fleet-scale shuffled random access and sequential scans, source data stored once and referenced everywhere, and the evaluation and adoption of a training-native storage format.
  • Make "GPU-hours lost to input wait" a first-class metric and drive it down on real jobs.

Observability and workload optimization

  • Build the monitoring and profiling toolchain at every layer a job touches - storage, node, GPU and interconnect, scheduler, and per-job metrics surfaced to the job owner - so that platform and users see the same picture.
  • Establish a workload taxonomy and per-type execution baselines (step time, utilization, communication fraction, input wait, checkpoint cost) that make regressions detectable and every optimization quantifiable.
  • Lay the groundwork for AI-assisted triage and optimization: metrics, logs and scheduler state accessible enough that an agent can be the first responder for failed jobs and propose improvements measured against the baselines.

Training framework and stack currency

  • Help build a thin, opinionated training framework layer over open-source distributed training libraries: one API for users, per-accelerator backends underneath, with golden images, validated launch recipes and a model-zoo CI that gates every release.
  • Run coordinated upgrade programs across the ML stack (distributed compute framework, Kubernetes operators, queueing, experiment tracking) on a cluster that is never idle.

Technical leadership

  • Define the platform roadmap with the team lead - build-versus-buy decisions, boundaries with the cluster, data and model teams, and prioritization by measured impact on goodput and cost.
  • Work directly with ML engineers to find where the platform slows them down or fails, and turn that into improvements to the scheduler, the data path, the tooling and the documentation.
  • Lead architecture across organizational boundaries, communicate recommendations to engineering leadership, and mentor the engineers building and operating the platform.

Qualifications

Required

  • 5+ years of software engineering experience, or equivalent demonstrated impact, with substantial distributed-systems work in production.
  • Staff-level technical leadership: you identify the problems worth solving, shape strategy across teams, make pragmatic trade-offs, and drive ambiguous initiatives from evidence to production.
  • Hands-on experience building or operating large-scale compute or ML infrastructure - and the failure modes that only show up at scale: hung collectives, retry storms, stranded capacity, silent input-bound jobs.
  • Deep expertise in at least one of the following areas, with working familiarity with the others:
  • Cluster scheduling and multi-tenancy - Kubernetes-based scheduling and resource management (Kueue, Volcano, Slurm, YARN or equivalent), including quota, fairness and preemption design.
  • Training data storage and I/O - object-store request behavior, caching tiers, shard formats, shuffle-versus-locality trade-offs, and measuring whether a job is input-bound.
  • Distributed training frameworks and accelerators - PyTorch distributed, Ray Train, JAX or equivalent, on GPU, TPU or Trainium: how process groups, collectives, sharding and checkpointing actually behave.
  • Observability and performance engineering - profiling and monitoring distributed workloads from the kernel and GPU up to the scheduler, and turning measurements into optimizations.
  • Strong programming skills in Python and experience with at least one additional relevant language such as Go, Rust or C++.
  • Experience operating on a major cloud provider (AWS preferred) and on Kubernetes.
  • Strong communication and developer empathy, with a track record of building platforms that engineers adopt and trust; self-directed in ambiguous problem spaces.

Bonus Points

  • Deep understanding of operating systems - memory management, I/O and network stack, scheduling, kernel-level debugging - as applied to debugging and optimizing distributed workloads.
  • Experience with large-model training and cross-GPU communication: collective communication, tensor/pipeline parallelism, NCCL performance at scale.
  • Experience with Ray and KubeRay internals, Kueue or Kubernetes scheduler extensions, or maintaining a patch set on top of an upstream project.
  • Experience with GPU profiling and performance tooling (torch profiler, Nsight Systems and Compute, DCGM, NCCL telemetry) or their counterparts on other accelerators.
  • Familiarity with training-native or columnar data formats (Lance, WebDataset, Parquet row groups) and with GPU-side video decode in the input path.
  • Experience with experiment tracking and model registry platforms (MLflow or equivalent) at scale.
  • Background in applying LLM agents to infrastructure operations, triage or performance tuning.

Pay Disclosure

The salary range for this role is $206,500-$258,100 for San Francisco Bay Area based applicants. This is the lowest to highest salary we in good faith believe we would pay for this role at the time of this posting. An employee's position within the salary range will be based on several factors including, but not limited to, specific competencies, relevant education, qualifications, certifications, experience, skills, geographic location, shift, and organizational needs.

We offer a comprehensive package of benefits for full-time and part-time employees, their spouse or domestic partner, and children up to age 26, including but not limited to paid vacation, paid sick leave, and a competitive portfolio of insurance benefits including life, medical, dental, vision, short-term disability insurance, and long-term disability insurance to eligible employees. You may also have the opportunity to participate in Rivian's 401(k) Plan and Employee Stock Purchase Program if you meet certain eligibility requirements. Full-time employee coverage is effective on their first day of employment. Part-time employee coverage is effective the first of the month following 90 days of employment. More information about benefits is available at rivianbenefits.com.



Equal Opportunity

Rivian is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, sex, sexual orientation, gender, gender expression, gender identity, genetic information or characteristics, physical or mental disability, marital/domestic partner status, age, military/veteran status, medical condition, or any other characteristic protected by law.

Rivian is committed to ensuring that our hiring process is accessible for persons with disabilities. If you have a disability or limitation, such as those covered by the Americans with Disabilities Act, that requires accommodations to assist you in the search and application process, please email us at candidateaccommodations@rivian.com.

Candidate Data Privacy and Technology

Rivian may collect, use and disclose your personal information or personal data (within the meaning of the applicable data protection laws) when you apply for employment and/or participate in our recruitment processes ("Candidate Personal Data"). This data includes contact, demographic, communications, educational, professional, employment, social media/website, network/device, recruiting system usage/interaction, security and preference information. Rivian may use your Candidate Personal Data for the purposes of (i) tracking interactions with our recruiting system; (ii) carrying out, analyzing and improving our application and recruitment process, including assessing you and your application and conducting employment, background and reference checks; (iii) establishing an employment relationship or entering into an employment contract with you; (iv) complying with our legal, regulatory and corporate governance obligations; (v) recordkeeping; (vi) ensuring network and information security and preventing fraud; and (vii) as otherwise required or permitted by applicable law.

Rivian may share your Candidate Personal Data with (i) internal personnel who have a need to know such information in order to perform their duties, including individuals on our People Team, Finance, Legal, and the team(s) with the position(s) for which you are applying; (ii) Rivian affiliates; and (iii) Rivian's service providers, including providers of background checks, staffing services, and cloud services.

Rivian may transfer or store internationally your Candidate Personal Data, including to or in the United States, Canada, the United Kingdom, and the European Union and in the cloud, and this data may be subject to the laws and accessible to the courts, law enforcement and national security authorities of such jurisdictions.

How We Use AI in Our Hiring Process: To ensure transparency, we want candidates to know that Rivian uses iCIMS Talent Cloud Artificial Intelligence (TCAI) and AI-enabled tools to assist with screening, reviewing, organizing and highlighting profiles and applications that match the key requirements for each role.

AI does not make hiring decisions: Qualified candidate applications are reviewed by a member of our team, and all decisions throughout the process are made by humans. We use AI to support efficiency and consistency, not to replace human judgment. We are committed to a fair, thoughtful, and equitable experience for every candidate.

Participation in AI profile matching is entirely voluntary. If you prefer that your profile not be used in this process, you can opt out at any time. Opting out means your profile will be excluded from automated matching and will not be surfaced for additional roles through this system. Your current application remains active and will not be affected in any way.

Please note that we are currently not accepting applications from third party application services.

Applied = 0

(web-9db6c7984-zzklj)