X-RemoteJob Icon X-RemoteJob

Member of Technical Staff, Performance & Capacity

Physical Superintelligence Competitive / DOE

Member of Technical Staff, Performance & Capacity

Physical Superintelligence Worldwide Sep 18, 2026
ATS VERIFIED

> ROLE OVERVIEW

We are seeking a Member of Technical Staff, Performance & Capacity to answer two questions honestly: how much compute do we actually need, and how well are we using what we have. Your job is to make both answers measurements rather than estimates, and then be held to them. Then make the same fleet produce more, quarter after quarter. We are a startup with roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale. Our mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit.

> CORE RESPONSIBILITIES

  • Measure the fleet instead of estimating it. Utilization, goodput, and training efficiency, taken on real workloads, so that capacity decisions rest on numbers someone actually observed. Where the platform wastes capacity, you find it and you close it.
  • Own the performance of research workflows end to end. A campaign's time is spent in kernels, collectives, the scheduler's placement decisions, the filesystem, and the queue. When it runs slow, you find which layer, and you fix what lives at the systems level or hand the owning team a diagnosis sharp enough to act on. The unit you optimize is the workflow, not the box.
  • Make the scheduler earn its fleet. The orchestration systems belong to the Distributed Systems team; their efficiency is yours: placement and bin-packing quality, preemption policy, the preemptible fraction as a measured quantity, and checkpointing economics. A workload that can yield on demand is cheaper to run, and worth money in a negotiation. You argue every policy change with a before and an after.
  • Own the decisions. The capacity model that turns a research workload into a defensible node count with stated assumptions and error bars, validated against observed demand. Which GPUs we buy, and what we ask providers for. Which serving engine the inference fleet runs, and which hardware a workload lands on. Yours to make, defended with measurements, and held to when real money is committed.

> HARD REQUIREMENTS & SPECS

  • Five or more years with GPU and large-scale compute workloads, including real multi-node experience: distributed training performance, interconnect and collective-communication behavior, and the patience to find where scale breaks down. Single-box optimization is not this job.
  • You know GPU performance characteristics at the system level: which workloads justify an H100 and which a B200, how to optimize across clusters rather than within one, and where the bottleneck actually sits. Systems view first, the weeds when the numbers demand it.
  • You understand AI training and inference performance deeply. On training: where step time goes, how to optimize across clusters rather than within one, and where the bottleneck actually sits. On inference: where latency goes, how to optimize across clusters rather than within one, and where the bottleneck actually sits.

Is the AI extraction inaccurate? Report an issue