X-RemoteJob Icon X-RemoteJob

Member of Technical Staff, ML Engineer

Physical Superintelligence Competitive / DOE

Member of Technical Staff, ML Engineer

Physical Superintelligence Worldwide Sep 18, 2026
ATS VERIFIED

> ROLE OVERVIEW

Physical Superintelligence is a startup building AI systems to discover new physics at scale. The company has roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute. The mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit. The company is creating the infrastructure to industrialize scientific discovery and usher in a new era of physics breakthroughs.

> CORE RESPONSIBILITIES

  • Own the training and inference infrastructure that Core AI depends on: distributed training jobs, GPU scheduling, and model-serving systems for both proprietary models and self-hosted inference.
  • Build the tools and abstractions AI researchers use to launch training runs, iterate on inference providers, and route workloads across models, so a researcher's time goes into the science instead of the plumbing.
  • Partner with Engineering on the shared platform: capacity planning, observability, and reliability for GPU and inference infrastructure, so training and serving hold up to the same production bar as everything else we ship.
  • Debug and harden the training and inference stack under real load. Egress failures, stalled retries, and routing edge cases are your problem to close, not someone else's ticket.
  • Stay hands-on. You write the code, not just the design doc, and you are the first call when a training job stalls or an inference path breaks.

> HARD REQUIREMENTS & SPECS

  • Three or more years building and operating ML training or inference infrastructure in production, at a company that trains or serves models at meaningful scale.
  • Hands-on experience with distributed training (multi-GPU or multi-node, using PyTorch, Ray, or comparable) and model-serving systems (vLLM, SGLang, Triton, or comparable).
  • Strong software engineering fundamentals. You can build a service that other engineers and researchers depend on every day, not a script that worked once.

Is the AI extraction inaccurate? Report an issue