Member of Technical Staff - ML Infrastructure Engineer
We're hiring an ML Infrastructure Engineer to own the infrastructure and tooling that our research team runs on every day. This is a hybrid DevOps / Developer Experience role: you'll stand up and operate large GPU clusters, build data pipelines that move massive medical datasets through training jobs efficiently, and create the tooling that lets our researchers move fast without thinking about plumbingML Infrastructure Engineer
About Voio
Voio is a healthcare AI company building frontier foundation models for medicine. We're based in San Francisco and our research team trains on extremely large multimodal medical datasets. The infrastructure this role owns goes directly into models aimed at improving clinical decision-making at scale.
The Role
We're hiring an ML Infrastructure Engineer to own the infrastructure and tooling that our research team runs on every day. This is a hybrid DevOps / Developer Experience role: you'll stand up and operate large GPU clusters, build data pipelines that move massive medical datasets through training jobs efficiently, and create the tooling that lets our researchers move fast without thinking about plumbing.
If you've ever sat next to an ML researcher, watched them lose half a day to a flaky GPU node or a misversioned dataset, and thought "I could fix this for everyone" — this role is for you.What you'll work on
- GPU cluster operations. Run and scale our GPU cluster across cloud providers. Job orchestration with SkyPilot, multi-node training reliability, capacity planning, and cost optimization. You'll be the person who makes a 64-GPU run "just work."
- Data infrastructure at scale. Manage petabyte-scale medical imaging and clinical data, such as DICOM, NIfTI, parquet, raw EHR. Build and maintain the data pipeline tiers (object store → NFS → NVMe scratch) so GPUs are never starved.
- Experiment tracking and developer experience. Own our W&B integration. Build the abstractions, CLIs, and templates that turn a complex distributed training job into a one-line submission. Make it easy for researchers to launch, track, compare, and resume experiments.
- Observability and on-call. Grafana dashboards, structured alert tiers routed through Slack and PagerDuty. Catch problems before researchers do.
- Reliability culture. Postmortems, runbooks, sane defaults. The bar is that no researcher is ever blocked on infrastructure for more than a few minutes.
What we're looking for
- Strong production experience operating GPU infrastructure at scale — multi-node distributed training, high-speed networking (InfiniBand or equivalent), storage tiering, and the kind of debugging that happens at 11pm when a long-running job dies on epoch 3.
- Comfort across the stack: Linux, Python, Kubernetes or equivalent orchestration, at least one major cloud (AWS / GCP / Azure), CI/CD.
- A genuine developer-experience instinct. You treat researchers as your primary users, and you measure success by how invisible the infrastructure becomes.
- Solid software engineering fundamentals — you can build tooling that other engineers rely on, with tests, docs, and a sane API.
Bonus points
- Production experience with SkyPilot, lakeFS, or W&B.
- Background in healthcare or biomedical data (DICOM, NIfTI, FHIR, HIPAA-aware infrastructure).
- Experience designing data versioning systems for mixed-modality datasets at scale.
- Open-source contributions to ML infrastructure projects.
SkyPilot · lakeFS · W&B · S3-compatible object stores · Grafana · Slack · PagerDuty · Python · Linux
Logistics
- Location: Berkeley-based
Apply for this job
*
indicates a required field