Back to jobs
New

Robotics Researcher — Manipulation (Omakase Zen)

US

 

 

Location: United States (SF Bay Area preferred)

Employment: Full-time

 

Aiming to be the world’s leading humanoid company, — humanoids that work reliably in the real world every day, built with mass-production discipline rather than demo discipline.

Three vertically integrated components:

Omakase D1— our own humanoid hardware

Omakase Zen — the manipulation intelligence foundation (this role)

Omakase OS — the orchestration software that runs robots in the field

Our moat is data. Through our own fleet and a partnership with a leading spot-labor platform, we have structurally exclusive access to real-world teleoperation and egocentric human data at a scale competitors cannot replicate. Zen turns that flywheel into manipulation foundation models.

 

 The Role

You will train the models that make our robots' hands work.

Today that means post-training VLA policies (π0.5-class, ACT) on our own teleoperation data and deploying them onto real robots doing real jobs in Japan. Next it means pre-training a cross-embodiment manipulation foundation model on our proprietary dataset.

You own the loop end to end: data curation → training → evaluation → real-robot deployment → failure analysis → back to data. We do not have a separate team that "puts the model on the robot." That handoff is where robot learning usually goes to die.

 

 What you would actually work on

Concrete problems currently open on our side:

Action-space representation. Our end-effector-space policies are systematically weaker on rotation than on translation. Until the representation is fixed, EE-space deployment is blocked. This is an open problem we would want your opinion on in the interview.

Cross-embodiment transfer. We collect human hand trajectories and retarget them to the robot. How much of that transfers, and in which representation, is not yet settled by our own measurements.

Evaluation that predicts real-robot success. Open-loop MSE on held-out episodes is cheap and weakly correlated with what happens on the robot. We build sim-based harnesses and structured real-robot trials, and we would like them to disagree less.

Making failures legible. When a policy misses a grasp, the useful question is which part of the pipeline was wrong — the data, the representation, the training, or the hardware. We invest in being able to answer that.

 

Responsibilities

- Train and post-train VLA / imitation-learning policies (diffusion and flow-matching action heads, ACT, π-class models) on real teleoperation data

- Design and run the pre-training effort for our cross-embodiment manipulation foundation model

- Build rigorous evaluation: open-loop metrics on held-out real data, sim-based eval harnesses, structured real-robot trials

- Own multi-node distributed training (we operate H100/H200 clusters) — throughput, dataloading, checkpointing

- Work with the data team on dataset schemas, quality gates, and retargeting (EE-space cross-embodiment representations)

- Deploy policies to real robots with the OS team and drive failure analysis back into data and training

 

 Required Qualifications

We set a high bar here.

- MS/PhD in ML, robotics, or CS — or an equivalent track record that speaks for itself

3+ years training large neural models, including at least one substantial robot-learning or VLA project you can walk us through in depth — what failed, what you measured, what you would do differently

- Demonstrated experience with imitation learning / VLA architectures (π0-class, OpenVLA, RT-class, ACT, diffusion policies). Fine-tuning through an API does not qualify

- Strong PyTorch or JAX; comfortable owning multi-GPU / multi-node training runs end to end

- Experience evaluating policies beyond loss curves: held-out real-data metrics, real-robot success rates, ablations

- Evidence of top-tier work: publications (CoRL / RSS / ICRA / NeurIPS / ICML) or policies you shipped onto physical robots in production or serious field trials

 

 Preferred

- Cross-embodiment training or action-space retargeting (EE-space representations, IK-aware pipelines)

- Teleoperation data collection systems; data-quality tooling for robot datasets

- Sim-to-real and simulation-based evaluation (Isaac, MuJoCo, Genesis)

- World models or video pre-training for robotics

- Japanese is not required — Zen operates in English

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


We kindly ask for your current and expected compensation to ensure alignment as we move forward in the process.

(※This information will not be shared with hiring managers and will not affect screening.)

 

選考を進めるにあたり、報酬条件のすり合わせのため、現在の年収およびご希望年収についてお伺いしております。

※本情報は採用担当者のみが確認し、採用マネージャーには共有されません。また、選考結果に直接影響するものではありません。

This information will help us ensure that we're aligned with your compensation expectations as we move forward in the process.

(※This information will not be shared with hiring managers and will not affect screening.)

 

選考を進めるにあたり、報酬条件のすり合わせのため、現在の年収およびご希望年収についてお伺いしております。

※本情報は採用担当者のみが確認し、採用マネージャーには共有されません。また、選考結果に直接影響するものではありません。

Select...