Job Application for Staff Data Engineer, AI at Biohub

Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere.

The Team

Biohub is a 501(c)(3) biomedical research organization building the first large-scale scientific initiative combining frontier AI with frontier biology to solve disease. We build the technology to help scientists around the world use AI-powered biology to study how cells operate, organize, and work as part of systems to understand why disease happens and how to correct it. With our compute capacity, AI research and engineering, and state-of-the-art technology for measuring, imaging, and programming biology, we are enabling scientists worldwide to use AI-powered biology to advance our understanding of human health.

The Opportunity

The role is part of the Data Engineering team, which focuses on owning the strategy, sourcing and implementation for data supporting AI research and development. Our goal is to maximize the speed, agility, and capability of biological AI research by connecting public data resources and Biohub's experimental platforms to AI systems. The data that trains biological frontier models comes in dozens of modalities (sequences, images, spatial coordinates, time series, molecular structures, metadata, publication artifacts) each with its own noise characteristics, biases, and information content. The question of how to represent this data for learning is one of the most important open problems in biological AI.

As a Staff Data Engineer at Biohub, you'll be designing systems that ingest data from public repositories, transform heterogeneous biological formats into AI-ready datasets, combine that with proprietary datasets, and deliver training datasets to researchers pushing the boundaries of what's possible in biological AI. The infrastructure you build will directly shape what our models can learn.

We're a small team with significant resources and long time horizons. We use AI tools aggressively in our own work—Claude Code, agents for workflow automation, LLMs for metadata extraction. We care about code quality, operational reliability, and building systems that scale. And we care about the biology: we want engineers who can recognize when a pipeline output is technically correct but scientifically wrong.

If you want to work at the intersection of large-scale infrastructure and frontier science, with real autonomy and the chance to build something genuinely new, we'd like to talk.

What You'll Do

Design and build data pipelines that process genomic and imaging data at petabyte scale
Solve performance and bandwidth challenges with creative engineering
Build agent-based systems for automated dataset curation, quality control, and workflow generation
Create tooling for data cataloging and registration that makes datasets discoverable and accessible
Collaborate with AI Research teams to translate model requirements into data specifications, and with our scientists to integrate public and internal data into large-scale ai-ready datasets
Improve pipeline reliability and observability, working toward 99%+ success rates without manual intervention

What You'll Bring

8+ years experience building reliable, operable data systems at petabyte scale
Strong software engineering fundamentals
Experience deploying distributed computing frameworks like Databricks, Spark, or Ray for large-scale data processing
Experience building and deploying large scale data platform solutions using infrastructure as code like Terraform or CDK
Experience with cloud infrastructure (AWS preferred) and on prem infrastructure
Comfort with ambiguity; ability to make progress when requirements are evolving
Interest in AI-native development practices and tooling
Nice to have: Background in computational biology, bioinformatics, or life sciences and experience with genomics datasets and formats (FASTQ, BAM, VCF) or imaging formats (OME-Zarr, HDF5)

Compensation

The future anticipated Redwood City, CA, and New York City, NY base pay range for a role in this field is $241,000–$301,000 annually. Compensation ranges will vary based on job-related skills, level of experience, and knowledge. Actual placement in range is based on job-related skills and experience, as evaluated throughout the interview process.

Better Together

As we grow, we’re excited to strengthen in-person connections and cultivate a collaborative, team-oriented environment. This role is a hybrid position requiring you to be onsite for at least 60% of the working month, approximately 3 days a week, with specific in-office days determined by the team’s manager. The exact schedule will be at the hiring manager's discretion and communicated during the interview process.

Benefits for the Whole You

We’re thankful to have an incredible team behind our work. To honor their commitment, we offer a wide range of benefits to support the people who make all we do possible.

Provides a generous employer match on employee 401(k) contributions to support planning for the future.
Paid time off to volunteer at an organization of your choice.
Funding for select family-forming benefits.
Relocation support for employees who need assistance moving

If you’re interested in a role but your previous experience doesn’t perfectly align with each qualification in the job description, we still encourage you to apply as you may be the perfect fit for this or another role.

#LI-Hybrid #LI-Onsite

First Name

Last Name

Country

Phone

Location (City)

Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf

LinkedIn Profile

This role is a hybrid role and will require you to be onsite approximately 3 days a week in office at our Redwood City headquarters or New York office. Will you be able to commit to this?

Select...

Do you know anyone that works at CZI or Biohub? If yes, who?

Are you currently eligible to work in the United States of America?

Select...

Do you now or in the future require visa sponsorship to continue working in the United States?

Select...

Are you actively interviewing and do you have any urgency (offer deadlines, in late stages of the hiring process at another company or companies)?

What is your earliest approximate start date?

Have you interviewed with Biohub or any Chan Zuckerberg entity in the past (i.e. Learning Commons, our education initiative)?

Select...

Do you have experience with event driven architecture or big data processing frameworks?

What's the largest-scale system you've worked on? What did you build, and how large was the dataset?

What kind of data have you built from day to day?

What kind of system do you use?

What is your strongest coding language?

Reasonable Accommodation, Data Privacy, Background Check and Artificial Intelligence Usage Notice.

Select...

Reasonable Accommodation Notice
The organization provides (and state and federal law requires) reasonable accommodations to be provided to qualified applicants with disabilities. Your recruiter will work with you during the interview process should you require any such accommodations. Examples of reasonable accommodation include making a change to the application process or work procedures, providing documents in an alternate format, using a sign language interpreter, or using specialized equipment.

Applicant Privacy Notice
To learn more about how we use the information you submit, please see our Privacy Notice for Job Applicants.

Background Check NoticeAs part of our hiring process, all offers of employment are contingent upon the successful completion of a background check. By submitting your application, you acknowledge that you will be required to undergo a background check prior to employment.

Artificial Intelligence Usage Notice
We use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications and analyzing resumes. These tools assist our recruitment team but do not replace human judgment. Hiring decisions are ultimately made by humans. If you have questions about this once your are in our hiring process, please contact your Recruiter.

Voluntary Self Identification

For reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in the organization’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Race & Ethnicity Definitions

Veteran Status

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. Classification of protected categories is as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Gender

Select...

Are you Hispanic/Latino

Select...

What is your Race/Ethnicity

Select...

Veteran Status

Select...

Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in Biohub’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Gender

Select...

Are you Hispanic/Latino?

Select...

Race & Ethnicity Definitions

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows:

Veteran Status

Select...

Voluntary Self-Identification of Disability

Form CC-305

Page 1 of 1

OMB Control Number 1250-0005

Expires 04/30/2026

Why are you being asked to complete this form?

We are a federal contractor or subcontractor. The law requires us to provide equal employment opportunity to qualified people with disabilities. We have a goal of having at least 7% of our workers as people with disabilities. The law says we must measure our progress towards this goal. To do this, we must ask applicants and employees if they have a disability or have ever had one. People can become disabled, so we need to ask this question at least every five years.

Completing this form is voluntary, and we hope that you will choose to do so. Your answer is confidential. No one who makes hiring decisions will see it. Your decision to complete the form and your answer will not harm you in any way. If you want to learn more about the law or this form, visit the U.S. Department of Labor’s Office of Federal Contract Compliance Programs (OFCCP) website at www.dol.gov/ofccp.

How do you know if you have a disability?

A disability is a condition that substantially limits one or more of your “major life activities.” If you have or have ever had such a condition, you are a person with a disability. Disabilities include, but are not limited to:

Alcohol or other substance use disorder (not currently using drugs illegally)
Autoimmune disorder, for example, lupus, fibromyalgia, rheumatoid arthritis, HIV/AIDS
Blind or low vision
Cancer (past or present)
Cardiovascular or heart disease
Celiac disease
Cerebral palsy
Deaf or serious difficulty hearing
Diabetes
Disfigurement, for example, disfigurement caused by burns, wounds, accidents, or congenital disorders
Epilepsy or other seizure disorder
Gastrointestinal disorders, for example, Crohn's Disease, irritable bowel syndrome
Intellectual or developmental disability
Mental health conditions, for example, depression, bipolar disorder, anxiety disorder, schizophrenia, PTSD
Missing limbs or partially missing limbs
Mobility impairment, benefiting from the use of a wheelchair, scooter, walker, leg brace(s) and/or other supports
Nervous system condition, for example, migraine headaches, Parkinson’s disease, multiple sclerosis (MS)
Neurodivergence, for example, attention-deficit/hyperactivity disorder (ADHD), autism spectrum disorder, dyslexia, dyspraxia, other learning disabilities
Partial or complete paralysis (any cause)
Pulmonary or respiratory conditions, for example, tuberculosis, asthma, emphysema
Short stature (dwarfism)
Traumatic brain injury

Disability Status

Select...

PUBLIC BURDEN STATEMENT: According to the Paperwork Reduction Act of 1995 no persons are required to respond to a collection of information unless such collection displays a valid OMB control number. This survey should take about 5 minutes to complete.

Staff Data Engineer, AI

The Team

The Opportunity

What You'll Do

What You'll Bring

Compensation

Better Together

Benefits for the Whole You

Apply for this job

Voluntary Self Identification

Voluntary Self-Identification

Voluntary Self-Identification of Disability