Back to jobs

Software Engineer-Data & ML Infrastructure

Company Introduction

At Bot Auto, we are revolutionizing the transportation of goods with our cutting-edge autonomous trucks, enhancing the quality of life for communities around the globe. With the agility of a start-up and the wisdom of seasoned experts, Bot Auto boasts a team that has achieved numerous world-firsts and unparalleled innovations. United by a shared vision, we create miracles and propel the future of transportation. Join us and transform your dreams into reality.

We are seeking a highly skilled and motivated Software Engineer to architect, develop, and scale a robust hybrid-cloud data and machine learning platform from the ground up. This role requires a hands-on coding expert with a deep understanding of distributed storage and compute systems, as well as extensive experience designing and implementing large-scale data and machine learning infrastructures. The ideal candidate will excel in creating efficient data ingestion, transformation, and lake formation pipelines, while managing both relational and NoSQL databases. A strong commitment to data privacy, security, and access control is essential, ensuring the integrity and accessibility of our data lake and database systems.

Key Responsibilities

Data Lake Infrastructure

  • Design and implement scalable data infrastructure, including data lakehouse systems leveraging S3, Data Lake, and Data Catalog, with support for diverse data formats such as Parquet, Avro, and JSON.
  • Integrate modern data lakehouse frameworks, including Delta Lake, Apache Hudi, and Apache Iceberg, to provide versioning, ACID compliance, and optimized storage for high-performance data analytics.
  • Architect, containerize, and orchestrate end-to-end data workflows using Kubernetes (K8s) and distributed computing frameworks, enabling efficient, large-scale data processing and transformation.

Machine Learning & Deep Learning Infrastructure

  • Develop and manage a robust feature (training data) store to support machine learning models with high-quality data features.
  • Design end-to-end data and ML pipelines, from data preparation to model deployment, enabling automated workflows for rapid experimentation and production.
  • Collaborate closely with research scientists to train and optimize deep learning models, supporting model benchmarking, validation, and continuous improvement.

Data Ingestion Framework

  • Build a resilient data ingestion framework to handle large, real-time data streams from core database replica logs, enterprise APIs, and GraphQL sources, facilitating data sharing and replication across systems.
  • Structure multi-layered data storage within the data lake, creating a comprehensive system that includes raw, curated, and warehousing layers to support various analytical and operational needs.

Core Database Management

  • Oversee and optimize core data storage solutions, including relational databases, NoSQL databases, and real-time databases, ensuring reliable and high-performance data access.
  • Implement and manage robust data backup, disaster recovery, and high-availability solutions to protect and maintain critical data assets.

Privacy, Security and Access Control

  • Secure Data Access: Implement access controls and encryption to safeguard data.
  • Privacy Compliance: Ensure data handling meets privacy standards and regulations.
  • Monitoring & Alerts: Set up systems to track data access and identify security risks.
  • Data Masking: Use techniques to protect sensitive data in analytics and ML workflows.

Qualifications

Required:

  • Educational Background: Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related field, or equivalent experience.
  • Experience in Data Infrastructure: Proven experience designing and implementing scalable data lakehouse architectures, including proficiency with cloud storage (e.g., S3), data cataloging, and modern data lake frameworks like Delta Lake, Apache Hudi, or Apache Iceberg.
  • Proficiency in Distributed Systems: Strong experience with distributed computing and container orchestration, specifically with Kubernetes (K8s), Spark, and other large-scale data processing frameworks.
  • Data Engineering & Machine Learning: Skilled in building and managing feature stores, data pipelines, and ML workflows, with experience in ML model training, benchmarking, and production deployment.
  • Strong Programming Skills: Proficiency in programming languages such as Python, SQL, and optionally Scala or Java, along with experience in data frameworks (e.g., Apache Spark, Kafka) and ETL tools.
  • Database Management: Solid understanding of relational and NoSQL databases, real-time databases, and core principles of database backup, recovery, and high-availability strategies.
  • Data Privacy and Security: Knowledgeable in data privacy standards and regulations (e.g., GDPR, CCPA), with hands-on experience implementing access controls, data encryption, and monitoring.
  • Analytical and Problem-Solving Skills: Demonstrated ability to troubleshoot complex data infrastructure issues, optimize performance, and apply innovative solutions in dynamic environments.
  • Collaboration and Communication: Strong interpersonal skills to work closely with cross-functional teams, including data scientists, ML researchers, and engineering teams.

Preferred:

  • Experience with the Autonomous Driving Industry.

Apply for this job

*

indicates a required field

Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in Bot Auto’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Select...
Select...
Race & Ethnicity Definitions

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Select...

Voluntary Self-Identification of Disability

Form CC-305
Page 1 of 1
OMB Control Number 1250-0005
Expires 04/30/2026

Why are you being asked to complete this form?

We are a federal contractor or subcontractor. The law requires us to provide equal employment opportunity to qualified people with disabilities. We have a goal of having at least 7% of our workers as people with disabilities. The law says we must measure our progress towards this goal. To do this, we must ask applicants and employees if they have a disability or have ever had one. People can become disabled, so we need to ask this question at least every five years.

Completing this form is voluntary, and we hope that you will choose to do so. Your answer is confidential. No one who makes hiring decisions will see it. Your decision to complete the form and your answer will not harm you in any way. If you want to learn more about the law or this form, visit the U.S. Department of Labor’s Office of Federal Contract Compliance Programs (OFCCP) website at www.dol.gov/ofccp.

How do you know if you have a disability?

A disability is a condition that substantially limits one or more of your “major life activities.” If you have or have ever had such a condition, you are a person with a disability. Disabilities include, but are not limited to:

  • Alcohol or other substance use disorder (not currently using drugs illegally)
  • Autoimmune disorder, for example, lupus, fibromyalgia, rheumatoid arthritis, HIV/AIDS
  • Blind or low vision
  • Cancer (past or present)
  • Cardiovascular or heart disease
  • Celiac disease
  • Cerebral palsy
  • Deaf or serious difficulty hearing
  • Diabetes
  • Disfigurement, for example, disfigurement caused by burns, wounds, accidents, or congenital disorders
  • Epilepsy or other seizure disorder
  • Gastrointestinal disorders, for example, Crohn's Disease, irritable bowel syndrome
  • Intellectual or developmental disability
  • Mental health conditions, for example, depression, bipolar disorder, anxiety disorder, schizophrenia, PTSD
  • Missing limbs or partially missing limbs
  • Mobility impairment, benefiting from the use of a wheelchair, scooter, walker, leg brace(s) and/or other supports
  • Nervous system condition, for example, migraine headaches, Parkinson’s disease, multiple sclerosis (MS)
  • Neurodivergence, for example, attention-deficit/hyperactivity disorder (ADHD), autism spectrum disorder, dyslexia, dyspraxia, other learning disabilities
  • Partial or complete paralysis (any cause)
  • Pulmonary or respiratory conditions, for example, tuberculosis, asthma, emphysema
  • Short stature (dwarfism)
  • Traumatic brain injury
Select...

PUBLIC BURDEN STATEMENT: According to the Paperwork Reduction Act of 1995 no persons are required to respond to a collection of information unless such collection displays a valid OMB control number. This survey should take about 5 minutes to complete.