Back to jobs
New

Research Data Scientist

Remote - United States

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role: 

We are looking for a highly skilled Research Data Scientist – GenAI/LLM to join our AI/LLM Delivery Unit and work on research-driven AI/ML initiatives involving Generative AI, Large Language Models (LLMs), NLP, multimodal AI, model evaluation, and AI data.

The role combines strong research and analytical capabilities with hands-on AI/ML expertise, requiring the candidate to design experiments, develop evaluation methodologies, analyze complex datasets, build research prototypes, and translate research findings into practical AI/ML solutions.

The ideal candidate will have a strong research orientation, excellent statistical and analytical skills, and the ability to work collaboratively with researchers, data scientists, AI/ML engineers, domain experts, and client-facing teams.

What You’ll Own:

AI/ML & Generative AI Research:

  • Conduct independent and collaborative research in Generative AI, LLMs, NLP, multimodal AI, machine learning, model evaluation, and AI data.
  • Formulate research questions and translate complex AI/ML problems into structured research methodologies and experiments.
  • Design, execute, and analyze experiments to evaluate and improve AI/ML models and solutions.
  • Build analytical models, prototypes, and research pipelines using Python and relevant ML frameworks.
  • Stay current with emerging research, methodologies, papers, and developments in GenAI, LLMs, NLP, multimodal models, and AI evaluation.

LLM & Model Evaluation:

  • Develop and implement LLM evaluation frameworks, benchmarks, datasets, and evaluation criteria.
  • Evaluate models for accuracy, robustness, bias, hallucination, reasoning, relevance, response quality, and other performance dimensions.
  • Conduct model benchmarking, error analysis, comparative analysis, and performance evaluation.
  • Work on areas such as RAG, SFT, RLHF/DPO, prompt engineering, fine-tuning, embeddings, and LLM optimization, as applicable.
  • Identify model and data gaps and recommend improvements to enhance model performance and reliability.

Data Science & Statistical Research:

  • Collect, clean, analyze, and interpret large and complex structured and unstructured datasets.
  • Perform EDA, statistical analysis, hypothesis testing, significance testing, correlation analysis, sampling, and error analysis.
  • Develop data-driven insights and identify patterns, trends, and relationships relevant to AI/ML research.
  • Apply appropriate statistical and quantitative methodologies to validate research findings

AI Data & Dataset Development:

  • Develop and evaluate datasets, sampling methodologies, taxonomies, annotation frameworks, data quality frameworks, and evaluation criteria for AI/ML models.
  • Analyze data quality and identify issues affecting model performance.
  • Collaborate with annotation, data engineering, and AI/ML teams to improve AI training and evaluation data.
  • Translate data and research findings into actionable recommendations for improving AI system

Research & Innovation:

  • Contribute to research papers, technical reports, whitepapers, patents, benchmarks, internal publications, and other research outputs, where applicable.
  • Identify opportunities to apply emerging research and technologies to real-world AI and data challenges.
  • Explore new methodologies, models, datasets, and evaluation approaches to improve AI capabilities.
  • Contribute to capability building and innovation within the AI/LLM practice

Collaboration & Stakeholder Engagement:

  • Work closely with researchers, data scientists, AI/ML engineers, data/annotation teams, domain experts, and delivery teams.
  • Present research findings, analytical insights, and technical recommendations to senior technical stakeholders.
  • Translate complex research and technical concepts into clear, actionable recommendations.
  • Where required, participate in client-facing technical discussions and presentations and help translate business requirements into AI/ML solutions.

You’ll Thrive in This Role If You Have:

  • Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Statistics, Mathematics, Computational Science, or a related discipline.
  • Bachelor’s/Master’s degree from IITs, NITs, or other premier engineering/research institutions is strongly preferred.
  • 4–7 years of hands-on research experience in AI/ML, Data Science, NLP, Generative AI, LLMs, or related areas.
  • Strong demonstrated research experience with the ability to independently formulate research questions, design experiments, analyze results, and communicate findings.
  • Demonstrated research track record through research publications, patents, conference presentations, open-source contributions, or significant AI/ML research projects.
  • Candidates with publications in reputed conferences/journals and a strong academic/research profile will be preferred
  • Strong proficiency in Python and SQL.
  • Strong hands-on experience with NumPy, Pandas, Scikit-learn, and preferably PyTorch/TensorFlow.
  • Strong understanding of:
    • Machine learning algorithms
    • Statistics and experimentation
    • Data analysis and feature engineering
    • Model evaluation and performance metrics
    • Hypothesis testing and statistical inference
  • Hands-on exposure to LLMs, NLP, Generative AI, and multimodal AI.
  • Experience with one or more of RAG, LLM evaluation, prompt engineering, fine-tuning, SFT, RLHF/DPO, embeddings, or model benchmarking.
  • Experience working with large-scale structured and unstructured datasets.
  • Familiarity with Git and cloud platforms such as AWS, Azure, or GCP is desirable.

  •  

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

 

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams. 

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in Innodata Inc.’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Select...
Select...
Race & Ethnicity Definitions

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Select...

Voluntary Self-Identification of Disability

Form CC-305
Page 1 of 1
OMB Control Number 1250-0005
Expires 07/31/2029

Why are you being asked to complete this form?

We are a federal contractor or subcontractor. The law requires us to provide equal employment opportunity to qualified people with disabilities. We have a goal of having at least 7% of our workers as people with disabilities. The law says we must measure our progress towards this goal. To do this, we must ask applicants and employees if they have a disability or have ever had one. People can become disabled, so we need to ask this question at least every five years.

Completing this form is voluntary, and we hope that you will choose to do so. Your answer is confidential. No one who makes hiring decisions will see it. Your decision to complete the form and your answer will not harm you in any way. If you want to learn more about the law or this form, visit the U.S. Department of Labor’s Office of Federal Contract Compliance Programs (OFCCP) website at www.dol.gov/ofccp.

How do you know if you have a disability?

A disability is a condition that substantially limits one or more of your “major life activities.” If you have or have ever had such a condition, you are a person with a disability. Disabilities include, but are not limited to:

  • Alcohol or other substance use disorder (not currently using drugs illegally)
  • Autoimmune disorder, for example, lupus, fibromyalgia, rheumatoid arthritis, HIV/AIDS
  • Blind or low vision
  • Cancer (past or present)
  • Cardiovascular or heart disease
  • Celiac disease
  • Cerebral palsy
  • Deaf or serious difficulty hearing
  • Diabetes
  • Disfigurement, for example, disfigurement caused by burns, wounds, accidents, or congenital disorders
  • Epilepsy or other seizure disorder
  • Gastrointestinal disorders, for example, Crohn's Disease, irritable bowel syndrome
  • Intellectual or developmental disability
  • Mental health conditions, for example, depression, bipolar disorder, anxiety disorder, schizophrenia, PTSD
  • Missing limbs or partially missing limbs
  • Mobility impairment, benefiting from the use of a wheelchair, scooter, walker, leg brace(s) and/or other supports
  • Nervous system condition, for example, migraine headaches, Parkinson’s disease, multiple sclerosis (MS)
  • Neurodivergence, for example, attention-deficit/hyperactivity disorder (ADHD), autism spectrum disorder, dyslexia, dyspraxia, other learning disabilities
  • Partial or complete paralysis (any cause)
  • Pulmonary or respiratory conditions, for example, tuberculosis, asthma, emphysema
  • Short stature (dwarfism)
  • Traumatic brain injury
Select...

PUBLIC BURDEN STATEMENT: According to the Paperwork Reduction Act of 1995 no persons are required to respond to a collection of information unless such collection displays a valid OMB control number. This survey should take about 5 minutes to complete.