Back to jobs
New

Principal Speech Data Linguist

Remote - United States

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role: 

As speech and audio models get better, the human role gets harder, not easier — it moves from producing transcripts to defining what a correct one is, adjudicating the cases models still get wrong, and designing the human-in-the-loop workflows that keep improving them. Innodata runs high-volume segmentation and transcription workflows for the customers and frontier labs building these models, and we are hiring a principal-level linguist to own the linguistic standards and quality behind that work — today, and as the workflows evolve alongside the models over the next two years. 

This is the applied-expert counterpart to our Speech & Audio Research Scientist. You set the standards the models are trained and measured against, and you understand the big picture: how different transcription and segmentation methods change what a model learns, and how that ripples into the speech and content-understanding systems our partners are building. You know the research and you know the tools — from IPA and acoustic analysis to forced alignment and the ASR engines our partners benchmark against — but your leverage is linguistic judgment and standard-setting at scale, not building models yourself. 

What You’ll Own:

  • You will own the linguistic foundation of Innodata's segmentation and transcription work across languages, domains, and use cases. Concretely, you will: 
  • Define transcription and segmentation standards, style guides, and annotation conventions — verbatim and clean/intelligent verbatim, IPA and phonetic transcription, timestamping and boundary segmentation, speaker labeling and diarization labels, disfluencies and non-speech events, code-switching, and orthographic conventions. 
  • Establish and run the quality frameworks behind that work: rubrics, error taxonomies, adjudication processes, inter-annotator agreement, and human QA at scale. 
  • Own the quality lifecycle for transcription and segmentation deliverables end to end — pre-processing and normalization of incoming data, quality checks at the point of acceptance, post-processing and pre-delivery validation against spec, and report creation and packaging for delivery — partnering with delivery operations on execution at scale. 
  • Design the human-in-the-loop workflows themselves — deciding where human review, correction, and adjudication add the most value as ASR quality rises, so our experts spend their time on what the models still can't do rather than on what they already can. 
  • Handle the linguistically hard cases models fail on — accented and dialectal speech, low-resource and multilingual audio, overlapping speech, domain jargon (medical, legal, technical), and noisy acoustic conditions. 
  • Partner with the Speech & Audio Research Scientist to turn model objectives into transcription and segmentation specifications, and to work out how different transcription methods — verbatim versus clean, phonetic versus orthographic, and how audio is segmented and labeled — affect the training and evaluation of ASR, TTS, and speech and content-understanding models. 
  • Train, calibrate, and mentor expert transcribers and reviewers, and build the onboarding and calibration that keep quality consistent as the work scales. 
  • Represent Innodata's transcription and segmentation approach to the customers and frontier labs we partner with, and contribute to the methodology and best-practice documentation that make our work legible to their teams. 

You’ll Thrive in This Role If You Have:

  • Substantial industry experience (typically 8+ years) in transcription, segmentation, and speech-data quality — enough that you have authored standards, not only followed them. This is a principal-level role, and we weight practical depth heavily. 
  • A Bachelor's degree in linguistics, phonetics, or computational linguistics, or a closely related field, is required — with a strong foundation in phonetics, phonology, and sociolinguistics so that IPA, prosody, disfluency, dialect, and register are native concepts. An advanced degree is preferred. 
  • A big-picture grasp of how transcription and segmentation choices flow downstream into modeling — how different methods change what speech and content-understanding models learn, and therefore which method fits which modeling objective. You can explain to a model builder why a transcription decision matters. 
  • Fluency in phonetic transcription and IPA, plus hands-on experience with acoustic and phonetic analysis of speech — spectrograms, formants, pitch and prosody, and segment boundaries — applied to real, messy speech data at scale (for example in Praat). 
  • Deep experience with audio segmentation and its conventions — utterance and turn boundaries, timestamping, and speaker and diarization labeling — across real-world audio. 
  • Hands-on fluency with the modern speech stack: Whisper and the commercial ASR engines your partners benchmark against (such as AssemblyAI, Deepgram, Rev, and Speechmatics), forced alignment (for example the Montreal Forced Aligner), and annotation tools such as ELAN. 
  • Comfort scripting for speech-data work — Python for batch processing, QA, and metrics such as inter-annotator agreement and WER, plus regular expressions and Praat scripting — enough to work fluently with data and pipelines without needing an engineer for every task. 
  • Practical data-management skills across the delivery lifecycle — pre-processing, acceptance-stage quality checks, post-processing, pre-delivery validation, and report creation and packaging — so deliverables leave the door correct, consistent, and well documented. 
  • Multilingual capability and hands-on experience with accented, dialectal, and code-switched speech; low-resource languages a strong plus. 
  • A point of view on how human-in-the-loop workflows should evolve as models improve — where humans stay in the loop, where they move up to adjudication and standard-setting, and how to measure the difference. 
  • Strong written and verbal communication, comfortable working directly with research scientists and interfacing with the customers and frontier labs we partner with. 
  • Bonus: responsible-AI considerations for speech, such as bias across accents and dialects and privacy and consent in voice data. 

 

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

 

 

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams. 

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...
Select...
Select...
Select...
Select...
Select...

If yes, you can always opt-out by replying STOP. 

If you opt-in to the service, Innodata Services and its affiliates may send you SMS messages to notify you of any updates to your application status and to engage in discussion throughout your application process.
 
Message frequency may vary depending on the type of communication. Standard message and data rates may apply according to your mobile carrier’s pricing plan.
 
You may opt out of receiving SMS messages at any time by replying STOP to any SMS message received from us. Once you opt out, you will no longer receive SMS communications from us unless you re-subscribe.
 
For assistance, reply HELP to any message or contact us at talent@innodata.com
 
Mobile information will not be shared with third parties or affiliates for marketing or promotional purposes

Privacy Policy

Select...

U.S. Standard Demographic Questions

We invite applicants to share their demographic background. If you choose to complete this survey, your responses may be used to identify areas of improvement in our hiring process.
Select...
Select...
Select...
Select...
Select...
Select...

Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in Innodata Inc.’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Select...
Select...
Race & Ethnicity Definitions

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Select...