NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Monitoring & Evaluation Specialist in United States of America
2 months ago
Apply with autofill
Apply with autofill
Innodatainc·2 months ago
2 months ago

QA / Evaluation Lead

Hybrid, United States of AmericaHybridSenior · 36+ yearsMonitoring & Evaluation Specialist

Sign up free to see how well your resume matches this role.

Boost your chances at innodatainc

How you compare FREE

?
Your scoreYour score: not yet known
→
16
Top 10%Top 10%: 16 out of 100

Top 10% of NextRaise users matched against Monitoring & Evaluation Specialist roles in United States.

Must-have skills for this role

  • iaa
  • data quality
  • secret clearance
  • python

PDF or DOCX · no account needed

Apply faster with autofill FREEinnodatainc uses Greenhouse - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Design and own the inter-annotator agreement (IAA) methodology for the Phase 1 demonstration corpus — metric selection (Cohen's kappa, Fleiss, Krippendorff's alpha), sampling design, adjudication workflow, and agreement thresholds
  • Define evaluation framework architecture: test and evaluation plans, IAA targets, drift detection gates, and model performance metrics per SOW Section 2.9
  • Configure and operate sampling-based quality control across the self-service and white-glove annotation paths during Phase D corpus production
  • Design and implement confidence-threshold escalation routing from automated annotation to senior-annotator adjudication
  • Validate quality scoring and IAA computation within the Innodata data layer
  • Support AI Solutions Engineer on evaluation design for SAM 2 and Frontier model API validation — define what 'good enough' looks like quantitatively
  • Produce evaluation framework documentation for the Phase 1 NPP closeout package, including per-DataCard documentation with the SA

What they're looking for

  • Bachelor's degree in Statistics, Data Science, Computer Science, or related quantitative field required; Master's degree preferred. Equivalent experience may substitute for degree on a 2-for-1 basis.
  • 6+ years total professional experience, 4+ years in data quality, evaluation methodology, or QA on AI/ML programs
  • IAA methodology expertise — Cohen's kappa, Fleiss' kappa, Krippendorff's alpha: hands-on, not theoretical
  • Evaluation framework design for AI/ML training data programs
  • QC process design: sampling methodology, escalation workflows, adjudication protocols
  • Python for QC tooling, metric computation, and statistical analysis
  • Active Secret clearance with TS/SCI eligibility

Nice to have

  • Master's degree
  • Prior DoD or IC data quality program experience
  • CVAT or equivalent annotation platform QC workflow configuration
  • Drift detection and model monitoring methodology
  • Experience with FMV / video annotation quality standards

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

About the Program: 

Innodata's Federal Practice builds the trusted data layer for critical infrastructure Trust & Safety work. Partnering with a leading systems integrator, we're delivering a modern, governed data services platform in a secure federal (IL4) environment. Over an intensive 20-week phase, you'll help stand up a data services storefront, a DataCard governance framework, synthetic data integration, and Databricks write-back capabilities.

About the Role: 

As the QA/Evaluation Lead, you'll own quality and evaluation across the platform. You'll design the evaluation framework that measures whether our data services and outputs meet the bar, build repeatable test and validation processes, and give the team an objective read on readiness at each milestone. Partnering with the Delivery Owner and engineering leads, you'll turn quality from an afterthought into a measurable, demonstrable strength. It's a role for someone who thinks rigorously about evaluation and takes pride in evidence-backed quality.

Key Responsibilities:

  • Design and own the inter-annotator agreement (IAA) methodology for the Phase 1 demonstration corpus — metric selection (Cohen's kappa, Fleiss, Krippendorff's alpha), sampling design, adjudication workflow, and agreement thresholds
  • Define evaluation framework architecture: test and evaluation plans, IAA targets, drift detection gates, and model performance metrics per SOW Section 2.9
  • Configure and operate sampling-based quality control across the self-service and white-glove annotation paths during Phase D corpus production
  • Design and implement confidence-threshold escalation routing from automated annotation to senior-annotator adjudication
  • Validate quality scoring and IAA computation within the Innodata data layer
  • Support AI Solutions Engineer on evaluation design for SAM 2 and Frontier model API validation — define what 'good enough' looks like quantitatively
  • Produce evaluation framework documentation for the Phase 1 NPP closeout package, including per-DataCard documentation with the SA

Must-Have Qualifications:

  • Bachelor's degree in Statistics, Data Science, Computer Science, or related quantitative field required; Master's degree preferred. Equivalent experience may substitute for degree on a 2-for-1 basis.
  • 6+ years total professional experience, 4+ years in data quality, evaluation methodology, or QA on AI/ML programs
  • IAA methodology expertise — Cohen's kappa, Fleiss' kappa, Krippendorff's alpha: hands-on, not theoretical
  • Evaluation framework design for AI/ML training data programs
  • QC process design: sampling methodology, escalation workflows, adjudication protocols
  • Python for QC tooling, metric computation, and statistical analysis
  • Active Secret clearance with TS/SCI eligibility

Nice-to-Have Qualifications:

  • Prior DoD or IC data quality program experience
  • CVAT or equivalent annotation platform QC workflow configuration
  • Drift detection and model monitoring methodology
  • Experience with FMV / video annotation quality standards

The expected hourly salary range for this position is $45 to $50 p/hour, based on experience, skills, and qualifications.

Note to Candidates: 

This role is not a project manager with QC responsibilities — it is a methodology expert who owns the intellectual framework behind data quality on a federal AI program. The right candidate can walk into a meeting with Government evaluators and explain exactly why the evaluation design produces trustworthy labels. That conversation is part of Phase 2 positioning.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams. 

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.

Company

Innodatainc
Hybrid, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Innodatainc's careers site·first seen 8 Aug 2026·last verified 9 Sept 2026·How we source jobs

Similar jobs

  • 2027 Internship Evaluation Engineer, Metric Prototyping at bedrock-roboticsSan Francisco, United States of America–match not yet calculated
  • Manager, Evaluation and Insights at aamcWashington DC, United States of America–match not yet calculated
  • Remote Patient Monitoring Specialist at vironix-aiUnited States–match not yet calculated
  • Sr. Manager, Security — Continuous Monitoring v 2.0 at DatabricksRemote - California–match not yet calculated
  • NOSC Monitoring Analyst at gditUSA FL MacDill AFB–match not yet calculated

Browse more jobs

  • Monitoring & Evaluation Specialist jobs in United States
  • Community Health Worker jobs in United States
  • Public Health Officer jobs in United States
  • Field Officer jobs in United States
  • Monitoring & Evaluation Specialist jobs in India
  • Monitoring & Evaluation Specialist jobs in United Kingdom