NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Monitoring & Evaluation Specialist in United States of America
1 day agoBe an early applicant
Apply with autofill
Apply with autofill
Innodatainc·1 day ago
1 day agoBe an early applicant

Quality Lead, Agentic AI Workflow Evaluation

In Office, United States of AmericaRemoteSenior · 36+ yearsMonitoring & Evaluation Specialist

Sign up free to see how well your resume matches this role.

Boost your chances at innodatainc

How you compare FREE

?
Your scoreYour score: not yet known
→
17
Top 10%Top 10%: 17 out of 100

Top 10% of NextRaise users matched against Monitoring & Evaluation Specialist roles in United States.

Must-have skills for this role

  • annotation
  • red-teaming
  • rlhf
  • python

PDF or DOCX · no account needed

Apply faster with autofill FREEinnodatainc uses Greenhouse - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Own the quality system for the engagement: audit design, sampling strategy, scoring standards, and how quality gets measured and reported
  • Build that system with the customer's quality leads where none exists, and where one does, operate it and recommend concrete improvements based on what the data shows
  • Re-score a sample of reviewer output as a second pass; identify error patterns rather than isolated mistakes
  • Run calibration sessions: surface disagreement, work it to resolution, and document the reasoning so the outcome holds for future cases
  • Maintain rubric health — flag criteria that are ambiguous, overlapping, or silent on cases the team keeps hitting, and drive revisions through the customer
  • Train and onboard new reviewers, including nesting plans, ramp criteria, and the judgment call on when someone is production-ready
  • Give the Engagement Manager the evidence behind performance conversations: who is drifting, on what, and whether coaching is working
  • Report quality trends to the Engagement Manager and, alongside them, to the customer
  • Deputize for the Engagement Manager on delivery operations during absences
  • Maintain information security, privacy, and facility access practices required by the customer's onsite environment

What they're looking for

  • Bachelor's degree or equivalent practical experience
  • 4+ years in quality assurance, quality management, or senior review work within annotation, evaluation, trust and safety, or a similarly judgment-intensive domain
  • Direct experience owning a quality function: you designed the audit, not just executed someone else’s
  • Significant experience with AI/ML evaluation work: annotation, red-teaming, RLHF, model or agent evaluation, or trust and safety review
  • Hands-on familiarity with agentic systems: tool use, multi-step task execution, sandboxed environments, and common failure modes
  • Demonstrated ability to run calibration with peers — including holding a position under disagreement and changing it when the argument is better
  • Strong written communication; able to document a scoring standard clearly enough that a reviewer can apply it and an auditor can check it
  • Comfortable in spreadsheets and in a dashboarding tool, with enough Python or SQL to pull and slice your own data (you will not be asked to build interfaces)
  • Experience training or onboarding reviewers into rubric-based work

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role: 

We are standing up a dedicated onsite team to evaluate complex, real-world agentic AI workflows for a frontier AI customer. Reviewers work through ambiguous, multi-step scenarios inside isolated test environments, assessing whether AI agents complete tasks safely, respect user intent and consent, and hold up under close scrutiny. The Quality Lead is the person accountable for whether that output is any good. 

This is a senior individual contributor role. You will not manage the reviewers — that sits with the Engagement Manager — but you set the standard they are held to. You own the audit sample, run calibration, keep the rubric usable as real cases stress it, and train reviewers into the work. You are also the deputy: when the Engagement Manager is out, the engagement runs on you. 

The quality approach here is not fully defined. We expect you to build it in partnership with the customer's quality leads, or at minimum to take what they have, run it honestly, and come back with specific recommendations for where it falls short. 

What You’ll Own:

  • Own the quality system for the engagement: audit design, sampling strategy, scoring standards, and how quality gets measured and reported 
  • Build that system with the customer's quality leads where none exists, and where one does, operate it and recommend concrete improvements based on what the data shows 
  • Re-score a sample of reviewer output as a second pass; identify error patterns rather than isolated mistakes 
  • Run calibration sessions: surface disagreement, work it to resolution, and document the reasoning so the outcome holds for future cases 
  • Maintain rubric health — flag criteria that are ambiguous, overlapping, or silent on cases the team keeps hitting, and drive revisions through the customer 
  • Train and onboard new reviewers, including nesting plans, ramp criteria, and the judgment call on when someone is production-ready 
  • Give the Engagement Manager the evidence behind performance conversations: who is drifting, on what, and whether coaching is working 
  • Report quality trends to the Engagement Manager and, alongside them, to the customer 
  • Deputize for the Engagement Manager on delivery operations during absences 
  • Maintain information security, privacy, and facility access practices required by the customer's onsite environment 

You’ll Thrive in This Role If You Have:

  • Bachelor's degree or equivalent practical experience 
  • 4+ years in quality assurance, quality management, or senior review work within annotation, evaluation, trust and safety, or a similarly judgment-intensive domain 
  • Direct experience owning a quality function: you designed the audit, not just executed someone else’s 
  • Significant experience with AI/ML evaluation work: annotation, red-teaming, RLHF, model or agent evaluation, or trust and safety review 
  • Hands-on familiarity with agentic systems: tool use, multi-step task execution, sandboxed environments, and common failure modes 
  • Demonstrated ability to run calibration with peers — including holding a position under disagreement and changing it when the argument is better 
  • Strong written communication; able to document a scoring standard clearly enough that a reviewer can apply it and an auditor can check it 
  • Comfortable in spreadsheets and in a dashboarding tool, with enough Python or SQL to pull and slice your own data (you will not be asked to build interfaces) 
  • Experience training or onboarding reviewers into rubric-based work 

The expected hourly salary range for this position is $75-85 p/hour, based on experience, skills, and qualifications.

 

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams. 

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.

Company

Innodatainc
In Office, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Innodatainc's careers site·first seen 20 Sept 2026·last verified 20 Sept 2026·How we source jobs

Similar jobs

  • Software Engineers: Paid Interview on AI Evaluation Tasks at teracUnited States–match not yet calculated
  • Evaluation Associate at cityofphiladelphiaPhiladelphia, United States of America–match not yet calculated
  • Manager II, Engineering - Database Monitoring (AI) at DatadogNew York, United States of America–match not yet calculated
  • Management Consultants: Paid AI Output Evaluation at teracUnited States–match not yet calculated
  • Manager, AV Evaluation Framework at generalmotorsSunnyvale, United States of America–match not yet calculated

Browse more jobs

  • Monitoring & Evaluation Specialist jobs in United States
  • Community Health Worker jobs in United States
  • Public Health Officer jobs in United States
  • Field Officer jobs in United States
  • Monitoring & Evaluation Specialist jobs in India
  • Monitoring & Evaluation Specialist jobs in France