NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Monitoring & Evaluation Specialist in United States of America
1 day agoBe an early applicant
Apply with autofill
Apply with autofill
Terac·1 day ago
1 day agoBe an early applicant

Software Engineers: Paid Interview on AI Evaluation Tasks

United StatesContractRemoteMid · 2-5 yearsMonitoring & Evaluation Specialist

Sign up free to see how well your resume matches this role.

Boost your chances at terac

How you compare FREE

?
Your scoreYour score: not yet known
→
17
Top 10%Top 10%: 17 out of 100

Top 10% of NextRaise users matched against Monitoring & Evaluation Specialist roles in United States.

Must-have skills for this role

  • code review
  • test automation
  • systems architecture
  • technical architecture

PDF or DOCX · no account needed

Apply faster with autofill FREEterac uses Ashby - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Review and assess the quality of realistic programming tasks
  • Evaluate coding environments and technical harnesses for AI testing
  • Walk us through potential flaws in the provided code structures
  • Provide feedback on the difficulty and realism of the challenges

What they're looking for

  • Professional experience as a software engineer or developer
  • Familiarity with building or verifying programming tasks
  • Experience with code review and evaluation harnesses
  • Comfortable discussing technical architecture and testing methodologies

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

What We're Researching

We're running a paid study on the quality and realism of coding environments designed to test AI agents. We are developing a comprehensive suite of programming tasks and evaluation harnesses to measure AI performance accurately. Your feedback directly shapes how these agents are benchmarked against real-world engineering standards.

How It Works

During this remote session, you will review several programming tasks and their corresponding evaluation harnesses. You will assess the technical accuracy, complexity, and realism of each coding challenge. We will ask you to walk through the logic of the evaluation environments and identify any potential flaws. Finally, you will provide feedback on how these environments compare to standard industry practices.

Who This Is For

We are looking for software engineers who have hands-on experience building, reviewing, or testing realistic programming tasks. We welcome full-stack developers, backend engineers, test automation engineers, and systems architects who are familiar with evaluation harnesses. Ideal candidates understand what makes a coding challenge robust, verifiable, and technically sound.

What You'll Do

  • Review and assess the quality of realistic programming tasks

  • Evaluate coding environments and technical harnesses for AI testing

  • Walk us through potential flaws in the provided code structures

  • Provide feedback on the difficulty and realism of the challenges

Who Should Apply

  • Professional experience as a software engineer or developer

  • Familiarity with building or verifying programming tasks

  • Experience with code review and evaluation harnesses

  • Comfortable discussing technical architecture and testing methodologies

Compensation

$75 per hour

 

Ready to participate?

Start your paid interview now

 

About Terac

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

 

Learn more at terac.com or on YouTube at @jointerac.

Company

Terac
United States

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Terac's careers site·first seen 20 Sept 2026·last verified 20 Sept 2026·How we source jobs

Similar jobs

  • Quality Lead, Agentic AI Workflow Evaluation at innodataincIn Office, United States of America–match not yet calculated
  • Evaluation Associate at cityofphiladelphiaPhiladelphia, United States of America–match not yet calculated
  • Manager II, Engineering - Database Monitoring (AI) at DatadogNew York, United States of America–match not yet calculated
  • Management Consultants: Paid AI Output Evaluation at teracUnited States–match not yet calculated
  • Manager, AV Evaluation Framework at generalmotorsSunnyvale, United States of America–match not yet calculated

Browse more jobs

  • Monitoring & Evaluation Specialist jobs in United States
  • Community Health Worker jobs in United States
  • Public Health Officer jobs in United States
  • Field Officer jobs in United States
  • Monitoring & Evaluation Specialist jobs in India
  • Monitoring & Evaluation Specialist jobs in France