NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Data Annotation Specialist in India
4 days ago
Apply with autofill
Apply with autofill
Weekday AI·4 days ago
4 days ago

PDF Annotation & Transcription Experts - Malayalam

IndiaContractRemoteMid · 2-5 yearsData Annotation Specialist

Sign up free to see how well your resume matches this role.

Boost your chances at Weekday AI

How you compare FREE

?
Your scoreYour score: not yet known
→
60
Top 10%Top 10%: 60 out of 100

Top 10% of NextRaise users matched against Data Annotation Specialist roles in India.

Must-have skills for this role

  • malayalam
  • annotation
  • transcription

PDF or DOCX · no account needed

Apply faster with autofill FREEWeekday AI uses Workable - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Open and check a task: pages are provided, so you do not source documents yourself. We find the PDFs and upload them for you. Before annotating, confirm the page is in Malayalam, is legible, has real content, and shows no personal details
  • Annotate structure: identify and bound every meaningful region of the page - document title, section heading, paragraph, list, table, figure, diagram, caption, formula, question, answer field - and assign each a component type and a reading-order index
  • Record relationships: link each region to the figure or table it belongs to through a parent component identifier
  • Transcribe faithfully: reproduce all text exactly as it appears in Malayalam script, including handwritten content, flagging any region where the source is not legible
  • Capture page metadata: language, document type, source, page dimensions, and flags for tables, formulas and handwriting
  • Review a colleague's work: every task is reviewed end to end by a second Malayalam expert, and experienced annotators take on that review

What they're looking for

  • You are a native Malayalam speaker with full command of the script, its diacritics and its conjunct forms
  • You have worked with documents: annotation, transcription, translation, localization, subtitling, proofreading, journalism, or regional-language data review
  • You are exact: character-level accuracy matters more here than speed, and a single wrong diacritic is a defect
  • You are systematic: you apply a taxonomy consistently across hundreds of pages rather than improvising per document
  • You are comfortable with unfamiliar layouts: multi-column newspapers, exam papers, handwritten forms

Nice to have

  • Regional-language AI data: annotation, labeling, grading, or bilingual evaluation for training datasets
  • Transcription and localization: MTPE, subtitling, bilingual QA, OCR correction or post-editing
  • Document production: typesetting, copy-editing, proofreading, or digitization of Malayalam-language material
  • Script and encoding: Unicode normalization, Malayalam input methods, numeral-form and character-form accuracy

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

This role is for one of our clients

Compensation: $12.68 per hour

Fluent Language Skills Required: Malayalam. Native fluency in Malayalam, including full command of Malayalam script and orthography, is required for this position. All annotation and transcription work is performed in Malayalam.

Why This Role Exists

Document understanding breaks down fastest in the languages that parsing and vision-language models rarely see. This project builds training data for exactly those languages: Malayalam, alongside four other Indic scripts, Japanese and Korean. Each task takes a real, publicly available PDF page and produces a complete structural map of that page, paired with a faithful transcription of every text region in the original script.

The dataset deliberately concentrates on the material models handle worst: handwriting, dense multi-column layouts, tables, diagrams, and mixed-script pages. Documents are drawn from newspapers, textbooks, examinations, and everyday formats such as flyers, forms, manuals, menus, brochures, notices and worksheets, so that the corpus reflects the real diversity of Malayalam documents rather than a narrow band of easily parsed ones.

Delivered work is human-authored throughout. Component identification, component typing, reading order and all transcription are performed by people, not generated by parsing models.

Requirements

What You'll Do

  • Open and check a task: pages are provided, so you do not source documents yourself. We find the PDFs and upload them for you. Before annotating, confirm the page is in Malayalam, is legible, has real content, and shows no personal details
  • Annotate structure: identify and bound every meaningful region of the page - document title, section heading, paragraph, list, table, figure, diagram, caption, formula, question, answer field - and assign each a component type and a reading-order index
  • Record relationships: link each region to the figure or table it belongs to through a parent component identifier
  • Transcribe faithfully: reproduce all text exactly as it appears in Malayalam script, including handwritten content, flagging any region where the source is not legible
  • Capture page metadata: language, document type, source, page dimensions, and flags for tables, formulas and handwriting
  • Review a colleague's work: every task is reviewed end to end by a second Malayalam expert, and experienced annotators take on that review

Who You Are

  • You are a native Malayalam speaker with full command of the script, its diacritics and its conjunct forms
  • You have worked with documents: annotation, transcription, translation, localization, subtitling, proofreading, journalism, or regional-language data review
  • You are exact: character-level accuracy matters more here than speed, and a single wrong diacritic is a defect
  • You are systematic: you apply a taxonomy consistently across hundreds of pages rather than improvising per document
  • You are comfortable with unfamiliar layouts: multi-column newspapers, exam papers, handwritten forms

Nice-to-Have Specialties

  • Regional-language AI data: annotation, labeling, grading, or bilingual evaluation for training datasets
  • Transcription and localization: MTPE, subtitling, bilingual QA, OCR correction or post-editing
  • Document production: typesetting, copy-editing, proofreading, or digitization of Malayalam-language material
  • Script and encoding: Unicode normalization, Malayalam input methods, numeral-form and character-form accuracy

What Success Looks Like

  • Every meaningful region on the page is captured, correctly bounded and correctly typed
  • Reading order reflects how the page is actually read, including across columns
  • Transcriptions match the source character for character, in Malayalam script rather than transliteration
  • Your tasks pass second-expert review the first time
  • Unsuitable pages are flagged up front rather than after thirty minutes of work

Company

Weekday AI
India

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Weekday AI's careers site·first seen 16 Sept 2026·last verified 16 Sept 2026·How we source jobs

Similar jobs

  • PDF Annotation & Transcription Experts - Telugu at Weekday AIIndia–match not yet calculated
  • PDF Annotation & Transcription Experts - Bengali at Weekday AIIndia–match not yet calculated
  • PDF Annotation & Transcription Experts - Gujarati at Weekday AIIndia–match not yet calculated
  • PDF Annotation & Transcription Experts - Odiya at Weekday AIIndia–match not yet calculated
  • Regulatory Labeling Manager at AmgenHyderabad, India–match not yet calculated

Browse more jobs

  • Data Annotation Specialist jobs in India
  • AI Engineer jobs in India
  • Machine Learning Engineer jobs in India
  • AI / ML Researcher jobs in India
  • Data Annotation Specialist jobs in United States
  • Data Annotation Specialist jobs in France