Jobs

LLM Model Response Evaluation

Lifted (an Upwork Company)

Apply on Himalayas
Location
Remote (United States)
Track
General AI
Posted
October 1, 2026
Source
Himalayas

Job description

  • Evaluating UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities.
  • The work may involve text, images, audio, video, HTML widgets, PDFs, or combinations of these modalities.
  • The work is domain-agnostic and may cover topics across arts, culture, history, science, engineering, and more.
  • Resources will be expected to independently research unfamiliar topics using trusted sources before making evaluation decisions.
  • Each task will include detailed project guidelines within the evaluation platform.
  • 3+ years of hands-on experience in LLM / GenAI data evaluation.
  • Bachelor's Degree required
  • Ability to research unfamiliar topics using trusted sources and make well-supported judgments.
  • Comfortable evaluating content across multiple modalities

Flexible and remote work

Variable workload: Accept or decline tasks based on your availability

No guaranteed hours: Workload may vary weekly

Our client, a global technology company that helps businesses build, train, and manage AI systems is looking for experts to evaluate model-generated content against defined quality rubrics such as factuality, consistency, aesthetics, and other evaluation criteria.

Originally posted on Himalayas

Skills

  • AI Evaluator
  • LLM Evaluator
  • AI Trainer
  • Content Evaluator
  • AI Annotation Specialist
  • AI Response Evaluation
  • LLM Evaluation
  • AI LLM Evaluation
  • Language Model Evaluation
  • AI Language Model Evaluation
  • AI model evaluation
  • AI ML Model Evaluation
  • AI Response Evaluator
  • AI Model Assessment

Similar roles

Get roles like this in your inbox

New agentic AI jobs, curated every Thursday. No spam.

Apply on Himalayas