LLM Engineering Expert – AI Evaluation
Unknown
- Location
- United States
- Track
- AI Agent Engineer
- Level
- Senior
- Posted
- October 7, 2026
- Source
- Indeed
Job description
About the role
Vraify is hiring experienced LLM Engineering Experts to create and validate challenging, simulation-based engineering design problems for evaluating advanced AI agents. This remote contractor opportunity is open only to candidates residing in the United States or Canada.
You will design multi-constraint tasks, configure open-source simulation tools, analyze agent execution logs, diagnose reasoning and tool-use failures, and build objective automated graders across electrical, mechanical, aerospace, control systems, systems engineering, and robotics.
Important experience note
This is a senior specialist role requiring at least 8 years of directly relevant engineering experience. To help candidates focus on opportunities aligned with their background, applications that do not meet this minimum will be automatically screened out. Freshers and early-career applicants are therefore not eligible for this role, and we warmly encourage them to consider opportunities better suited to their current experience level.
Responsibilities
- Author original, self-contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective autograders.
- Build, run, and validate problem environments using open-source simulation tools and custom Python test benches.
- Evaluate coding-agent outputs and execution logs across repeated trials to identify systemic failure modes.
- Refine task difficulty using empirical model-performance data without introducing ambiguity or missing information.
- Collaborate with AI researchers, pod leads, and domain experts to integrate rigorous benchmarks into the model-evaluation pipeline.
Required qualifications
- Master’s degree or PhD in Electrical Engineering, Mechanical Engineering, Aerospace Engineering, or a closely related field.
- 8+ years of hands-on engineering design experience.
- Proficiency with at least one relevant open-source simulation package, such as ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, build123d, OpenModelica, Cantera, or Gmsh.
- Strong Python scripting skills.
- Hands-on experience with modern LLMs or coding agents and evaluation concepts such as pass@k, failure-mode analysis, and nondeterministic behavior.
- Ability to audit trajectory logs and isolate core reasoning or tool-use failures.
- Strong attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
- Stable high-speed internet and a personal desktop or laptop suitable for remote work.
Engagement details
- Remote contractor engagement for up to 24 weeks, starting immediately.
- Expected commitment: 40 hours per week with at least 4 hours of overlap with Pacific Time.
- Weekend on-call availability required; part-time engagement may be considered.
- Compensation: USD $500 per completed task; each task is expected to require approximately 30–40 hours.
- Shortlisted candidates will complete a calibration and delivery review process.
Eligible locations
United States and Canada only.
Work Location: Remote