LLM Engineering Expert
Unknown
- Location
- Remote
- Track
- Prompt Engineer
- Posted
- October 6, 2026
- Source
- Indeed
Job description
LLM Engineering Expert
Location: United States — Remote
Employment type: Contract
Duration: Up to 24 weeks
Start date: Immediately
Availability: 40 hours per week with 4 hours of overlap with Pacific Standard Time (PST)
Compensation: $500 per completed task
Openings: 100
About the role
We are seeking experienced engineering professionals to create and validate challenging, simulation-based engineering design problems that help train and evaluate advanced AI agents. The work spans Electrical, Mechanical, Control Systems, Aerospace, Systems, and Robotics engineering.
You will develop complex tasks that require AI agents to interpret requirements, navigate design trade-offs, use open-source simulation tools, diagnose failures, and iterate toward valid solutions. You will also analyze execution logs and build automated, objective graders to improve model performance.
Responsibilities
- Author original, self-contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective automated graders.
- Build, run, and validate simulation environments using open-source tools and custom Python test benches.
- Analyze coding-agent outputs and execution logs across repeated trials to identify reasoning and tool-use failures, including misinterpreted simulator feedback, premature convergence, and physically impossible designs.
- Refine task difficulty based on observed model performance while keeping requirements clear and complete.
- Collaborate with AI researchers and engineering domain experts to integrate rigorous benchmarks into model evaluation workflows.
Required qualifications
- Bachelor’s degree or equivalent practical experience.
- 10+ years of hands-on engineering experience.
- Proficiency with at least one domain-relevant open-source simulation package, such as ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, build123d, OpenModelica, Cantera, or Gmsh.
- Strong Python scripting skills.
- Hands-on experience with modern large language models or coding agents and evaluation concepts such as pass@k, failure-mode analysis, and nondeterministic behavior.
- Ability to audit execution logs and isolate reasoning and tool-use failures.
- Strong attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
- Availability for 40 hours per week, including 4 hours of PST overlap and weekend on-call availability.
- A personal desktop or laptop with a stable, high-speed internet connection.
Preferred qualifications
- Master’s degree or PhD in Electrical, Mechanical, or Aerospace Engineering.
- Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role.
Engagement details
- Fully remote contractor assignment based in the United States.
- Contract duration of up to 24 weeks, with an immediate start.
- Compensation of $500 per completed task. Tasks are expected to take approximately 30–40 hours to complete.
Work Location: Remote