Jobs

Machine Learning Engineer – Evals

Mango Global Technologies

Apply on the company site
Location
New York, NY, US
Track
ML Engineer
Salary
$220K–$300K / yr
Posted
October 8, 2026
Source
Indeed

Job description

Machine Learning Engineer – Evals

Location: New York, NY

Work arrangement: On-site, 5 days/week

Employment: Full-time

Compensation: $220,000–$300,000 base + competitive equity

Openings: 1

Visa sponsorship: Not available — candidates must already have unrestricted authorization to work in the United States.

About the Role

We are looking for a Machine Learning Engineer – Evals to own the evaluation of a sophisticated AI memory and identity system end to end.

You will define how we determine whether evolving entity representations are actually improving, build the infrastructure required to measure those improvements, and use evaluation results to identify and fix failures.

This is not a role for someone who simply runs established benchmarks. You will be expected to design novel evaluations from first principles, work with imperfect or absent ground truth, analyze model and agent traces, and turn research questions into production-quality evaluation systems.

The ideal candidate combines strong ML engineering skills with genuine research ability and enjoys working in an ambiguous, fast-moving environment.

What You’ll Do

Design evaluation systems

Define what “better” means for representations that evolve over time.

Design novel evaluation methodologies where traditional ground truth is unavailable or insufficient.

Develop scoring frameworks, metrics, rubrics, judge models, and diagnostic methodologies.

Continuously evolve evaluation definitions as models and product capabilities change.

Build evaluation infrastructure

Build production-quality Python pipelines and evaluation harnesses.

Manage datasets, labels, experiment versions, model versions, judges, reruns, and regression testing.

Create infrastructure that allows researchers and engineers to answer new evaluation questions quickly.

Integrate experiment tracking and observability into evaluation workflows.

Analyze results and improve the system

Analyze model outputs, agent trajectories, and execution traces.

Identify meaningful failure modes rather than optimizing for superficial metrics.

Translate evaluation findings into concrete engineering or research improvements.

Run experiments, compare results, document findings, and drive the resulting fixes.

Build simulation agents at scale

Develop and operate simulation environments capable of modeling large numbers of entities.

Increase the throughput and reliability of evaluation runs.

Build systems that enable rapid iteration across models, prompts, representations, and evaluation strategies.

Own the full evaluation loop

You will own the process from

Research question → evaluation design → data → implementation → execution → analysis → diagnosis → rerun → recommendation

There is no expectation that this work will be handed off between multiple teams.

Required Qualifications

Dealbreakers

3+ years of professional experience designing and owning LLM evaluations.

Demonstrated experience designing, building, and running LLM evaluations end to end.

Production-quality Python experience, including testing, maintainability, debugging, and release practices.

Experience creating evaluation methodologies rather than exclusively running existing benchmarks.

Strong ability to investigate model/agent failures using traces, outputs, datasets, and experimental results.

Ability to work on-site in New York City five days per week.

Must already have unrestricted U.S. work authorization; visa sponsorship is not available.

Required Technical Skills

Python

PyTorch

Hugging Face / Transformers

LLM evaluation and benchmarking

Experiment design and statistical/quantitative analysis

Evaluation datasets and labeling methodologies

Model/LLM-as-a-judge evaluation

ML experimentation and debugging

Preferred Technical Skills

Weights & Biases

Hydra

SQL

Kafka or other streaming/event-driven infrastructure

Evaluation/observability frameworks

Agent simulation and multi-agent systems

Large-scale data pipelines

Nice to Have

Experience with AI memory, identity, personalization, agent evaluation, or agentic optimization.

Research experience in LLMs, NLP, machine learning, evaluation, or related areas.

First-author/co-author research at NeurIPS, ICML, ICLR, ACL, EMNLP, or comparable venues.

Alternatively, a high-quality open-source project, benchmark, framework, or technical publication demonstrating comparable research ability.

Degree in Computer Science, Mathematics, Statistics, Machine Learning, or another quantitative discipline.

Who Will Not Be a Strong Fit

Candidates who have only run established benchmarks without designing their own evaluations.

Engineers whose LLM experience is limited to application development without evaluation ownership.

Candidates without substantial hands-on Python engineering experience.

Candidates without evidence of research thinking, novel experimentation, published work, or comparable open-source work.

Candidates unwilling or unable to work five days per week from the New York City office.

Compensation

$220,000–$300,000 base salary + competitive equity

Final compensation will depend on experience, technical depth, research contributions, and interview performance.

Hiring Process

The successful candidate should expect technical discussions focused on

Designing an evaluation from scratch.

Diagnosing an LLM/agent failure from traces.

Building a production evaluation pipeline.

Research methodology and experimental design.

Python/ML engineering.

Previous research or open-source work.

Pay: $220,000.00 - $300,000.00 per year

Benefits

  • Employee assistance program

Application Question(s)

  • Can you work full-time on-site in New York City 5 days per week, and do you currently have unrestricted U.S. work authorization without requiring visa sponsorship?
  • Have you either (a) published ML/LLM research at a recognized conference/journal or (b) released substantial open-source research/ML work of comparable quality?
  • Do you have hands-on experience with PyTorch AND Hugging Face/Transformers?
  • Do you have

Similar roles

Get roles like this in your inbox

New agentic AI jobs, curated every Thursday. No spam.

Apply on the company site