Machine Learning Engineer – Evals
Mango Global Technologies
- Location
- New York, NY, US
- Track
- ML Engineer
- Salary
- $220K–$300K / yr
- Posted
- October 8, 2026
- Source
- Indeed
Job description
Machine Learning Engineer – Evals
Location: New York, NY
Work arrangement: On-site, 5 days/week
Employment: Full-time
Compensation: $220,000–$300,000 base + competitive equity
Openings: 1
Visa sponsorship: Not available — candidates must already have unrestricted authorization to work in the United States.
About the Role
We are looking for a Machine Learning Engineer – Evals to own the evaluation of a sophisticated AI memory and identity system end to end.
You will define how we determine whether evolving entity representations are actually improving, build the infrastructure required to measure those improvements, and use evaluation results to identify and fix failures.
This is not a role for someone who simply runs established benchmarks. You will be expected to design novel evaluations from first principles, work with imperfect or absent ground truth, analyze model and agent traces, and turn research questions into production-quality evaluation systems.
The ideal candidate combines strong ML engineering skills with genuine research ability and enjoys working in an ambiguous, fast-moving environment.
What You’ll Do
Design evaluation systems
Define what “better” means for representations that evolve over time.
Design novel evaluation methodologies where traditional ground truth is unavailable or insufficient.
Develop scoring frameworks, metrics, rubrics, judge models, and diagnostic methodologies.
Continuously evolve evaluation definitions as models and product capabilities change.
Build evaluation infrastructure
Build production-quality Python pipelines and evaluation harnesses.
Manage datasets, labels, experiment versions, model versions, judges, reruns, and regression testing.
Create infrastructure that allows researchers and engineers to answer new evaluation questions quickly.
Integrate experiment tracking and observability into evaluation workflows.
Analyze results and improve the system
Analyze model outputs, agent trajectories, and execution traces.
Identify meaningful failure modes rather than optimizing for superficial metrics.
Translate evaluation findings into concrete engineering or research improvements.
Run experiments, compare results, document findings, and drive the resulting fixes.
Build simulation agents at scale
Develop and operate simulation environments capable of modeling large numbers of entities.
Increase the throughput and reliability of evaluation runs.
Build systems that enable rapid iteration across models, prompts, representations, and evaluation strategies.
Own the full evaluation loop
You will own the process from
Research question → evaluation design → data → implementation → execution → analysis → diagnosis → rerun → recommendation
There is no expectation that this work will be handed off between multiple teams.
Required Qualifications
Dealbreakers
3+ years of professional experience designing and owning LLM evaluations.
Demonstrated experience designing, building, and running LLM evaluations end to end.
Production-quality Python experience, including testing, maintainability, debugging, and release practices.
Experience creating evaluation methodologies rather than exclusively running existing benchmarks.
Strong ability to investigate model/agent failures using traces, outputs, datasets, and experimental results.
Ability to work on-site in New York City five days per week.
Must already have unrestricted U.S. work authorization; visa sponsorship is not available.
Required Technical Skills
Python
PyTorch
Hugging Face / Transformers
LLM evaluation and benchmarking
Experiment design and statistical/quantitative analysis
Evaluation datasets and labeling methodologies
Model/LLM-as-a-judge evaluation
ML experimentation and debugging
Preferred Technical Skills
Weights & Biases
Hydra
SQL
Kafka or other streaming/event-driven infrastructure
Evaluation/observability frameworks
Agent simulation and multi-agent systems
Large-scale data pipelines
Nice to Have
Experience with AI memory, identity, personalization, agent evaluation, or agentic optimization.
Research experience in LLMs, NLP, machine learning, evaluation, or related areas.
First-author/co-author research at NeurIPS, ICML, ICLR, ACL, EMNLP, or comparable venues.
Alternatively, a high-quality open-source project, benchmark, framework, or technical publication demonstrating comparable research ability.
Degree in Computer Science, Mathematics, Statistics, Machine Learning, or another quantitative discipline.
Who Will Not Be a Strong Fit
Candidates who have only run established benchmarks without designing their own evaluations.
Engineers whose LLM experience is limited to application development without evaluation ownership.
Candidates without substantial hands-on Python engineering experience.
Candidates without evidence of research thinking, novel experimentation, published work, or comparable open-source work.
Candidates unwilling or unable to work five days per week from the New York City office.
Compensation
$220,000–$300,000 base salary + competitive equity
Final compensation will depend on experience, technical depth, research contributions, and interview performance.
Hiring Process
The successful candidate should expect technical discussions focused on
Designing an evaluation from scratch.
Diagnosing an LLM/agent failure from traces.
Building a production evaluation pipeline.
Research methodology and experimental design.
Python/ML engineering.
Previous research or open-source work.
Pay: $220,000.00 - $300,000.00 per year
Benefits
- Employee assistance program
Application Question(s)
- Can you work full-time on-site in New York City 5 days per week, and do you currently have unrestricted U.S. work authorization without requiring visa sponsorship?
- Have you either (a) published ML/LLM research at a recognized conference/journal or (b) released substantial open-source research/ML work of comparable quality?
- Do you have hands-on experience with PyTorch AND Hugging Face/Transformers?
- Do you have