Senior Lead Software Engineer - AI/ML Developer
AI Platform / SRE
39 open positions across the agentic economy.
An AI Platform or Site Reliability Engineer builds and runs the internal platform that AI product teams ship on: Kubernetes and cloud infrastructure, CI/CD, observability, capacity and cost management, and the reliability engineering that keeps inference-backed products inside their SLOs. The role differs from AI Infrastructure in altitude: infrastructure roles go deep on GPUs, training clusters, and serving runtimes; platform and SRE roles own the paved road every team at the company uses. Hiring is broad across AI-native companies of every size. Strong candidates bring production SRE experience plus curiosity about the failure modes stochastic systems add.
As of September 2026, AgenticCareers tracks 39 open AI Platform / SRE positions across 28 companies. The median advertised salary is $193K–$265K per year (based on 25 listings with disclosed pay). 18% of roles are remote. The most active employers are Harvey, Anthropic, Ashby.
Full AI Platform / SRE salary dataAI Platform / SRE
Software
AI Platform / SRE
AI Platform / SRE
Locations: New York, New York ## Job description About this role Your team The Aladdin Data Platform Engineering team builds and operates the foundational data and developer platforms that power the Aladdin ecosystem. Our portfolio includes the Enterprise Data Platform (EDP), Data Services and supporting infrastructure, and Aladdin Studio, the builder ecosystem for Aladdin. Together, these platforms enable the Aladdin community to acquire, manage, govern, process, and di
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE · Executive
AI Platform / SRE
AI · ai agent
AI Platform / SRE
AI · llm · ai agent
rag
mlops
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Platform / SRE
AI Agent
AI Platform / SRE
An AI Platform or Site Reliability Engineer builds and runs the internal platform that AI product teams ship on: Kubernetes and cloud infrastructure, CI/CD, observability, capacity and cost management, and the reliability engineering that keeps inference-backed products inside their SLOs. The role differs from AI Infrastructure in altitude: infrastructure roles go deep on GPUs, training clusters, and serving runtimes; platform and SRE roles own the paved road every team at the company uses. Hiring is broad across AI-native companies of every size. Strong candidates bring production SRE experience plus curiosity about the failure modes stochastic systems add.
As of September 2026, AgenticCareers tracks 39 open AI Platform / SRE positions across 28 companies, updated daily.
The median advertised AI Platform / SRE salary is $193K–$265K per year, based on 25 listings with disclosed pay.
18% of the AI Platform / SRE roles currently tracked on AgenticCareers are remote.
The most active employers hiring AI Platform and SRE engineers right now are Harvey, Anthropic, Ashby.
How to Deploy AI Agents in Production
Deploying AI agents in production surfaces a unique set of challenges that most tutorials skip entirely: this guide addresses reliability, cost, latency, and observability head-on.
From DevOps to AI Agent Engineering: A Transition Guide
Your infrastructure, reliability, and observability skills transfer directly. Here is the fastest path from DevOps to AI agent engineering in 2026.
The AI Agent Observability Stack: 5 Platforms Compared
You cannot improve what you cannot measure. This guide compares the five leading AI agent observability platforms across tracing, evaluation, cost tracking, and production monitoring for 2026.
1,700+ curated AI and agentic jobs from top companies
Curated every Thursday. No spam.