Software Engineer I - AI Observability
- Location
- Portugal and 2 more locations
- Track
- AI Ops / DevOps
- Level
- Junior
- Posted
- October 7, 2026
- Source
- Greenhouse
Job description
What is The Role
Elastic Observability is building a system that monitors production environments the way an expert SRE would: it learns what normal looks like across logs, metrics, and traces, generates its own detection queries, identifies when something meaningful changes, and launches autonomous investigations — without waiting for a human to notice first.
The goal is to close the gap between the alert that fires and the answer that matters, at a scale no on-call rotation can match. You will work on the Kibana UI, where
Autonomous investigations meet the engineers who rely on them, making agent reasoning legible enough that an SRE can trust it, challenge it, and act on it.
You will also work on the context and detection pipelines that turn raw telemetry into structured signals, the analysis that separates noise from events worth investigating, and the multi-agent system that digs into root causes autonomously and drives remediation.
What You Will Be Doing
You will work across our AI SRE solution. Over time, that will include the following
- User-facing surfaces: Build the Kibana interfaces where signals, investigations, and findings reach users: the views an SRE uses to follow an investigation, inspect the evidence behind a conclusion, and drive the remediation that follows. This is most of the work today, though investigations increasingly need to reach engineers wherever they already work
- Agentic investigation: Contribute to the multi-agent system that pursues root cause autonomously: build the tools agents use to investigate, shape the context they reason over, and help decide when an agent has gathered enough evidence to conclude
- Detection and signal generation: Contribute to the systems that establish baselines across high-cardinality telemetry, generate their own detection queries, and identify state changes worth investigating
- Evaluation infrastructure: Agent output quality is the product here. Help build offline and online evaluations, and contribute to dataset construction and curation
- Cross-functional collaboration: Partner with UX designers on how autonomous investigations are presented, product managers on scope and sequencing, and data scientists on analysis methods
What You Bring
- At least two years developing applications with TypeScript, React and Node.js
- Using AI to accelerate development, debug systems, and optimize code, while still owning the outcomes
- Knowledge of LLM-based systems — agent orchestration, tool use, retrieval, context management
- A bachelor's degree in Computer Science (or a related technical field), or equivalent hands-on experience
- A creative individual with an enthusiastic interest in technology and software engineering
- A willingness to learn in a dynamic and distributed environment
Bonus Points
- Knowledge of self-directed agentic execution (multi-turn reasoning loops) and evaluating output quality
- Evaluating non-deterministic systems by designing evaluation datasets, offline and online evaluations
- Knowledge of Observability and Elasticsearch
- Experience working in distributed teams
- Experience in open source
The typical starting salary range for this role is
€38.700 — €61.300 EUR