Jobs

Software Engineer, Site Reliability (SRE)

Sierra

Apply on Ashby
Location
San Francisco, CA
Track
AI Platform / SRE
Salary
$230K–$390K / yr
Posted
October 21, 2025
Source
Ashby

Job description

What you'll do

As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and building the foundation of reliability, observability, and scalability across Sierra’s AI-driven infrastructure. You’ll partner closely with our core engineering and product teams to ensure our systems are highly available, efficient, and built for growth.

  • Own Sierra’s observability stack—monitoring, alerting, logging, and tracing—to give engineers clear visibility into system health and performance.
  • Partner with product and platform engineers to design systems that are reliable and scalable from day one—not as an afterthought.
  • Design and implement scalable, reliable, and secure cloud infrastructure (AWS) using Terraform and modern DevOps tooling.
  • Improve the reliability and scalability of our LLM deployments, ensuring robust, performant, and cost-effective operation.
  • Lead improvements to deployment pipelines, CI/CD tooling, and incident management processes to reduce downtime and response time.
  • Define the foundation of SRE practices at Sierra, influencing culture, tooling, and best practices across the engineering org.

What you'll bring

  • 5+ years of hands-on experience in Site Reliability or Infrastructure engineering roles for complex SaaS or cloud-based systems.
  • Experience designing for availability, scalability, and reliability at both infrastructure and application layers.
  • Deep experience with Terraform, AWS services, container orchestration, and cloud networking (including IAM and VPC architecture).
  • Strong background in observability systems (e.g., Prometheus, Grafana, Datadog, or similar).
  • Experience working with enterprise customers and familiarity with their compliance and networking needs along with integration patterns.
  • Comfortable working in fast-moving environments and collaborating across product, ML, and core engineering teams.
  • Degree in Computer Science or a related field, or equivalent professional experience.

Even better...

  • Experience with LLM infrastructure — optimizing inference performance, managing fine-tuned models, or large-scale model deployment.
  • Past experience in an early-stage startup environment, especially defining SRE culture and tooling from scratch.
  • Familiarity with incident management automation or self-healing infrastructure patterns.

Similar roles

Get roles like this in your inbox

New agentic AI jobs, curated every Thursday. No spam.

Apply on Ashby