Software Engineer, Site Reliability (SRE)
- Location
- San Francisco, CA
- Track
- AI Platform / SRE
- Salary
- $230K–$390K / yr
- Posted
- October 21, 2025
- Source
- Ashby
Job description
What you'll do
As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and building the foundation of reliability, observability, and scalability across Sierra’s AI-driven infrastructure. You’ll partner closely with our core engineering and product teams to ensure our systems are highly available, efficient, and built for growth.
- Own Sierra’s observability stack—monitoring, alerting, logging, and tracing—to give engineers clear visibility into system health and performance.
- Partner with product and platform engineers to design systems that are reliable and scalable from day one—not as an afterthought.
- Design and implement scalable, reliable, and secure cloud infrastructure (AWS) using Terraform and modern DevOps tooling.
- Improve the reliability and scalability of our LLM deployments, ensuring robust, performant, and cost-effective operation.
- Lead improvements to deployment pipelines, CI/CD tooling, and incident management processes to reduce downtime and response time.
- Define the foundation of SRE practices at Sierra, influencing culture, tooling, and best practices across the engineering org.
What you'll bring
- 5+ years of hands-on experience in Site Reliability or Infrastructure engineering roles for complex SaaS or cloud-based systems.
- Experience designing for availability, scalability, and reliability at both infrastructure and application layers.
- Deep experience with Terraform, AWS services, container orchestration, and cloud networking (including IAM and VPC architecture).
- Strong background in observability systems (e.g., Prometheus, Grafana, Datadog, or similar).
- Experience working with enterprise customers and familiarity with their compliance and networking needs along with integration patterns.
- Comfortable working in fast-moving environments and collaborating across product, ML, and core engineering teams.
- Degree in Computer Science or a related field, or equivalent professional experience.
Even better...
- Experience with LLM infrastructure — optimizing inference performance, managing fine-tuned models, or large-scale model deployment.
- Past experience in an early-stage startup environment, especially defining SRE culture and tooling from scratch.
- Familiarity with incident management automation or self-healing infrastructure patterns.