Jobs

Software Engineer, Infrastructure

Sierra

Apply on Ashby
Location
San Francisco, CA
Track
AI Infrastructure
Salary
$230K–$390K / yr
Posted
June 5, 2026
Source
Ashby

Job description

What you’ll do

As a Software Engineer, Infrastructure at Sierra, you will be responsible for designing, building, and maintaining the core systems that make our AI platform possible. You’ll focus on making Sierra’s infrastructure secure, reliable, and scalable, enabling product teams to deliver with speed and confidence.

  • Ensure the reliability, scalability, and performance of our platform and LLM inference serving as we rapidly grow traffic.
  • Build and maintain cloud infrastructure using Terraform to ensure scalable, secure, and reproducible environments.
  • Create and maintain a self-serve infrastructure platform that enables the rest of engineering to deploy and operate services.
  • Own and evolve CI/CD pipelines and release management, enabling fast, reliable deployments for Sierra’s platform.
  • Architect and operate distributed systems that leverage distributed databases, retrieval systems, and ML models.
  • Develop and maintain core data serving abstractions along with authentication and security features (SSO, RBAC, authentication controls).
  • Navigate and integrate our stack with enterprise customer environments in scalable and maintainable ways.
  • Enhance observability tooling (metrics, logging, tracing) to provide deep visibility into platform health and performance.
  • Lead and participate in incident management, improving system resilience through proactive monitoring, root cause analysis, and postmortems.

What you’ll bring

  • Strong software engineering background with 5–7+ years of hands-on development experience in highly technical products.
  • A strong inclination towards building automation, tooling, and platform, along with designing maintainable systems.
  • Proven experience with cloud platforms (AWS, GCP, or Azure) and infrastructure as code (Terraform preferred).
  • Hands-on expertise in CI/CD systems, release management, and container orchestration (e.g., Docker, Kubernetes).
  • Experience with observability tools (Prometheus, Grafana, Datadog, OpenTelemetry, etc.).
  • Experience in incident response and operating distributed systems in production.
  • Degree in Computer Science or related field, or equivalent professional experience

Even better…

  • Production experience working with LLMs and machine learning models.
  • Background in distributed systems, running SaaS services at scale, and agentic architecture.
  • Familiarity with security and authentication protocols (OAuth, SSO, mTLS).
  • Previous experience in a fast-paced startup environment or platform/infra-focused team.

Similar roles

Get roles like this in your inbox

New agentic AI jobs, curated every Thursday. No spam.

Apply on Ashby