Senior Forward Deployed Engineer (DevOps)
- Location
- Hybrid
- Level
- Senior
- Salary
- $194K–$266K / yr
- Posted
- October 7, 2026
- Source
- Greenhouse
Job description
About Us
Available Locations
Hybrid | Austin, San Francisco
About the Role
Cloudflare's Senior Forward Deployed Infrastructure Engineers work where Cloudflare's infrastructure meets customer impact.
As a Forward Deployed Infrastructure Engineer, you will be embedded with one of Cloudflare's most strategic customers, collaborating to ensure modern cloud compute capacity is available, reliable, and ready when needed. You will guide new capacity from delivery to production, partnering with Cloudflare's data center and hardware teams to ensure systems pass validation, come online smoothly, and uphold Cloudflare's high standards.
You will serve as a bridge connecting Cloudflare's infrastructure, hardware, and data center teams with the customer's engineering organization. You'll ensure teams work from a shared vision, clear operational roadblocks, and feed real-world insights back into how Cloudflare builds and evolves its platform.
This role is ideal for engineers who want to
- Stay deeply technical and hands-on with real production infrastructure.
- Guide outcomes from hardware delivery through running workloads.
- Shape how Cloudflare delivers compute capacity at scale.
- Build long-term, trust-based partnerships with world-class engineering teams.
Responsibilities
- Capacity Onboarding & Delivery: Lead bringing new compute capacity online for the customer, collaborating on planning, provisioning, validation, and handover to production. Drive turn-ups that are predictable, efficient, and repeatable.
- Hardware & Spec Validation: Partner with teams to ensure delivered systems meet agreed specifications. Define acceptance criteria and validation workflows, including burn-in, benchmarking, and health checks across compute, storage, and networking. Work alongside on-site data center teams to execute checks and resolve issues.
- Standards & Operational Readiness: Champion Cloudflare's build, configuration, security, and observability standards prior to capacity launch. Ensure comprehensive monitoring, alerting, and runbooks are established for all production systems.
- Cross-Team Alignment: Act as the central point of collaboration between Cloudflare's infrastructure, hardware engineering, data center operations, supply chain, and network teams, alongside the customer's engineering group.
- Technical Accountability: Serve as the primary technical point of contact for the customer's infrastructure, guiding account strategy and technical direction for supporting Cloudflare resources.
- Reliability & Incident Response: Engage in operational rhythms with the customer and Cloudflare, including standups, capacity reviews, change management, and incident response. Facilitate blameless root-cause analysis and post-incident improvements
.
- Automation & Tooling: Develop production-quality automation and tooling for provisioning, validation, and fleet management to continually improve rollout efficiency.
- Capacity Planning & Forecasting: Partner with the customer to understand future demand, forecast capacity needs, and proactively identify constraints.
- Feedback Loop & Product Influence: Surface real-world operational issues and platform opportunities with Cloudflare's Infrastructure, Hardware, and Product teams to inform future roadmaps.
- Strategic Relationship Management: Cultivate trusted technical relationships with senior stakeholders (Staff+ engineers, Directors, and VPs) across both organizations.
- On-Site Presence: Spend time regularly on-site at customer offices to foster close communication and partnership with their engineering teams.
Desirable Skills, Knowledge, and Experience
At Cloudflare, we know that great candidates come from diverse backgrounds with non-linear paths. You don’t need to tick every single box to be right for this role. If you are excited about building a better Internet and ready to make an impact, please apply.
- Infrastructure & SRE Background: Solid experience in infrastructure engineering, SRE, or production engineering, with a track record of running and scaling large-scale production systems.
- Capacity & Fleet Operations: Practical experience bringing server capacity online, encompassing provisioning, imaging, configuration management, and fleet-scale lifecycle management.
- Hardware Fluency: Familiarity with server hardware, firmware, BMC/IPMI/Redfish, and diagnostics. Ability to set validation standards, triage issues remotely, and partner with on-site technicians. Experience with GPUs or specialized accelerators is a plus.
- Linux & Networking: Strong understanding of Linux systems and core networking principles (TCP/IP, BGP, DNS, load balancing, and data center topologies).
- Automation & Development: Proficiency in languages such as Go, Python, or Rust, along with experience using infrastructure-as-code and configuration management tools (e.g., Terraform, Ansible, Salt).
- Production Ownership: Experience supporting mission-critical services, including on-call rotations, incident response, SLO management, and reliability engineering.
- Observability: Familiarity with monitoring, metrics, and logging frameworks (e.g., Prometheus, Grafana, ClickHouse) to maintain system health proactively.
- Containers & Orchestration: Exposure to Kubernetes or similar platforms and workload scheduling on newly provisioned capacity.
- AI-Augmented Workflows: Interest or experience in leveraging modern AI tools to streamline automation, debugging, and documentation.
- Program-Level Coordination: Demonstrated ability to coordinate complex, multi-team delivery efforts across regions and time zones with clear communication and alignment.
- Proactive Problem-Solving: A self-motivated approach to identifying bottlenecks—whether technical or procedural—and collaborating with partners to implement durable fixes.
- Stakeholder Communication: Ability to engage comfortably in strategic technical discussions with executive-level stakeholders, translating infrastructure context into bus