What You'll Do
Own the company's cloud and infrastructure cost optimization strategy across both production and R&D environments.
Monitor and optimize spending across AWS, AI providers, Datadog, Cloudflare, and other infrastructure vendors.
Build dashboards, alerts, and automated monitoring to provide real-time visibility into infrastructure costs and usage.
Partner closely with the Data team to analyze cost drivers, identify optimization opportunities, and improve resource efficiency.
Work with Engineering, DevOps, and Research to design cost-efficient architectures without compromising reliability or developer productivity.
Drive FinOps best practices, including budgeting, forecasting, cost allocation, capacity planning, and infrastructure optimization.
Build internal tools and automation that help engineers understand and optimize the cost of the systems they build.
What We're Looking For
4+ years of experience in DevOps, Infrastructure, Site Reliability Engineering, FinOps, or a similar technical role.
Strong hands-on DevOps background with experience managing production cloud infrastructure.
Experience working in a fast-growing startup, with the ability to operate in a dynamic, fast-paced environment.
Deep understanding of AWS and modern cloud-native infrastructure, including Kubernetes, Docker, Terraform, and CI/CD.
Experience optimizing cloud costs and working with infrastructure platforms such as Datadog, Cloudflare, and AI providers.
Strong analytical and problem-solving skills, with the ability to turn infrastructure data into actionable insights.
Experience building monitoring, dashboards, alerts, and automation around infrastructure usage and cost.
Strong ownership mindset and the ability to work cross-functionally with Engineering, Research, Data, and Product teams.
Curiosity about AI infrastructure and a passion for building scalable, efficient systems.
