Let’s get started
By clicking ‘Next’, I agree to the Terms of Service
and Privacy Policy
Jobs / Job page
Head of Infrastructure image - Rise Careers
Job details

Head of Infrastructure

About FluidStack

Over the last several years, AI has transformed from a niche technology to one that is surpassing human-level performance on a diverse range of tasks. As a consequence, the amount of compute used for training and inference has grown 5x year-on-year. Plans are being made for training runs that cost upwards of one billion USD and companies are planning to spend hundreds of billions on AI infrastructure in the coming few years.

FluidStack is accelerating this trend, building and operating GPU supercomputers for top AI labs, governments, and enterprises. Our customers include Mistral, Poolside, Black Forest Labs, Meta, and more.

Our team is lean, highly motivated, and focused on providing the world’s best supercomputing experience. We put our customers first in everything we do, and we hold ourselves and each other to insanely high standards. We expect you to care deeply about the work you do, the products you build, and the experience our customers have in every interaction with us.

About the Role

FluidStack is hiring a Head of Infrastructure to lead deployments of 10,000+ GPU supercomputers globally. Reporting directly to the co-founder/president, you will lead our engagements with OEMs, data centers, ISPs, and all relevant infrastructure partners. You will own sourcing, procurement, and be responsible for the timely deployment of some of the largest GPU supercomputers in the world.

You will be in charge of building a world-class deployment team to deliver multi-thousand GPU clusters in a matter of days. This is a unique opportunity to build the infrastructure function from the ground up in an extremely fast-paced environment, as well as a chance to shape the future of AI.

You are expected to have exceptional technical and interpersonal communication skills. You should be able to concisely and accurately share knowledge, in both written and verbal form, with teammates, customers, and suppliers.

Focus

  • You are responsible for the entire supply chain, from the original sourcing of each individual component up to the handover of the burned-in cluster to our customer.

  • You own relationships with OEMs, with the responsibility of continuously improving delivery timelines and costs across the entire supply chain.

  • You will design and build AI clusters, combining your deep knowledge in the area with high-level customer requirements and our past deployment learnings.

  • You are responsible for sourcing additional data center capacity to support our rapidly scaling supercomputer business.

  • You will hire and manage a small but effective “swat team” of deployment engineers, responsible for the world’s fastest setup, burn-in, and delivery of reliable GPU clusters.

  • You will partner with engineering, sales, finance, and legal to always have infrastructure ready one step ahead of our customer needs.

  • You will be required to travel significantly, be it to conferences/trade shows, data centers, customer sites, OEM factories, etc.

About You

An ideal candidate meets at least the following requirements:

  • 3+ years of related experience deploying GPU clusters; 5+ years deploying infrastructure at global scale.

  • On-site experience physically setting up hardware in data centers.

  • Strong relationships with, and history of procuring from, compute and storage OEMs, data centers, ISPs, and others.

  • Experience with InfiniBand or RoCE networking deployments.

  • An understanding of the software that runs on these clusters: Kubernetes/SLURM, PyTorch/Jax, etc.

  • Extreme attention to detail and ability to prioritize and deliver in a fast-paced environment.

  • Highly proactive with an extreme sense of urgency and ownership.

  • Strong engineering background, preferably in Computer Engineering, Electrical Engineering, Computer Science, Software Engineering, Math, Operations Research, Logistics, or similar fields.

Exceptional candidates have one or more of the following experiences:

  • You have designed, built, and operated a 4000+ GPU cluster.

  • You have build tooling to manage bare metal hardware via MaaS, Netbox, or similar tooling.

  • You have deployed and managed petabyte scale all-flash storage systems, including DDN, VAST, and/or Weka; or Ceph, LUSTRE, or similar open source tools.

Benefits

  • Competitive total compensation package (cash + equity).

  • Retirement or pension plan, in line with local norms.

  • Health, dental, and vision insurance.

  • Generous PTO policy, in line with local norms.

  • FluidStack is remote first, but has offices in London, New York, and SF. For all other locations, we provide access to WeWork, if desired.

  • All-paid business travel to hardware and data center shows around the world.

FluidStack Glassdoor Company Review
5.0 Glassdoor star iconGlassdoor star iconGlassdoor star iconGlassdoor star iconGlassdoor star icon
FluidStack DE&I Review
5.0 Glassdoor star iconGlassdoor star iconGlassdoor star iconGlassdoor star iconGlassdoor star icon
CEO of FluidStack
FluidStack CEO photo
Unknown name
Approve of CEO
What You Should Know About Head of Infrastructure, FluidStack

Join FluidStack as the Head of Infrastructure, where you’ll play a pivotal role in shaping the future of AI! Our company is dedicated to building and operating GPU supercomputers for the foremost AI labs, governments, and enterprises globally. In this leadership position, you’ll oversee the deployment of over 10,000 GPU supercomputers, working closely with OEMs, data centers, ISPs, and other infrastructure partners. Reporting directly to our co-founder, you’ll have the autonomy to craft a world-class deployment team tasked with delivering these powerful clusters in record time. Your day-to-day will involve handling the entire supply chain—from sourcing individual components to ensuring timely handovers of clusters to customers. You’ll partner with various teams within FluidStack, leveraging your technical savvy and exceptional communication skills to ensure that infrastructure keeps pace with our growing customer needs. We're looking for someone with a strong background in deploying GPU clusters at a global scale, plus hands-on experience in data centers. The ideal candidate will possess proactive leadership abilities and exceptional attention to detail in this fast-paced environment. If you're ready to take on a challenge that not only affects our clients but also shapes the AI ecosystem, FluidStack is the place for you! We offer competitive compensation and the flexibility of remote-first work culture, with access to WeWork offices when needed. Embark on this exciting journey and lead our infrastructure initiatives at FluidStack!

Frequently Asked Questions (FAQs) for Head of Infrastructure Role at FluidStack
What are the responsibilities of the Head of Infrastructure at FluidStack?

As the Head of Infrastructure at FluidStack, you will be responsible for managing the deployment of over 10,000 GPU supercomputers worldwide. This includes overseeing the entire supply chain from sourcing components to delivering clusters to clients, developing relationships with OEMs, and hiring a dedicated team of deployment engineers. Your role will be central to ensuring the infrastructure is ready ahead of customer needs, driving efficiency, and participating in cross-functional collaborations.

Join Rise to see the full answer
What qualifications are required for the Head of Infrastructure position at FluidStack?

Candidates should have a minimum of 3 years of experience deploying GPU clusters and 5 years in infrastructure deployment on a global scale. A strong engineering background—ideally in Computer Engineering, Electrical Engineering, or related fields—is also essential. Familiarity with hardware setups in data centers, as well as knowledge of network deployment technologies like InfiniBand or RoCE, will greatly enhance your candidacy.

Join Rise to see the full answer
How does FluidStack support the ongoing learning and development of its Head of Infrastructure?

FluidStack places a strong emphasis on continuous learning and personal growth for its leadership team, including the Head of Infrastructure. You will have the opportunity to attend hardware and data center shows globally, participate in relevant conferences, and engage with industry leaders to stay at the forefront of technology trends. We also encourage collaboration and knowledge sharing within our dynamic team environment.

Join Rise to see the full answer
What is the work culture like at FluidStack for the Head of Infrastructure role?

FluidStack fosters a remote-first work culture that emphasizes flexibility, collaboration, and high standards of performance. The Head of Infrastructure will work with a motivated team that prioritizes customer satisfaction and innovation in AI computing. You're encouraged to take ownership of your role and contribute actively to the company’s mission while engaging with various teams to achieve exceptional results.

Join Rise to see the full answer
What are the travel requirements for the Head of Infrastructure at FluidStack?

The Head of Infrastructure position at FluidStack will involve significant travel to conferences, data centers, OEM factories, and customer sites across the globe. This travel will be essential for building relationships with partners and ensuring that deployments are executed successfully in various locations. FluidStack covers all travel expenses, allowing you to focus on your critical role.

Join Rise to see the full answer
Common Interview Questions for Head of Infrastructure
Can you describe your experience with deploying GPU clusters?

When answering this question, share specific examples of GPU clusters you've deployed, detailing your role in the process from planning to execution. Discuss the size and scale of the clusters, any challenges faced, and how you addressed them. Highlight your understanding of both the technical aspects and project management skills involved.

Join Rise to see the full answer
What strategies do you use for sourcing hardware components?

Discuss the methods you employ for sourcing hardware components efficiently, such as building relationships with OEMs and suppliers. Be sure to mention any tools or databases you use to streamline this process and how you ensure resource availability aligns with project timelines.

Join Rise to see the full answer
How do you manage a deployment team in a fast-paced environment?

Explain your approach to building and leading teams, emphasizing your leadership style and communication skills. Share how you foster collaboration and productivity among team members, ensuring everyone is aligned with the project goals and maintains a high standard of output.

Join Rise to see the full answer
How do you prioritize tasks in a multi-project setting?

Outline your prioritization process, which might include assessing project timelines, resources, and impact on customer satisfaction. Demonstrate how your methodical approach ensures that critical tasks are completed on time while adjustable to accommodate changes or emergencies.

Join Rise to see the full answer
What experience do you have with InfiniBand or RoCE networking?

Detail your technical knowledge and hands-on experience with InfiniBand or RoCE networking technologies. Discuss specific projects where you've implemented these solutions, including any outcomes or improvements that resulted from your work.

Join Rise to see the full answer
How do you ensure the quality of the deployed clusters?

Talk about the processes and checks you implement for quality assurance during the deployment phase. This may involve testing protocols, performance benchmarks, and feedback loops with your team and clients to ensure satisfaction with the end product.

Join Rise to see the full answer
Can you provide an example of a successful relationship you've built with an OEM?

Provide a compelling narrative of how you established a successful partnership with an OEM. Focus on your negotiation tactics, the benefits to your organization, and the long-term impacts of the collaboration on project success and supply chain efficiency.

Join Rise to see the full answer
In your opinion, what is key to managing infrastructure for AI initiatives?

Discuss the importance of alignment with AI project requirements, scalability, and responsiveness to changing technologies. Share insights on how you anticipate infrastructure needs based on current trends and future scaling within the AI field.

Join Rise to see the full answer
How do you handle setbacks during hardware deployments?

Share your approach to problem-solving when faced with setbacks, emphasizing your ability to remain calm under pressure, analyze the root causes, and implement solutions quickly. Provide an example of a setback you've encountered and how you successfully navigated it.

Join Rise to see the full answer
What tools do you find essential for infrastructure management?

List the tools and software that you rely on for infrastructure management, support for automation, monitoring, or orchestration. Explain how these tools enhance your efficiency and contribute to overall project success.

Join Rise to see the full answer
Similar Jobs
Posted yesterday
Photo of the Rise User
Posted yesterday
Inclusive & Diverse
Rise from Within
Mission Driven
Diversity of Opinions
Work/Life Harmony
Customer-Centric
Social Impact Driven
Passion for Exploration
Family Medical Leave
Maternity Leave
Paternity Leave
Family Coverage (Insurance)
Medical Insurance
Dental Insurance
Vision Insurance
Mental Health Resources
Life insurance
Disability Insurance
Health Savings Account (HSA)
Flexible Spending Account (FSA)
Photo of the Rise User
Mission Driven
Social Impact Driven
Passion for Exploration
Reward & Recognition
Photo of the Rise User
Posted 5 days ago
Photo of the Rise User
Posted 5 days ago
Photo of the Rise User
AECOM Remote Stevens Point, WI, United States
Posted 4 days ago
Photo of the Rise User
Posted 9 days ago
Photo of the Rise User
Posted 22 hours ago
MATCH
Calculating your matching score...
FUNDING
DEPARTMENTS
SENIORITY LEVEL REQUIREMENT
TEAM SIZE
No info
EMPLOYMENT TYPE
Full-time, remote
DATE POSTED
January 9, 2025

Subscribe to Rise newsletter

Risa star 🔮 Hi, I'm Risa! Your AI
Career Copilot
Want to see a list of jobs tailored to
you, just ask me below!