Member of Technical Staff - Research Infrastructure Engineer
Mätch Vc — B2B · B2B Software · Industrial
About the role
About Black Forest Labs
We’re the team behind Latent Diffusion, Stable Diffusion, and FLUX—foundational technologies that changed how the world creates images and video. We’re creating the generative models that power how people make images and video—tools used by millions of creators, developers, and businesses worldwide. Our FLUX models are among the most advanced in the world, and we’re just getting started.
Headquartered in Freiburg, Germany with a growing presence in San Francisco, we’re scaling fast while staying true to what makes us different: research excellence, open science, and building technology that expands human creativity.
Why This Role
We're looking for engineers to build and maintain the engine that powers our mission to develop visual intelligence. From maintaining and scaling clusters, to building research platforms to accelerate the rate of innovation, this team operates with large breadth and depth. We build the systems to make multi-week/month long training possible, to orchestrate resources at scale, and at the same time efficiently, enabling the next breakthrough model. If you’re obsessed with distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement, this team would be perfect for you.
What You’ll Work On
Maintain research infrastructure, ensuring health, and optimizing components to extract peak performance from the system (both on application, and infrastructure side)
Scale infrastructure to meet growing research demands while maintaining reliability and performance
Collaborate with research teams to deeply understand their infrastructure needs, and design solutions that balance performance with cost efficiency.
Identify and resolve performance bottlenecks and capacity hotspots through deep analysis of distributed systems at scale.
Build and evolve telemetry and monitoring systems to provide deep visibility into infrastructure performance, utilization, and costs across our cloud and datacenter fleets.
Participate in on-call rotations and incident response to maintain system reliability
Technical Focus
Python, Bash, Go
Kubernetes
Nvidia GPU drivers, and operators
OTel, Prometheus
What We’re Looking For
Experience building or operating large-scale training platforms
Worked with large scale compute clusters (GPUs)
Proven ability to debug performance and reliability issues across large distributed fleets
Strong problem-solving skills and ability to work independently
Strong communication skills and the ability to work effectively with both internal and external partners
Deep knowledge of modern cloud infrastructure including Kubernetes, Infrastructure as Code, AWS, and GCP
Experience with SLURM
Experience building or operating large-scale training platforms
How We Work Together
We’re a distributed team with real offices that people actually use. Depending on your role, you’ll either join us in Freiburg or SF at least 2 days a week (or one full week every other week), or work remotely with a monthly in-person week to stay connected. We’ll cover reasonable travel costs to make this possible. We think in-person time matters, and we’ve structured things to make it accessible to all. We’ll discuss what this will look like for the role during our interview process.
Everything we do is grounded in four values:
Obsessed. We are a frontier research lab. The science has to be right, the understanding deep, the product beautiful.
Low Ego. The work speaks. The best idea wins, no matter who said it. Credit is shared. Nobody is above any task.
Bold. We take the ambitious bet. We ship, we do not wait for conditions to be perfect.
Kind. People over politics. We treat each other with genuine warmth. Agency without empathy creates chaos.
If this sounds like work you’d enjoy, we’d love to hear from you.
Base Annual Salary:
EU €100,000 - €230,000 + Equity
US $150,000 - $300,000 + Equity
This role is based in our Freiburg / San Francisco office. We operate a hybrid model and cover reasonable travel costs — relocation is encouraged but not required. We do expect a meaningful in-person presence, and we'll discuss what that looks like for your situation during the process.
What the index says about this role
- First seen by JobLarper — Aug 2, 2026, 6 days ago. Older postings collect hundreds of applicants — a tailored résumé matters more the longer a role has been live.
- What DevOps & Site Reliability Engineer roles ask for — across 1,405 indexed openings: Cloud (50%), Python (41%), REST/APIs (33%), Go (31%), Backend (24%). This posting names Cloud, Python, REST/APIs, Go, Backend.
- Mätch Vc is hiring actively — 4 open roles indexed.
Derived from the 27,000 roles JobLarper indexes daily from official company boards — not from the job description above.
More open roles at Mätch Vc
- Member of Technical Staff - Research EngineerSan Francisco (USA), Freiburg (Germany) · Senior+
- Forward Deployed Robotics EngineerFreiburg (Germany) · Mid
- Member of Technical Staff - Model Serving / API Backend EngineerSan Francisco (USA) · Senior+
All 4 open roles at Mätch Vc →
Similar DevOps & Site Reliability Engineer roles at other companies
- Finance Systems Engineer, TaxAnthropic · San Francisco, CA | Seattle, WA
- Senior Site Reliability Engineer (SRE & Platform Reliability)Affirm · Remote Poland · Remote
- Senior SRE - VOIPRingCentral · Bangalore, India
- Senior OT / Edge DevOps Engineer (m/f/x)Reverion · Eresing · München
- Principal AI Platform EngineerSentinelOne · Brno, South Moravian, Czech Republic
- Lead Cloud DevOps Engineer (Oakland, CA Office)Fictiv · Oakland, CA Office
- Principal Site Reliability EngineerDell Technologies · Bengaluru, Karnataka, India
- Senior Application Support Engineer / Site Reliability Engineer (SRE)DTCC · Boston, MA, United States
Browse all devops & site reliability engineer jobs in europe — 176 open roles across 114 companies.