Site Reliability Engineer (AI Start-up)
Salary: CHF 100'000 - 140'000 per year
Requirements:
- Experience designing and operating distributed systems at scale, with a solid understanding of failure modes, capacity planning and the trade-offs between consistency, availability and latency
- Strong observability and incident response skills: you build monitoring that catches problems before customers do, lead structured incident responses and drive lasting fixes, not just restarts
- Hands-on experience with infrastructure as code, CI/CD pipelines and container orchestration in production
- A security mindset: encryption, network isolation and secure handling of enterprise data are a default for you, not an afterthought
- Based in, or ready to relocate to, Prague, Berlin, Lisbon, Porto or Valencia, with the right to work in the EU (no visa sponsorship)
Responsibilities:
- You'll join a fast-growing, well-funded AI start-up backed by one of Europe's top venture capital firms. They're building an AI operations platform for large enterprises, currently focused on retail and consumer goods. Their AI agents carry out business-critical work wherever it happens, across SAP, spreadsheets, supplier portals, email and APIs, with heavy use of browser and computer automation. Customers rely on these agents for critical workflows, so reliability directly affects their business.
- What you'll do
- This is not a traditional ops role that follows runbooks. You'll build the reliability practice, not maintain one:
- Evolve observability, infrastructure topology and release processes to keep up with high-velocity, AI-assisted development
- Make sure misbehaving AI agents are spotted early: understand their failure modes and find the causes of high latency
- Build tooling that sounds the alarm before customers notice, so the on-call team can assess impact and mitigate issues in minutes, not hours
- Make sure post-mortem actions are followed through across engineering, even if that means adding some friction
- Decide what to automate first, balance reliability against shipping speed, and make incident calls with incomplete information
- The team
- The engineering team has around 20 people and is growing. About half are based at the home base in Prague, and the rest work from small hubs across Europe, all within two hours of Prague time. The culture is built on ownership, direct feedback and customer focus: the team ships weekly, fixes forward, and uses AI tools wherever they help.
Technologies:
- AI
- AI Agents
- AWS
- Azure
- CI/CD
- Cloud
- Datadog
- GCP
- Kubernetes
- LLM
- Network
- REST
- SAP
- Security
- Terraform
- Docker
- PostgreSQL
- Python
- Redis
- TypeScript
More:
Rockstar Recruiting is hiring on behalf of a fast-growing, well-funded AI start-up, backed by one of Europe's top venture capital firms, that is building an AI operations platform for large enterprises, with customers among Europe's biggest retailers. This is a senior Site Reliability Engineer role for someone who knows what good looks like and wants to build a reliability practice from the ground up: owning observability, incident response, security and the reliability of production AI agents, on a cloud-native stack built on GCP, Kubernetes, Terraform and Datadog.
The team works from small hubs across Europe rather than fully remote, so the role is open to candidates based in, or ready to relocate to, Prague, Berlin, Lisbon, Porto or Valencia, with the right to work in the EU. Please get in touch with Tijana of Rockstar Recruiting to discuss further details: tl@rockstar.jobs
last updated 40 week of 2026
Original source: https://swissdevjobs.ch/jobs/Rockstar-Recruiting-AG-Site-Reliability-Engineer-AI-Start-up