Job
- Level
- Senior
- Ort
- Hamburg, Stuttgart
- Arbeitsmodell
- Hybrid, Onsite
- Job Feld
- IT, DevOps, Security
- Anstellung
- Vollzeit
- Vertragsart
- Unbefristetes Dienstverhältnis
Job Zusammenfassung
In dieser Rolle übernimmst du die Verantwortung für die Zuverlässigkeit von Systemen, automatisierst manuelle Prozesse, definierst SLOs, betreibst den Kubernetes-Cluster und implementierst Sicherheitsmaßnahmen.
Job Technologien
Deine Rolle im Team
- You join the Platform team and own the reliability of the systems our AI agents run on.
- This is a reliability engineering role: you treat operations as a software problem, so where others run a manual procedure, you write the automation that makes it unnecessary.
- You define what "healthy" means in numbers, measure it, and hold the line on it in production.
- When something breaks, you bring the system back and then make sure it cannot break the same way twice.
- Define and own SLOs and SLIs for the platform and manage error budgets against them.
- Carry on-call, act as incident commander, and run blameless post-incident reviews that produce real follow-up.
- Run production readiness and capacity planning ahead of demand, not after the page fires.
- Run and harden the Kubernetes platform (Helm, GitOps, service mesh) and the cloud underneath it (Terraform, multi-region).
- Own observability: metrics, logs, and distributed tracing, so problems surface before users feel them.
- Eliminate toil through automation, self-healing systems, and automated remediation.
- Drive cost visibility and FinOps practice across cloud and LLM spend.
- Bake security into the platform: least-privilege access, secrets management, policy as code, and vulnerability management.
- Use coding agents to build automation, write infrastructure code, and reason through failure modes faster.
- Automate operational toil and incident workflows with AI in the loop.
- Collaborate with the product team to give real-world feedback on Blockbrain's own tools from an operator's perspective.
- Stay curious about emerging AI capabilities and apply them to platform and reliability work.
Unsere Erwartungen an dich
Ausbildung
- Degree in computer science or a related field, or equivalent hands-on experience.
Qualifikationen
- Clear and calm under pressure.
- Can coordinate an incident and write a post-mortem others learn from.
- Works in English; German is a plus.
- Treats infrastructure as code and operations as a software discipline.
- Genuinely enjoys automating manual work away.
- Measures success in incidents that did not happen.
- Fixes root causes, not symptoms.
- Plans capacity and reliability work ahead of demand and balances on-call, project work, and toil reduction.
- We care about what you can operate, not the certificate.
- Kubernetes, Terraform / IaC, CI/CD (e.g. GitHub Actions), observability (Prometheus, Grafana, distributed tracing), a major cloud (AWS, Azure, or GCP), scripting (Python, Go, or TypeScript), secrets management, and policy as code.
- Strong ownership, blameless culture, a bias toward automation, calm in incidents, and a security-by-default mindset.
Erfahrung
- 5+ years in DevOps, SRE, or platform engineering, ideally operating production SaaS at scale.
- Hands-on experience running Kubernetes in production is essential.
Unser Angebot
- Full-time position on-site in Stuttgart, Hamburg or Munich (3 days per week) with flexible working hours.
- Deutschland-Ticket, Wellpass fitness membership, access to the latest AI tools, and regular company and team off-sites.
- International team with exceptional talents.
- Steep learning curve in a fast-growing AI startup.
- A high level of personal responsibility and the freedom to actively shape processes.
- MacBook, iPhone, headset, and all the tools you need to perform at your best.
Themen mit denen du dich im Job beschäftigst
Job Standorte
Das ist dein Arbeitgeber
Blockbrain
Blockbrain fokussiert sich auf die Automatisierung von Dokumentenprozessen und die Verbesserung des Wissensmanagements. Die Plattform steigert die Effizienz und optimiert die Nutzung von Unternehmenswissen.
Description
- Unternehmenstyp
- Startup
- Arbeitsmodell
- Hybrid, Onsite
- Branche
- Internet, IT, Telekom