Job
- Level
- Senior
- Ort
- München
- Arbeitsmodell
- Hybrid, Onsite
- Job Feld
- IT, Software, DevOps
- Anstellung
- Vollzeit
- Vertragsart
- Unbefristetes Dienstverhältnis
- Gehalt
- 75.000 bis 90.000€ Brutto/Jahr
Job Zusammenfassung
In diesem Job entwickelst du automatisierte Lösungen zur Überwachung und Verbesserung der Produktionspipeline, erstellst zuverlässige Alarmierungen und baust KI-Agenten, die Probleme proaktiv diagnostizieren und beheben.
Job Technologien
Deine Rolle im Team
- Our crawlers collect regulatory updates from 80+ sources and push them through an automated pipeline: extraction, embedding, search, classification, consolidation and translation. As we add regions, keeping it healthy has become a job of its own.
- We want an engineer who makes production tell us what's wrong before customers do, fixes what can be fixed automatically, and builds AI agents that diagnose the rest. When an issue does reach a developer, it should arrive with a root cause and a suggested fix.
- Not a ticket-driven ops role: you'll write production Python from week one and own platform reliability alongside a core-team developer.
- Cut the noise. Separate transient failures (network blips, timeouts, errors that vanish on rerun) from real ones, and build alerting the team trusts.
- Build agentic incident response: agents that gather Sentry issues, logs, metrics and recent deploys, classify the failure, apply known fixes or open draft PRs, and brief the right developer.
- Monitor the data, not just the infrastructure. A job that succeeds but extracts nothing is still a failure. Track freshness and completeness, e.g. 'every source checked on time', 'every document produced text and embeddings'.
- Make the pipeline self-healing: retries with backoff, idempotent and resumable jobs, deadletter handling, clear escalation when automation gives up.
- Keep agents safe: scoped permissions, audit trails, human approval for risky actions, and measuring how often they're right.
- Continuously audit our Infrastructure and Identify opportunities to make it more efficient and save costs.
- Catch memory, timeout and cost problems across Cloud Run before they become silent OOM kills.
- Add structured logging, metrics, tracing and sensible Sentry grouping, with infrastructure as code and CI/CD.
Unsere Erwartungen an dich
Qualifikationen
- A track record of turning an ignored alert channel into one people act on.
- Solid GCP (or similar), containers and CI/CD.
- Strong grasp of distributed-system failure modes: retries, idempotency, partial failure, backpressure.
- Pragmatism, and clear communication in a small team.
Erfahrung
- 5+ years in software engineering, SRE or platform roles, with real production Python.
- Hands-on experience building with LLMs or agents in production, and judgment about when a plain if-statement is better.
Unser Angebot
- Hybrid work culture: join us in our Munich office (min. 2 days/week).
- Flexible working hours.
- 26+4 vacation days per year (4 fixed 'company rest days' over Christmas).
- 30 days of 'workation' per year, within the EU and selected countries.
- High autonomy and flat hierarchies.
- EGYM Wellpass for unlimited access to fitness courses and gyms.
- Udemy access for educational videos.
Benefits
Work-Life-Integration
Themen mit denen du dich im Job beschäftigst
Job Standorte
Das ist dein Arbeitgeber
Certivity
Certivity ist ein RegTech-Startup mit Sitz in München, das eine KI-gestützte Plattform entwickelt, um Unternehmen in regulierten Branchen zu unterstützen. Diese Plattform wandelt regulatorische Dokumente in strukturierte Informationen um und hilft Teams, regulatorische Änderungen zu überwachen. Seit der Gründung im Jahr 2021 hat das Unternehmen mehr als 15.000 Nutzer in 10 Ländern gewonnen.
Description
- Unternehmenstyp
- Startup
- Branche
- Internet, IT, Telekom