Job
- Level
- Senior
- Ort
- Berlin
- Arbeitsmodell
- Hybrid, Onsite
- Job Feld
- IT, DevOps, Back End
- Anstellung
- Vollzeit
- Vertragsart
- Unbefristetes Dienstverhältnis
Job Zusammenfassung
In dieser Position verantwortest du den Aufbau und die kontinuierliche Verbesserung einer Produktionsbeobachtungsplattform, die Metriken, Logs und Traces vereint, um das Monitoring und die Incident-Diagnose für Produktteams zu optimieren.
Job Technologien
Deine Rolle im Team
- Own and continuously improve a unified, production-grade observability stack covering metrics, logs, and traces, giving every team a consistent, self-service way to understand and operate the health of their services.
- Build and maintain trusted alerting by tuning alerts against clear SLOs/SLIs, reducing noise, improving ownership and ensuring real issues surface early.
- Define, document, and drive adoption of instrumentation standards covering metric naming, label cardinality, structured logging, and distributed tracing with OpenTelemetry.
- Enable product teams to diagnose and resolve incidents faster through well-designed dashboards, runbooks, hands-on support, and practical observability guidance.
- Run the observability platform like a product, owning its roadmap, interfaces, reliability, scalability, and cost across areas such as retention, sampling, and cardinality.
- Provide clear technical direction for observability across the engineering organisation, making pragmatic architectural trade-offs and helping teams adopt platform capabilities effectively.
Unsere Erwartungen an dich
Qualifikationen
- Instrumentation & telemetry: OpenTelemetry adoption, metric and label design, structured logging, and distributed tracing across complex, multi-team systems.
- SLO/SLI & error budgets: dashboards and alerting that surface real problems early while minimizing noise and false positives.
- Platform as a product: treating other engineering teams as customers, with clear interfaces and a roadmap shaped by their needs.
- Alerting & reliability attitude: comfortable with on-call, driving blameless learning from incidents, and continuously tuning signal quality instead of letting alert fatigue set in.
- A background that includes time spent writing and shipping production software, bringing engineering instincts to building tooling and automation rather than only operating what already exists.
- Proficiency in at least one backend language (for example Python or Go) for automation, tooling, and building internal observability capabilities.
- Excellent written and verbal communication skills in English, able to explain complex system behavior clearly to both engineers and non-technical stakeholders.
Erfahrung
- 6+ years in Platform Engineering, Observability, or infrastructure roles, with end-to-end ownership of a production observability stack (for example the LGTM stack - Loki, Grafana, Tempo, Mimir/Prometheus - or an equivalent metrics/logs/traces platform) at scale.
- Demonstrated experience owning the long-term consequences of your own architectural and tooling decisions - having lived with what you built through its maintenance, upgrades, and failure modes, not just its initial rollout.
Unser Angebot
- Employment is subject to applicable security screening (incl. SUPO).
Themen mit denen du dich im Job beschäftigst
Job Standorte
Das ist dein Arbeitgeber
ICEYE
ICEYE Oy ist ein innovatives Unternehmen aus Finnland, das sich auf die Entwicklung und den Betrieb von SAR-Satelliten spezialisiert hat. Es liefert wertvolle Erdbeobachtungsdaten für diverse Anwendungen.
Description
- Unternehmenstyp
- Startup
- Arbeitsmodell
- Hybrid, Onsite
- Branche
- Luft-, Raumfahrt