Logo Cloudfactory GmbH

Senior Site Reliability Engineer

Job

  • Level
    Senior
  • Ort
    Berlin
  • Arbeitsmodell
    Hybrid, Onsite
  • Job Feld
    IT, DevOps, Back End
  • Anstellung
    Vollzeit
  • Vertragsart
    Unbefristetes Dienstverhältnis

Job Zusammenfassung

In dieser Position überwachst du die Produktionssysteme, optimierst ML- und LLM-Infrastruktur und entwickelst tollere Lösungen zur Automatisierung, Sicherheit und Zuverlässigkeit der Plattformen in einer globalen Umgebung.

Job Technologien

Deine Rolle im Team

  • As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly.
  • You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security.
  • The SRE team owns the foundation of AI Platform's Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features.
  • We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org.
  • Reliability of platform(includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services.
  • Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks we run everywhere else.
  • Company-wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team.
  • Developer tooling and automation that compounds - reusable GitHub Actions, GitOps workflows, Terraform modules - so every engineer ships faster.
  • Reusable components packaging common open-source tools (Grafana, Istio, CloudNative stack, and ML tooling such as model registries and feature stores) for teams to deploy in any environment.
  • Secure-by-default infrastructure - baking security, compliance audits, cost governance, and audit trails into the platform in close partnership with our lead/backend/staff engineers.

Unsere Erwartungen an dich

Qualifikationen

  • Good proficiency in Python or Go or general scripting for automation and tooling(automation with higher language preferred).
  • AI is already in your daily loop - Agentic tooling (Claude Code, Codex, Droid, internal skills) is part of how you ship and not what you are experimenting with.
  • First-principles reasoning - Reasoning from constraints and failure modes naming the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off).
  • At least one infrastructure build you owned end to end - with the outcome metric attached (deploy time, MTTR, cost, adoption, availability).
  • Cross-functional strength. Track record working with product, backend/frontend teams to pull through collective initiative.
  • Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management.
  • Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe).
  • MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents).
  • LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request).
  • Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails.
  • Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
  • Problem Solving: Ability to break down complex problems into simple, actionable solutions.
  • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
  • Availability: Willingness to support processes for 24x7 operational support.

Erfahrung

  • 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes.
  • Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud (AWS preferred).

Unser Angebot

  • At CloudFactory, we believe that work should be more than just a job-it should be a platform for growth, impact, and community.
  • Here, you'll earn with purpose, learn every day, and serve a mission that truly matters.
  • If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we'd love to have you on this journey!

Benefits

Work-Life-Integration

Themen mit denen du dich im Job beschäftigst

Job Standorte

  • Standort Berlin

    Deutschland

Das ist dein Arbeitgeber

Cloudfactory GmbH

Cloudfactory GmbH

Die CloudFactory GmbH, als Tochtergesellschaft des internationalen Unternehmens CloudFactory, fokussiert sich auf IT-Dienstleistungen, darunter Beratung, Projektleitung und System Engineering. Mit mehreren Standorten in Deutschland und der Schweiz bietet sie umfassende Lösungen im IT-Bereich an.

Description

  • Unternehmenstyp
    Etablierte Firma
  • Arbeitsmodell
    Hybrid, Onsite
  • Branche
    Internet, IT, Telekom
Logo Cloudfactory GmbH

Senior Site Reliability Engineer

Ort
Berlin
Arbeitsmodell
Hybrid, Onsite
Diversität
Für alle Personen geeignet (m/w/d)
Nur Englisch
Nur Englisch erforderlich

Weitere Jobs