Logo Cloudfactory GmbH

Senior Site Reliability Engineer

Job

  • Level
    Senior
  • Job Feld
    IT, DevOps, Back End
  • Anstellung
    Vollzeit
  • Vertragsart
    Unbefristetes Dienstverhältnis
  • Ort
    Berlin
  • Arbeitsmodell
    Hybrid, Onsite
  • Job Zusammenfassung

    In dieser Position überwachst du die Produktionssysteme, optimierst ML- und LLM-Infrastruktur und entwickelst tollere Lösungen zur Automatisierung, Sicherheit und Zuverlässigkeit der Plattformen in einer globalen Umgebung.

    Job Technologien

    Deine Rolle im Team

    • As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly.
    • You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security.
    • The SRE team owns the foundation of AI Platform's Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features.
    • We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org.
    • Reliability of platform(includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services.
    • Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks we run everywhere else.
    • Company-wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team.
    • Developer tooling and automation that compounds - reusable GitHub Actions, GitOps workflows, Terraform modules - so every engineer ships faster.
    • Reusable components packaging common open-source tools (Grafana, Istio, CloudNative stack, and ML tooling such as model registries and feature stores) for teams to deploy in any environment.
    • Secure-by-default infrastructure - baking security, compliance audits, cost governance, and audit trails into the platform in close partnership with our lead/backend/staff engineers.

    Unsere Erwartungen an dich

    Qualifikationen

    • Good proficiency in Python or Go or general scripting for automation and tooling(automation with higher language preferred).
    • AI is already in your daily loop - Agentic tooling (Claude Code, Codex, Droid, internal skills) is part of how you ship and not what you are experimenting with.
    • First-principles reasoning - Reasoning from constraints and failure modes naming the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off).
    • At least one infrastructure build you owned end to end - with the outcome metric attached (deploy time, MTTR, cost, adoption, availability).
    • Cross-functional strength. Track record working with product, backend/frontend teams to pull through collective initiative.
    • Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management.
    • Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe).
    • MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents).
    • LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request).
    • Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails.
    • Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
    • Problem Solving: Ability to break down complex problems into simple, actionable solutions.
    • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
    • Availability: Willingness to support processes for 24x7 operational support.

    Erfahrung

    • 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes.
    • Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud (AWS preferred).

    Unser Angebot

    • At CloudFactory, we believe that work should be more than just a job-it should be a platform for growth, impact, and community.
    • Here, you'll earn with purpose, learn every day, and serve a mission that truly matters.
    • If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we'd love to have you on this journey!

    Benefits

    Work-Life-Integration

    Themen mit denen du dich im Job beschäftigst

    Job Standorte

    • Standort Berlin

      Deutschland

    Das ist dein Arbeitgeber

    Cloudfactory GmbH

    Cloudfactory GmbH

    Die CloudFactory GmbH, als Tochtergesellschaft des internationalen Unternehmens CloudFactory, fokussiert sich auf IT-Dienstleistungen, darunter Beratung, Projektleitung und System Engineering. Mit mehreren Standorten in Deutschland und der Schweiz bietet sie umfassende Lösungen im IT-Bereich an.

    Description

  • Unternehmenstyp
    Etablierte Firma
  • Arbeitsmodell
    Hybrid, Onsite
  • Branche
    Internet, IT, Telekom
  • Logo Cloudfactory GmbH

    Senior Site Reliability Engineer

    Ort
    Berlin
    Arbeitsmodell
    Hybrid, Onsite
    Diversität
    Für alle Personen geeignet (m/w/d)
    Nur Englisch
    Nur Englisch erforderlich

    Weitere Jobs