Job
- Level
- Senior
- Job Feld
- IT, DevOps, Back End
- Anstellung
- Vollzeit
- Vertragsart
- Unbefristetes Dienstverhältnis
- Ort
- Berlin
- Arbeitsmodell
- Hybrid, Onsite
Job Zusammenfassung
In dieser Position überwachst du die Produktionssysteme, optimierst ML- und LLM-Infrastruktur und entwickelst tollere Lösungen zur Automatisierung, Sicherheit und Zuverlässigkeit der Plattformen in einer globalen Umgebung.
Job Technologien
Deine Rolle im Team
- As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly.
- You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security.
- The SRE team owns the foundation of AI Platform's Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features.
- We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org.
- Reliability of platform(includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services.
- Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks we run everywhere else.
- Company-wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team.
- Developer tooling and automation that compounds - reusable GitHub Actions, GitOps workflows, Terraform modules - so every engineer ships faster.
- Reusable components packaging common open-source tools (Grafana, Istio, CloudNative stack, and ML tooling such as model registries and feature stores) for teams to deploy in any environment.
- Secure-by-default infrastructure - baking security, compliance audits, cost governance, and audit trails into the platform in close partnership with our lead/backend/staff engineers.
Unsere Erwartungen an dich
Qualifikationen
- Good proficiency in Python or Go or general scripting for automation and tooling(automation with higher language preferred).
- AI is already in your daily loop - Agentic tooling (Claude Code, Codex, Droid, internal skills) is part of how you ship and not what you are experimenting with.
- First-principles reasoning - Reasoning from constraints and failure modes naming the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off).
- At least one infrastructure build you owned end to end - with the outcome metric attached (deploy time, MTTR, cost, adoption, availability).
- Cross-functional strength. Track record working with product, backend/frontend teams to pull through collective initiative.
- Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management.
- Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe).
- MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents).
- LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request).
- Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails.
- Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
- Problem Solving: Ability to break down complex problems into simple, actionable solutions.
- Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
- Availability: Willingness to support processes for 24x7 operational support.
Erfahrung
- 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes.
- Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud (AWS preferred).
Unser Angebot
- At CloudFactory, we believe that work should be more than just a job-it should be a platform for growth, impact, and community.
- Here, you'll earn with purpose, learn every day, and serve a mission that truly matters.
- If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we'd love to have you on this journey!
Benefits
Work-Life-Integration
Themen mit denen du dich im Job beschäftigst
Job Standorte
Das ist dein Arbeitgeber
Cloudfactory GmbH
Die CloudFactory GmbH, als Tochtergesellschaft des internationalen Unternehmens CloudFactory, fokussiert sich auf IT-Dienstleistungen, darunter Beratung, Projektleitung und System Engineering. Mit mehreren Standorten in Deutschland und der Schweiz bietet sie umfassende Lösungen im IT-Bereich an.
Description
- Unternehmenstyp
- Etablierte Firma
- Arbeitsmodell
- Hybrid, Onsite
- Branche
- Internet, IT, Telekom