Job
- Level
- Erfahren
- Job Feld
- IT, System, Security
- Anstellung
- Vollzeit
- Vertragsart
- Befristetes Dienstverhältnis
- Ort
- Tübingen
- Arbeitsmodell
- Hybrid, Onsite
Job Zusammenfassung
In dieser Rolle entwirfst und betreibst du Hochleistungsrechencluster, entwickelst Sicherheitsarchitekturen und automatisierst die Bereitstellung in einer anspruchsvollen HPC-Umgebung für Machine Learning.
Job Technologien
Deine Rolle im Team
- Design and operate their HPC clusters across four data centers, including scheduler (SLURM), parallel filesystems, networks and accelerators, ensuring high availability and throughput for research workloads.
- Conceive and establish the security architecture of the Machine Learning Science Cloud and harden the HPC environment.
- Evolve the automated provisioning and configuration of heterogeneous compute, storage and network nodes (e.g. image-based provisioning, node lifecycle).
- Run patch and vulnerability management - risk assessment across heterogeneous systems.
- Build and operate logging, monitoring and intrusion detection, and integrate HPC telemetry into both operational dashboards and incident-response workflows.
- Lead incident response for the clusters: detection, containment, forensic support and post-incident review.
- Automate operations and security policy as code (Ansible/IaC).
- Advice researchers on efficient, secure cluster usage (job scheduling, data handling, access workflows) and derive requirements for our further roadmap from their scientific workloads.
Unsere Erwartungen an dich
Ausbildung
- Masters degree in Computer Science or a related field.
Qualifikationen
- In-depth IT security knowledge: system hardening, network security, applied cryptography and IAM - with the ability to derive architectural decisions from a threat model, not only to apply given baselines.
- Hands-on HPC background: Slurm, parallel file systems (Weka, Lustre, Ceph), GPU workloads and high-speed networks (InfiniBand, 400G Ethernet).
- A plus: security frameworks (ISO 27001, BSI Grundschutz), container security (Apptainer/Singularity/Docker) or offensive-security fundamentals.
- Independent, structured working style and good communication in English; German is a plus.
- A collaborative, user-facing mindset - comfortable supporting and advising researchers and translating their needs into platform design.
Erfahrung
- Solid Linux experience in production environments (RHEL/Almalinux/Ubuntu).
- Experience with virtualization for management-plane and infrastructure services (Proxmox).
- Strong scripting and automation skills (Bash, Python, Ansible) and experience with configuration management / Infrastructure-as-Code.
Unser Angebot
- Technically deep, architecturally open work that directly enables cutting-edge machine-learning research.
- Flexible working hours and the option to work partially from home.
- Working in an English-speaking, international team of HPC experts.
- A small, senior team with flat hierarchy where responsibility is split by domain.
- Ownership of a technical domain in a production environment of real scale.
- Professional development, conference attendance and real influence on our roadmap.
Themen mit denen du dich im Job beschäftigst
Job Standorte
Das ist dein Arbeitgeber
Cyber Valley GmbH
Die Cyber Valley GmbH fungiert als zentrale Management- und Koordinationsstelle des Cyber-Valley-Innovationscampus in der Region Stuttgart/Tübingen. Sie bietet eine Plattform für Networking, Veranstaltungen und Programme im Bereich Künstliche Intelligenz und Robotik und fördert den Aufbau sowie die Sichtbarkeit des Ökosystems.
Description
- Unternehmenstyp
- Etablierte Firma
- Arbeitsmodell
- Hybrid, Onsite
- Branche
- Internet, IT, Telekom