Job
- Level
- Senior
- Job Feld
- Software, Data
- Anstellung
- Vollzeit
- Vertragsart
- Unbefristetes Dienstverhältnis
- Ort
- Berlin
- Arbeitsmodell
- Onsite
Job Zusammenfassung
In dieser Position arbeitest du an der Optimierung von AI-Workloads auf NVIDIA-Plattformen, analysierst Leistungsdaten, löst Clusterprobleme und unterstützt Kunden bei der Skalierung hochperformanter Lösungen.
Job Technologien
Deine Rolle im Team
- Collaborating with NVIDIA's training framework developers and product teams to stay ahead of the latest features and help partners to adopt them effectively.
- Assisting with deployment, debugging, and improving the efficiency of AI workloads on extensive NVIDIA platforms.
- Benchmarking new framework features, analyzing performance, and sharing actionable insights with both customers and internal teams.
- Working directly with external customers to solve cluster performance and stability issues, identify bottlenecks, and implement effective solutions.
- Build expertise and guide customers in scaling workloads efficiently and reliably on the latest generation of NVIDIA GPUs.
- Contributing to Europe's Sovereign AI initiative by helping customers implement advanced resiliency features within AI training pipelines.
Unsere Erwartungen an dich
Qualifikationen
- Strong programming skills in at least one of the following languages: C, C++, or Python.
- Solid understanding of CPU and GPU architectures, CUDA, parallel filesystems, and high-speed interconnects.
- Proficient knowledge of training pipelines and frameworks, encompassing their internal operations and performance attributes.
Erfahrung
- BS, MS, PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or a related engineering field-or equivalent practical experience.
- 8+ years of experience in accelerated computing technologies at cluster scale, ideally including work with NVIDIA platforms.
- Practical experience identifying and resolving bottlenecks in large-scale training workloads or parallel applications.
- Hands-on experienced in profiling and debugging large parallel applications.
- Experienced in working with large compute clusters with an understanding of their internal scheduling and resource management mechanisms (e.g. SLURM or Cloud based clusters).
Themen mit denen du dich im Job beschäftigst
Job Standorte
Das ist dein Arbeitgeber
Nvidia
NVIDIA hat sich in den letzten zwei Jahrzehnten immer wieder neu erfunden. Die Erfindung der GPU durch NVIDIA 1999 löste das Wachstum des PC-Spielemarktes aus, definierte moderne Computergrafiken neu und revolutionierte paralleles Computing.
Description
- Gründungsjahr
- 1999
- Sprachen
- Englisch
- Unternehmenstyp
- Etablierte Firma
- Arbeitsmodell
- Full Remote, Hybrid, Onsite
- Branche
- Handel