Veeda AI Deutschlandweit vor 4 Tagen

Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)

Jetzt bewerben Bewerbung ohne Konto fortsetzen
Aus der Stellenanzeige

Vollständige Stellenanzeige von Veeda AI

Originaltext · vollständig und lesefreundlich formatiert

ph3bAbout Us /b /h3pVeeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence.

If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. /ph3bResponsibilities /b /h3ullipbDistributed Training Systems Scalability: /b Design, optimize, and maintain high-throughput distributed training systems across large-scale GPU clusters for multi-modal foundation models. /p /lilipbPrecision Numerical Stability: /b Debug, diagnose, and resolve subtle numerical instability issues (underflow/overflow, loss spikes, gradient explosion, and mixed-precision divergence) in FP16, BF16, FP8, and custom quantization schemes. /p /lilipbFault Diagnostics Recovery: /b Build advanced fault-detection mechanisms and automated diagnostics to rapidly pinpoint and isolate silent data corruption (SDC), hardware hang/deadlock, memory leaks, and "card-freeze" issues during large training runs. /p

/lilipbPerformance Profiling Optimization: /b Profile distributed communication bottlenecks, memory usage, and kernel execution to improve overall FLOPS utilization across multi-node, multi-GPU training jobs. /p /lilipbDeveloper Tooling Infrastructure: /b Develop resilient checkpointing systems, rapid fault-recovery pipelines, and execution telemetry to keep researcher productivity high and hardware downtime minimal. /p /li /ulh3bRequirements /b /h3ullipYou have a Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field. /p /lilipYou have deep hands-on experience with deep learning training frameworks (e.g., PyTorch) and distributed training paradigms (FSDP, Megatron-LM, DeepSpeed, Tensor Parallelism, Pipeline Parallelism). /p /lilipYou have proven experience in numerical precision analysis, low-precision training

(BF16/FP8), and debugging complex loss divergence/stability issues in massive training runs. /p /lilipYou have strong root-cause analysis skills for hardware/software interaction bugs, including stuck CUDA kernels, NCCL timeouts, GPU hardware faults, and silent training corruptions. /p /lilipYou have strong programming skills in Python and C++/CUDA, with a deep understanding of low-level GPU architectures and memory hierarchies. /p /li /ulh3bNice to Have /b /h3ullipYou have experience running or porting large-scale training workloads on bAMD GPUs /b (ROCm platform) or bGoogle TPUs /b (JAX/XLA stack). /p /lilipYou have contributed to low-level training infrastructure, custom CUDA/Triton kernels, or distributed training open-source projects. /p /lilipYou have built resilient fault-tolerant training frameworks with dynamic node re-queueing and rapid checkpointing/saving mechanisms. /p /li

/ul /p

Bereit?

Bewerbung für Veeda AI fortsetzen · kein Konto nötig.

Jetzt bewerben
Beim Arbeitgeber

Aktuell die einzige offene Stelle bei Veeda AI.

Neue Stellen kommen monatlich dazu — schau gerne später noch mal rein.

Ähnliche Stellen

Wenn dir dieser Job gefällt, schau dir auch an:

Weiter stöbern:

Kostenfrei starten