Das ist der Job
Apply expertise in data modeling, normalization, and semantic cataloging for AI/ML workloads.
Darum lohnt es sich
Responsibilities Provision and configure large GPU clusters and compute resources for LLM training, finetuning, and inference workloads. Develop and optimize LLM model‑serving infrastructure, including deployment and optimization of various inference frameworks.
Lead model lifecycle management including versioning, checkpointing and reproducibility across training and inference deployments. Design and champion robust evaluation frameworks to assess model performance, accuracy, and reliability, ensuring AI systems are consistently at production‑ready standards.
Identify and address GPU utilization and GPU memory efficiency bottlenecks and apply techniques like quantization, batching, and caching. Architect and maintain data platforms and pipelines specifically designed to support LLMs, Retrieval‑Augmented Generation (RAG), and AI Agentic Systems at scale.
Deliver production‑ready code with a focus on performance, maintainability, and testing rigor, ensuring the ability to ship fast without compromising quality. Define and enforce best practices for MLOps/DataOps surrounding LLMs, including monitoring, observability, and zero‑touch recovery mechanisms for AI services.
Document architectural designs thoroughly and communicate technical decisions clearly to stakeholders.