Machine Learning Researcher – RL and Agentic Systems
Aktuelle Original-Stellenanzeige
Quelle: StudySmarter Stellenbestand · Status: aktiv · Bewerbung über das zentrale StudySmarter-Formular.
Die ganze Ausschreibung von Jobtailor
Automatisch strukturiert · Originaltext unformatiert geliefert
Das ist der Job
Improve internal infrastructure for reproducible experimentation, benchmark management, and evaluation quality.
Darum lohnt es sich
Build quality scorecards and evaluation methods that make dataset strengths, weaknesses, and failure modes legible across teams. Collaborate closely with research and engineering teams to identify data bottlenecks, improve evaluation methodology, and shape internal best practices around task-grounded AI training data.
Responsibilities Design and build datasets, tasks, and environments for benchmarking agentic systems and multi-step model behavior. Translate real-world workflows into structured tasks, interaction traces, trajectories, stateful environments, and verifiable outcomes that can be used to evaluate advanced AI systems.
Develop frameworks that assess diversity, realism, coverage, fidelity, informativeness, and downstream usefulness of datasets for agentic systems. Evaluate planning, tool use, robustness, recovery from failure, task completion, and generalization behavior in RL-style or agentic environments.
Connect model failures back to concrete dataset, environment, or task-design gaps and recommend improvements grounded in empirical evidence. Contribute to tools and systems that automate dataset validation, environment generation, rollout analysis, benchmark construction, and evaluation workflows.
Represent DataLab’S perspective in cross-functional discussions around dataset quality, benchmark design, and frontier agentic-system evaluation. Requirements PhD or equivalent Master’s Degree + 4+ years industry experience in machine learning, computer science, statistics, engineering, mathematics, economics, or related quantitative fields.
Strong understanding of AI model training pipelines, evaluation methodology, and the role of data in shaping model performance. Experience working with large, unstructured, or semi-structured datasets used to train or evaluate ML systems.
Experience with reinforcement learning, sequential decision-making, agentic systems, tool-using models, or multi-step model evaluation. Experience designing tasks, benchmarks, environments, simulations, or evaluation frameworks for real-world model behavior.
Strong intuition for realism, coverage, difficulty, fidelity, and meaningful outcome structure in datasets. Strong experimental design, evaluation, benchmarking, and data-validation skills. High ownership and ability to independently identify and solve high-impact problems. #J-18808-Ljbffr
Bereit?
Bewerbung wird direkt an Jobtailor uebergeben - kein Konto noetig.