We're building an AI-native claims operation. Our thesis: the moat isn't the model — it's the system around it that makes AI reliable enough to trust with real operational work. Building that system is this role. We're an agile, senior team automating residential property claims, with agents running live cases today. You'll own the agentic pipeline and the evaluation and feedback systems that make it measurably good — and, being full-stack, the interfaces where humans work alongside it.
- Our multi-agent pipeline (LangGraph, LangChain) — the agents that read a claim, decide the next step, and communicate with every party
- Context and prompt engineering — getting the right knowledge to the agents and keeping their behaviour reliable
- The evaluation and testing systems that let us measure agent quality and ship changes with confidence — this is the actual product
- The feedback systems that turn real-world signals into measurable improvement over time
- Model routing and experimentation across providers
- The full-stack surfaces where our team reviews and works with the agents
Your profile
- Have 3+ years in software engineering, with real hands-on LLM/agent work — not just calling an API, but building systems that stay reliable when the model doesn't
- Are strong in Python and comfortable enough across the stack to ship a clean internal UI when it's needed
- Think empirically about AI: you reach for evals, not vibes, and you know "it works on my three examples" isn't evidence
- Have built — or badly wanted to build — proper evaluation and feedback systems for an LLM product
- Are comfortable with ambiguity and high ownership in a small team
Nice to have
- Experience with LangGraph/LangChain, Langfuse, or agent-evaluation tooling
- Prompt/context engineering at production scale
- RAG, retrieval, or knowledge-base design
- Insurance, fintech, or another regulated domain
- German (helpful, not required — we work in English)