Senior AI Platform Engineer (Python, AWS, Data Pipelines)
Vollständige Stellenanzeige von Aether Biomedical
Originaltext · vollständig und lesefreundlich formatiert
Location: remote with the first day onboarding in Warsaw and occasional visits once per quarterRate: 170 pln/h on b2b
Project Overview
We are building CaaS (Content as a Service) a platform that transforms publisher content (PDF textbooks and Excel manifests) into structured, enriched, AI-ready data. The platform processes content once and exposes it through a unified service layer used by multiple downstream applications.
Key use cases
- RAG-based Teacher Assistant
- Editorial tooling
- Future AI-powered student-facing products
The goal of this role is to design, build, and maintain a scalable data and AI platform that ingests, processes, enriches, and serves content reliably across multiple environments and consumers.
Responsibilities
Data Engineering & PipelinesBuild and maintain multi-stage data ingestion pipelinesDesign and implement idempotent, restartable batch processing workflowsUse S3 as core storage layer for raw and processed dataImplement pipeline stages including:Content ingestion and book identity assignmentPDF-to-markdown conversion (AI OCR)Table of contents and structure extractionHierarchical chunkingEmbedding generationAI / LLM ProcessingUse LLMs and OCR models to extract structured data from PDFsDesign prompts and context strategies for consistent outputsGenerate structured metadata and enrich content for downstream use casesData Storage & ConsistencyMaintain PostgreSQL (Aurora) as system of recordDesign and maintain SQL schemas and versioned migrationsEnsure data consistency across:S3PostgreSQL (Aurora)Vector database (Weaviate)Implement reconciliation logic across distributed systemsRetrieval &
Vector SearchWork with Weaviate for vector search and semantic retrievalSupport RAG-based applicationsDesign data organization strategies (by subject, country, and client)APIs & IntegrationBuild REST APIs using FastAPIExpose content as a service for multiple downstream applicationsIntegrate with internal and external systemsEngineering PracticesWrite strongly typed Python code (mypy)Follow CI/CD processes with automated checks (ruff, pytest)Work across dev / staging / production environmentsDebug distributed data inconsistencies
Key Requirements
- Must-have
- Strong Python development experience (production systems)
- AWS experience (S3, Glue, Aurora)
- Experience with data pipelines (ETL / batch processing)
- Strong SQL and PostgreSQL experience
- Experience with schema design and migrations
- Nice-to-have
- Experience with LLMs in production (OCR, content processing, enrichment)
- Prompt engineering / context engineering
- Experience with vector databases (Weaviate, Pinecone, Qdrant, pgvector)
- Knowledge of embeddings, semantic search, and RAG
- Experience with FastAPI
- Experience with Airflow / MWAA
- Experience building data platforms serving multiple consumers
Bereit?
Bewerbung für Aether Biomedical fortsetzen · kein Konto nötig.