Senior AI Platform Engineer (Python, AWS, Data Pipelines)

Jetzt bewerben Bewerbung ohne Konto fortsetzen
Aus der Stellenanzeige

Vollständige Stellenanzeige von Aether Biomedical

Originaltext · vollständig und lesefreundlich formatiert

Location: remote with the first day onboarding in Warsaw and occasional visits once per quarterRate: 170 pln/h on b2b

Project Overview

We are building CaaS (Content as a Service) a platform that transforms publisher content (PDF textbooks and Excel manifests) into structured, enriched, AI-ready data. The platform processes content once and exposes it through a unified service layer used by multiple downstream applications.

Key use cases

  • RAG-based Teacher Assistant
  • Editorial tooling
  • Future AI-powered student-facing products

The goal of this role is to design, build, and maintain a scalable data and AI platform that ingests, processes, enriches, and serves content reliably across multiple environments and consumers.

Responsibilities

Data Engineering & PipelinesBuild and maintain multi-stage data ingestion pipelinesDesign and implement idempotent, restartable batch processing workflowsUse S3 as core storage layer for raw and processed dataImplement pipeline stages including:Content ingestion and book identity assignmentPDF-to-markdown conversion (AI OCR)Table of contents and structure extractionHierarchical chunkingEmbedding generationAI / LLM ProcessingUse LLMs and OCR models to extract structured data from PDFsDesign prompts and context strategies for consistent outputsGenerate structured metadata and enrich content for downstream use casesData Storage & ConsistencyMaintain PostgreSQL (Aurora) as system of recordDesign and maintain SQL schemas and versioned migrationsEnsure data consistency across:S3PostgreSQL (Aurora)Vector database (Weaviate)Implement reconciliation logic across distributed systemsRetrieval &

Vector SearchWork with Weaviate for vector search and semantic retrievalSupport RAG-based applicationsDesign data organization strategies (by subject, country, and client)APIs & IntegrationBuild REST APIs using FastAPIExpose content as a service for multiple downstream applicationsIntegrate with internal and external systemsEngineering PracticesWrite strongly typed Python code (mypy)Follow CI/CD processes with automated checks (ruff, pytest)Work across dev / staging / production environmentsDebug distributed data inconsistencies

Key Requirements

  • Must-have
  • Strong Python development experience (production systems)
  • AWS experience (S3, Glue, Aurora)
  • Experience with data pipelines (ETL / batch processing)
  • Strong SQL and PostgreSQL experience
  • Experience with schema design and migrations
  • Nice-to-have
  • Experience with LLMs in production (OCR, content processing, enrichment)
  • Prompt engineering / context engineering
  • Experience with vector databases (Weaviate, Pinecone, Qdrant, pgvector)
  • Knowledge of embeddings, semantic search, and RAG
  • Experience with FastAPI
  • Experience with Airflow / MWAA
  • Experience building data platforms serving multiple consumers

Bereit?

Bewerbung für Aether Biomedical fortsetzen · kein Konto nötig.

Jetzt bewerben
Weitere Stellen bei diesem Arbeitgeber

Aether Biomedical hat 6 weitere offene Stellen:

Alle 7 Stellen bei Aether Biomedical ansehen →
Ähnliche Stellen

Wenn dir dieser Job gefällt, schau dir auch an:

Weiter stöbern:

Kostenfrei starten