JustJoin.IT Praca zdalna Senior

Senior AI Platform Engineer (Python, AWS, Data Pipelines)

emagine Polska

⚲ Warszawa

160 - 170 PLN netto (B2B)

Wymagania

  • AWS
  • PostgreSQL
  • LLM
  • SQL
  • Python

Opis stanowiska

Location: remote with the first day onboarding in Warsaw and occasional visits once per quarterRate: 170 pln/h on b2b

Project Overview
We are building CaaS (Content as a Service) — a platform that transforms publisher content (PDF textbooks and Excel manifests) into structured, enriched, AI-ready data.
The platform processes content once and exposes it through a unified service layer used by multiple downstream applications.
Key use cases:
• RAG-based Teacher Assistant
• Editorial tooling
• Future AI-powered student-facing products
The goal of this role is to design, build, and maintain a scalable data and AI platform that ingests, processes, enriches, and serves content reliably across multiple environments and consumers.

Responsibilities

Data Engineering & Pipelines
• Build and maintain multi-stage data ingestion pipelines
• Design and implement idempotent, restartable batch processing workflows
• Use S3 as core storage layer for raw and processed data
• Implement pipeline stages including:• Content ingestion and book identity assignment
• PDF-to-markdown conversion (AI OCR)
• Table of contents and structure extraction
• Hierarchical chunking
• Embedding generation

AI / LLM Processing
• Use LLMs and OCR models to extract structured data from PDFs
• Design prompts and context strategies for consistent outputs
• Generate structured metadata and enrich content for downstream use cases

Data Storage & Consistency
• Maintain PostgreSQL (Aurora) as system of record
• Design and maintain SQL schemas and versioned migrations
• Ensure data consistency across:• S3
• PostgreSQL (Aurora)
• Vector database (Weaviate)

• Implement reconciliation logic across distributed systems

Retrieval & Vector Search
• Work with Weaviate for vector search and semantic retrieval
• Support RAG-based applications
• Design data organization strategies (by subject, country, and client)

APIs & Integration
• Build REST APIs using FastAPI
• Expose content as a service for multiple downstream applications
• Integrate with internal and external systems

Engineering Practices
• Write strongly typed Python code (mypy)
• Follow CI/CD processes with automated checks (ruff, pytest)
• Work across dev / staging / production environments
• Debug distributed data inconsistencies

Key Requirements

Must-have
• Strong Python development experience (production systems)
• AWS experience (S3, Glue, Aurora)
• Experience with data pipelines (ETL / batch processing)
• Strong SQL and PostgreSQL experience
• Experience with schema design and migrations

Nice-to-have
• Experience with LLMs in production (OCR, content processing, enrichment)
• Prompt engineering / context engineering
• Experience with vector databases (Weaviate, Pinecone, Qdrant, pgvector)
• Knowledge of embeddings, semantic search, and RAG
• Experience with FastAPI
• Experience with Airflow / MWAA
• Experience building data platforms serving multiple consumers

🔍 Dekoder Ogłoszenia

🔴
occasional visits once per quarter
Może oznaczać konieczność podróży służbowych do Warszawy kilka razy w roku, co może być uciążliwe.
🔴
AI-ready data
Dane będą wymagały znaczącego przetworzenia i przygotowania, zanim będą faktycznie użyteczne dla modeli AI.
🔴
AI OCR
Technologia OCR (optyczne rozpoznawanie znaków) będzie wykorzystywana, ale jej skuteczność w przypadku PDF-ów może być zmienna i wymagać dopracowania.
🔴
Design prompts and context strategies for consistent outputs
Oznacza to, że będziesz musiał eksperymentować z różnymi sposobami formułowania zapytań do modeli LLM, aby uzyskać powtarzalne i użyteczne wyniki.
🟡
Maintain PostgreSQL (Aurora) as system of record
Oczekuje się, że będziesz odpowiedzialny za utrzymanie i ewentualne rozwijanie bazy danych, co może wykraczać poza samo projektowanie.