Data/AI Engineer
⚲ Warszawa
Do uzgodnienia
Wymagania
- Python
- Snowflake
- ETL
- AI
Opis stanowiska
Billennium is a global technology company with over 20 years of experience, committed to innovation and empowering businesses. As an employer, we offer a supportive, growth-focused environment where collaboration and creativity thrive. Join us to shape the future of technology together!
Role Purpose
We are looking for a Data / AI Engineer to build and own the upstream data pipelines that prepare large volumes of documents and business data for AI-powered applications.
You will transform diverse, complex data sources into structured, high-quality datasets ready for analytics, retrieval-augmented generation (RAG), and large language model (LLM) workflows. You will ensure that information is consistently collected, cleaned, enriched, processed, and made available for intelligent search and AI-driven solutions.
You will work closely with data scientists and AI engineers to enable reliable, scalable, and production-ready data foundations.
Key Responsibilities
• Own end-to-end data flows from raw documents and source systems to structured data, embeddings, and searchable data stores
• Build and maintain ETL/ELT pipelines for extracting, transforming, and loading data
• Ingest and process data from multiple internal and external sources
• Extract metadata and structure unstructured documents for AI and analytics use cases
• Clean, normalize, and prepare large document collections for LLM and RAG applications
• Develop automated pipelines for document processing, chunking, embedding generation, and indexing
• Ensure data quality, consistency, versioning, and pipeline reliability
• Support the development and optimization of AI solutions using embeddings, vector search, semantic search, and retrieval techniques
• Help design scalable data architectures for growing data volumes
• Collaborate with data scientists, engineers, and business stakeholders to deliver data-driven capabilities
Required Skills & Experience
• Strong Python skills, including experience with data processing libraries such as pandas
• Strong SQL skills and experience working with structured data
• Experience building ETL/ELT pipelines and data processing workflows
• Experience working with document processing and unstructured data
• Familiarity with Docker and modern data platforms or warehouses
• Experience working with databases such as Snowflake, MongoDB, or similar technologies
• Understanding of data modelling principles and data preparation techniques
• Analytical mindset with strong attention to detail
• Ability to work effectively with messy, evolving, and complex data sources
• Clear, direct, and collaborative communication style
Nice to Have
• Experience with OCR technologies and document extraction
• Experience with MongoDB Atlas or similar platforms
• Basic understanding of embeddings, RAG architectures, vector databases, and semantic search
• Experience supporting AI or machine learning projects
Perks and benefits:
•
Comprehensive benefits - enjoy Udemy for Business, private medical care, Multisport card, veterinary package, language lessons, and shopping vouchers.
• Flexibility - adaptable working hours and remote/hybrid work options to suit your lifestyle & location.
• Career growth - access opportunities for professional development and learning, including perks related to our official partnerships with global IT giants: Microsoft, AWS, Snowflake, Salesforce & more.
• Global collaboration - work with a diverse, international team.
• Innovative environment - be part of a forward-thinking and growth-oriented workplace.
• Engaging community – Work with passionate professionals and participate in team-building events, hackathons, and CSR initiatives to make an impact beyond work.
• Team-building events including our company tradition (annual company event in Mazury).
• A pleasant surprise to start your journey with us in the form of a welcome pack.
Recruitment process:
• HR call
• Technical Interview
• Interview with the dedicated Client
• Final decision/ Feedback
Sounds interesting? Click "Apply" and have a chance to hear more!
Your skills & experience
Our offer
Your role
Seniority
Area of expertise
Technology stack (Tags)
Others (pictures etc.)
Role Purpose
We are looking for a Data / AI Engineer to build and own the upstream data pipelines that prepare large volumes of documents and business data for AI-powered applications.
You will transform diverse, complex data sources into structured, high-quality datasets ready for analytics, retrieval-augmented generation (RAG), and large language model (LLM) workflows. You will ensure that information is consistently collected, cleaned, enriched, processed, and made available for intelligent search and AI-driven solutions.
You will work closely with data scientists and AI engineers to enable reliable, scalable, and production-ready data foundations.
Key Responsibilities
• Own end-to-end data flows from raw documents and source systems to structured data, embeddings, and searchable data stores
• Build and maintain ETL/ELT pipelines for extracting, transforming, and loading data
• Ingest and process data from multiple internal and external sources
• Extract metadata and structure unstructured documents for AI and analytics use cases
• Clean, normalize, and prepare large document collections for LLM and RAG applications
• Develop automated pipelines for document processing, chunking, embedding generation, and indexing
• Ensure data quality, consistency, versioning, and pipeline reliability
• Support the development and optimization of AI solutions using embeddings, vector search, semantic search, and retrieval techniques
• Help design scalable data architectures for growing data volumes
• Collaborate with data scientists, engineers, and business stakeholders to deliver data-driven capabilities
Required Skills & Experience
• Strong Python skills, including experience with data processing libraries such as pandas
• Strong SQL skills and experience working with structured data
• Experience building ETL/ELT pipelines and data processing workflows
• Experience working with document processing and unstructured data
• Familiarity with Docker and modern data platforms or warehouses
• Experience working with databases such as Snowflake, MongoDB, or similar technologies
• Understanding of data modelling principles and data preparation techniques
• Analytical mindset with strong attention to detail
• Ability to work effectively with messy, evolving, and complex data sources
• Clear, direct, and collaborative communication style
Nice to Have
• Experience with OCR technologies and document extraction
• Experience with MongoDB Atlas or similar platforms
• Basic understanding of embeddings, RAG architectures, vector databases, and semantic search
• Experience supporting AI or machine learning projects
Perks and benefits:
•
Comprehensive benefits - enjoy Udemy for Business, private medical care, Multisport card, veterinary package, language lessons, and shopping vouchers.
• Flexibility - adaptable working hours and remote/hybrid work options to suit your lifestyle & location.
• Career growth - access opportunities for professional development and learning, including perks related to our official partnerships with global IT giants: Microsoft, AWS, Snowflake, Salesforce & more.
• Global collaboration - work with a diverse, international team.
• Innovative environment - be part of a forward-thinking and growth-oriented workplace.
• Engaging community – Work with passionate professionals and participate in team-building events, hackathons, and CSR initiatives to make an impact beyond work.
• Team-building events including our company tradition (annual company event in Mazury).
• A pleasant surprise to start your journey with us in the form of a welcome pack.
Recruitment process:
• HR call
• Technical Interview
• Interview with the dedicated Client
• Final decision/ Feedback
Sounds interesting? Click "Apply" and have a chance to hear more!
Your skills & experience
Our offer
Your role
Seniority
Area of expertise
Technology stack (Tags)
Others (pictures etc.)
🔍 Dekoder Ogłoszenia
🔴
build and own the upstream data pipelines
Oczekuje się, że będziesz samodzielnie projektować, wdrażać i utrzymywać całe potoki danych, co może oznaczać dużą odpowiedzialność i brak wsparcia.
🔴
transform diverse, complex data sources into structured, high-quality datasets
Może to oznaczać pracę z bardzo nieuporządkowanymi i trudnymi do przetworzenia danymi, co wymaga znacznego wysiłku w zakresie czyszczenia i normalizacji.
🟡
work closely with data scientists and AI engineers
Choć brzmi to jak współpraca, może oznaczać, że będziesz głównie realizować zadania narzucone przez inne zespoły, bez większego wpływu na architekturę.
🟡
ensure that information is consistently collected, cleaned, enriched, processed, and made available
Podkreśla szeroki zakres obowiązków związanych z jakością i dostępnością danych, co może oznaczać bardzo pracochłonne zadania.
🔴
Own end-to-end data flows
Podobnie jak w pierwszym punkcie, sugeruje pełną odpowiedzialność za cały cykl życia danych, od źródła do gotowego produktu.