Senior Data Engineer (Python, GCP, Airflow)
⚲ Kraków
14 500 - 22 350 PLN (PERMANENT)
Wymagania
- Python
- GCP
- BigQuery
- Apache Airflow
- ETL
- Spark
- NoSQL
- Relational database
- Apache Beam (nice to have)
Opis stanowiska
O projekcie:
What will you do?
You will join a modern data platform initiative focused on building scalable data solutions in Google Cloud Platform. The project involves designing robust data pipelines, developing data models, and ensuring reliable and high-quality data for business decision-making. You will collaborate with cross-functional teams to deliver efficient and maintainable data architecture.
We offer
- Hybrid work in Krakow (2 office days per week)- Working in a highly experienced and dedicated team- Benefit package tailored to your needs (medical, sport, lunch subsidy, life insurance, etc.)- Online training and certifications- Access to e‑learning platform- Social events
Wymagania:
- Strong proficiency in Python for backend or data processing- Hands-on experience with Google Cloud Platform including BigQuery, Cloud Storage, Compute Engine or GKE- Experience with Apache Airflow for workflow orchestration- Strong experience with ETL and data integration pipelines- Experience with Spark- Knowledge of NoSQL databases- Strong knowledge of relational databases- Strong understanding of data architecture- Good understanding of data quality- Strong understanding of data integration
Nice to have skills:
- Familiarity with Apache Beam
Codzienne zadania:
- Design, develop, and maintain scalable data pipelines and ETL processes
- Build and manage Apache Airflow DAGs for workflow orchestration
- Develop data processing logic using Python and frameworks like DBT or Spark
- Work with BigQuery and Cloud SQL to store and transform data
- Implement and maintain data architecture and data models
- Ensure data quality, integrity, and reliability
- Collaborate with business stakeholders to translate requirements into data solutions
- Maintain technical documentation and contribute to data standards
- Support deployment and monitoring of pipelines in GCP environments
- Participate in code reviews and improve engineering best practices
What will you do?
You will join a modern data platform initiative focused on building scalable data solutions in Google Cloud Platform. The project involves designing robust data pipelines, developing data models, and ensuring reliable and high-quality data for business decision-making. You will collaborate with cross-functional teams to deliver efficient and maintainable data architecture.
We offer
- Hybrid work in Krakow (2 office days per week)- Working in a highly experienced and dedicated team- Benefit package tailored to your needs (medical, sport, lunch subsidy, life insurance, etc.)- Online training and certifications- Access to e‑learning platform- Social events
Wymagania:
- Strong proficiency in Python for backend or data processing- Hands-on experience with Google Cloud Platform including BigQuery, Cloud Storage, Compute Engine or GKE- Experience with Apache Airflow for workflow orchestration- Strong experience with ETL and data integration pipelines- Experience with Spark- Knowledge of NoSQL databases- Strong knowledge of relational databases- Strong understanding of data architecture- Good understanding of data quality- Strong understanding of data integration
Nice to have skills:
- Familiarity with Apache Beam
Codzienne zadania:
- Design, develop, and maintain scalable data pipelines and ETL processes
- Build and manage Apache Airflow DAGs for workflow orchestration
- Develop data processing logic using Python and frameworks like DBT or Spark
- Work with BigQuery and Cloud SQL to store and transform data
- Implement and maintain data architecture and data models
- Ensure data quality, integrity, and reliability
- Collaborate with business stakeholders to translate requirements into data solutions
- Maintain technical documentation and contribute to data standards
- Support deployment and monitoring of pipelines in GCP environments
- Participate in code reviews and improve engineering best practices
🔍 Dekoder Ogłoszenia
🔴
modern data platform initiative focused on building scalable data solutions in Google Cloud Platform
Projekt może być na wczesnym etapie rozwoju, co oznacza potencjalne problemy z architekturą i potrzebę budowania od podstaw.
🔴
collaborate with cross-functional teams
Może oznaczać konieczność częstych spotkań i uzgadniania wymagań z wieloma różnymi działami, co może spowalniać pracę.
🔴
Benefit package tailored to your needs
Pakiet benefitów może być standardowy, a 'dopasowanie do potrzeb' oznacza wybór spośród ograniczonej listy opcji.
🔴
Strong proficiency in Python for backend or data processing
Może oznaczać, że będziesz musiał zajmować się zarówno budowaniem API, jak i skomplikowanymi procesami ETL, co wymaga szerokiego zakresu umiejętności.
🟡
Develop data processing logic using Python and frameworks like DBT or Spark
Wymóg użycia DBT lub Spark może sugerować, że te technologie są kluczowe, ale niekoniecznie oznacza, że będziesz miał swobodę wyboru narzędzi.