Senior Data Engineer
⚲ Białystok, Olsztyn, Gdańsk, Szczecin, Poznań, Warszawa, Lublin, Wrocław, Kraków, Łódź
Do uzgodnienia
Wymagania
- Git
- CI/CD
- ETL
- Data
- Apache
- GCP
- BigQuery
- SQL
Opis stanowiska
Key Responsibilities
•
Architect & Model: Design, implement, and maintain scalable data models using Star, Snowflake, and Data Vault methodologies.
• Optimize BigQuery: Write, refactor, and optimize complex, highly efficient SQL queries with a strong focus on cost reduction, execution time, and concurrency.
• Build ELT/ETL Pipelines: Design, test, and deploy robust data ingestion and transformation pipelines (primarily using GCP Data Fusion / Apache Airflow).
• Integrate Diverse Sources: Connect and wrangle data from various sources, including RESTful/SOAP APIs and SFTP servers, handling formats like JSON, CSV, and XML.
• Drive DataOps Culture: Collaborate within an Agile framework using Git and CI/CD pipelines to ensure automated testing and continuous delivery of data solutions.
Technical Requirements
• 3+ years of expert-level SQL & BigQuery hands-on experience (complex optimization, cost control, data integrity).
• 3+ years of experience in ELT/ETL pipeline development using GCP Data Fusion (highly preferred) or Apache Airflow.
• Strong expertise in Data Modeling & Architecture – deep understanding of Kimball/Inmon concepts and hands-on experience with Data Vault (major plus).
• API & Data Ingestion proficiency – proven track record of ingesting and parsing JSON, XML, and CSV data from REST/SOAP endpoints.
• Git & CI/CD fluency – practical knowledge of version control and continuous testing/delivery tools in cloud environments.
•
Architect & Model: Design, implement, and maintain scalable data models using Star, Snowflake, and Data Vault methodologies.
• Optimize BigQuery: Write, refactor, and optimize complex, highly efficient SQL queries with a strong focus on cost reduction, execution time, and concurrency.
• Build ELT/ETL Pipelines: Design, test, and deploy robust data ingestion and transformation pipelines (primarily using GCP Data Fusion / Apache Airflow).
• Integrate Diverse Sources: Connect and wrangle data from various sources, including RESTful/SOAP APIs and SFTP servers, handling formats like JSON, CSV, and XML.
• Drive DataOps Culture: Collaborate within an Agile framework using Git and CI/CD pipelines to ensure automated testing and continuous delivery of data solutions.
Technical Requirements
• 3+ years of expert-level SQL & BigQuery hands-on experience (complex optimization, cost control, data integrity).
• 3+ years of experience in ELT/ETL pipeline development using GCP Data Fusion (highly preferred) or Apache Airflow.
• Strong expertise in Data Modeling & Architecture – deep understanding of Kimball/Inmon concepts and hands-on experience with Data Vault (major plus).
• API & Data Ingestion proficiency – proven track record of ingesting and parsing JSON, XML, and CSV data from REST/SOAP endpoints.
• Git & CI/CD fluency – practical knowledge of version control and continuous testing/delivery tools in cloud environments.
🔍 Dekoder Ogłoszenia
🔴
Architect & Model: Design, implement, and maintain scalable data models using Star, Snowflake, and Data Vault methodologies.
Oczekuje się, że będziesz projektować i wdrażać modele danych od podstaw, a nie tylko pracować z istniejącymi.
🔴
Optimize BigQuery: Write, refactor, and optimize complex, highly efficient SQL queries with a strong focus on cost reduction, execution time, and concurrency.
Duży nacisk na optymalizację istniejących zapytań, co może oznaczać pracę z nieoptymalnym kodem i konieczność jego gruntownego przepisywania.
🔴
Drive DataOps Culture
Może oznaczać konieczność aktywnego promowania i wdrażania praktyk DataOps, a nie tylko ich stosowania.
🔴
3+ years of expert-level SQL & BigQuery hands-on experience
Wymagane jest bardzo głębokie i praktyczne doświadczenie, nie tylko teoretyczna wiedza.
🔴
GCP Data Fusion (highly preferred) or Apache Airflow
Chociaż Airflow jest wymieniony, preferencja dla Data Fusion sugeruje, że to właśnie to narzędzie będzie głównym obszarem pracy.