T-Hub - Data Engineer with Java
⚲ Warszawa
Do uzgodnienia
Wymagania
- Data
- BigQuery
- IAM
- Apache
- Java
Opis stanowiska
WHAT TASKS AWAIT YOU?
• Design and implement batch and streaming pipelines; select Beam/Spark/Flink based on requirements.
• Data transformation ops for analytics and downstream services.
• Optimize jobs (performance, cost) and implement robust testing and observability.
• Build and manage Airflow orchestration.
• Troubleshoot production issues, lead incident resolution, and drive root-cause fixes.
• Contribute to team standards, templates, and reusable libraries.
• Using Java and Apache Beam as the main programming language and library for the streaming pipelines.
• Designs and ships a new pipeline or major refactor with measurable latency/cost improvements
• Hands-on, close cooperation with the Data Engineering Lead, resulting in quick onboarding and successful code base understanding
• Introduces or improves testing/observability patterns adopted by the team.
WHAT SKILLS WILL BE APPRECIATED?
• Strong experience with Google Cloud Platform, including Dataflow, BigQuery performance optimization (partitioning and clustering), Pub/Sub, and Dataproc, as well as Cloud Monitoring and Logging.
• Good understanding of IAM, VPC fundamentals, and service accounts.
• Deep expertise in Apache Beam, including advanced transformations, windowing and triggering, side inputs and outputs, as well as state and timers for streaming pipelines.
• Experience in optimizing resource sizing and performance on Dataflow, including high-throughput and low-latency pipelines, custom IOs, and debugging at runner level.
• Strong experience with Apache Spark, including optimization of joins, shuffles, and partitioning, handling schema evolution, and debugging jobs using metrics.
• Familiarity with Apache Flink and ability to implement streaming pipelines.
• Knowledge of Docker and Kubernetes for job packaging, especially in Spark and Flink environments.
• • Advanced SQL skills, including complex queries, performance tuning, and incremental data processing patterns.
• Ability to define and enforce data quality checks and acceptance criteria.
• Hands-on experience with Airflow, including building and maintaining complex DAGs.
• Proven ownership of production-grade data pipelines with focus on reliability and scalability.
• Extensive proficiency in Java at an expert level, including concurrency, immutability, and performance profiling.
• Experience in building reusable libraries and enforcing code quality, testing standards, and review practices.
• Strong skills in system design and API design, including performance-critical code.
• Experience in designing and architecting batch and streaming data systems, including selecting technologies and defining standards and guardrails.
• Ability to provide technical leadership, drive engineering strategy, and ensure scalability, security, and cost efficiency.
• Proven track record of building platform components such as libraries, templates, governance frameworks, and data quality solutions.
• Experience with CI/CD for data (e.g. GitHub Actions, Cloud Build) and infrastructure-as-code tools like Terraform.
• Minimum of 6 years of experience in data engineering with demonstrated architectural leadership.
• Track record of delivering large-scale data systems and influencing cross-team outcomes.
• Ability to set architectural direction, ensure compliance and security, and drive measurable platform improvements.
• Willingness to travel at least four times per year.
OUR OFFER FOR YOU
Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.
You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.
• Design and implement batch and streaming pipelines; select Beam/Spark/Flink based on requirements.
• Data transformation ops for analytics and downstream services.
• Optimize jobs (performance, cost) and implement robust testing and observability.
• Build and manage Airflow orchestration.
• Troubleshoot production issues, lead incident resolution, and drive root-cause fixes.
• Contribute to team standards, templates, and reusable libraries.
• Using Java and Apache Beam as the main programming language and library for the streaming pipelines.
• Designs and ships a new pipeline or major refactor with measurable latency/cost improvements
• Hands-on, close cooperation with the Data Engineering Lead, resulting in quick onboarding and successful code base understanding
• Introduces or improves testing/observability patterns adopted by the team.
WHAT SKILLS WILL BE APPRECIATED?
• Strong experience with Google Cloud Platform, including Dataflow, BigQuery performance optimization (partitioning and clustering), Pub/Sub, and Dataproc, as well as Cloud Monitoring and Logging.
• Good understanding of IAM, VPC fundamentals, and service accounts.
• Deep expertise in Apache Beam, including advanced transformations, windowing and triggering, side inputs and outputs, as well as state and timers for streaming pipelines.
• Experience in optimizing resource sizing and performance on Dataflow, including high-throughput and low-latency pipelines, custom IOs, and debugging at runner level.
• Strong experience with Apache Spark, including optimization of joins, shuffles, and partitioning, handling schema evolution, and debugging jobs using metrics.
• Familiarity with Apache Flink and ability to implement streaming pipelines.
• Knowledge of Docker and Kubernetes for job packaging, especially in Spark and Flink environments.
• • Advanced SQL skills, including complex queries, performance tuning, and incremental data processing patterns.
• Ability to define and enforce data quality checks and acceptance criteria.
• Hands-on experience with Airflow, including building and maintaining complex DAGs.
• Proven ownership of production-grade data pipelines with focus on reliability and scalability.
• Extensive proficiency in Java at an expert level, including concurrency, immutability, and performance profiling.
• Experience in building reusable libraries and enforcing code quality, testing standards, and review practices.
• Strong skills in system design and API design, including performance-critical code.
• Experience in designing and architecting batch and streaming data systems, including selecting technologies and defining standards and guardrails.
• Ability to provide technical leadership, drive engineering strategy, and ensure scalability, security, and cost efficiency.
• Proven track record of building platform components such as libraries, templates, governance frameworks, and data quality solutions.
• Experience with CI/CD for data (e.g. GitHub Actions, Cloud Build) and infrastructure-as-code tools like Terraform.
• Minimum of 6 years of experience in data engineering with demonstrated architectural leadership.
• Track record of delivering large-scale data systems and influencing cross-team outcomes.
• Ability to set architectural direction, ensure compliance and security, and drive measurable platform improvements.
• Willingness to travel at least four times per year.
OUR OFFER FOR YOU
Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.
You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.
🔍 Dekoder Ogłoszenia
🔴
Designs and ships a new pipeline or major refactor with measurable latency/cost improvements
Oczekuje się, że kandydat będzie samodzielnie projektował i wdrażał nowe rozwiązania, a nie tylko pracował nad istniejącymi.
🔴
Troubleshoot production issues, lead incident resolution, and drive root-cause fixes.
Duża część pracy będzie polegać na gaszeniu pożarów i rozwiązywaniu problemów w środowisku produkcyjnym, co może być stresujące.
🟡
Contribute to team standards, templates, and reusable libraries.
Oprócz bieżących zadań, oczekuje się aktywnego udziału w tworzeniu i utrzymaniu wewnętrznych narzędzi i dobrych praktyk zespołu.
🟡
Hands-on, close cooperation with the Data Engineering Lead, resulting in quick onboarding and successful code base understanding
Może oznaczać, że proces onboardingu będzie intensywny i wymagał będzie szybkiego przyswojenia dużej ilości informacji, ale też bliskiej współpracy z liderem.
🟡
Introduces or improves testing/observability patterns adopted by the team.
Oczekuje się proaktywności w zakresie poprawy jakości kodu i monitorowania, co może wykraczać poza standardowe obowiązki inżyniera danych.