Databricks Data Engineer
⚲ Warszawa, Gdańsk, Kraków
30 240 - 36 960 PLN netto (B2B)
Wymagania
- Python
- SQL
- Databricks
Opis stanowiska
Data Engineer with Databricks
We are looking for a Data Engineer with strong Databricks experience to join a data & AI project focused on building a modern, scalable data platform supporting analytics, machine learning and AI use cases.
What you’ll do
• Design, build and maintain scalable data pipelines using Databricks and Apache Spark.
• Integrate data from multiple sources and build reliable, reusable data flows.
• Develop and optimize data processing solutions for large volumes of data.
• Work closely with Data Scientists, Analysts and business stakeholders to deliver data products supporting ML, BI and analytics.
• Contribute to the development of a unified data platform and high-quality data layer.
• Ensure data pipelines are reliable, scalable, performant and easy to maintain.
• Optimize data processing and infrastructure with a focus on performance, scalability and cost efficiency.
• Implement and maintain data quality, monitoring and data engineering best practices.
• Support the development and evolution of modern lakehouse and cloud data architectures.
Your experience
•
Hands-on experience with Databricks is required.
• Strong experience in Data Engineering and building production-grade data pipelines.
• Strong SQL and PySpark / Apache Spark skills.
• Experience working with large datasets and distributed data processing.
• Good understanding of modern data platform, lakehouse and data architecture concepts.
• Experience with cloud environments such as Azure, AWS or GCP.
• Experience with Delta Lake and data orchestration tools is a strong advantage.
• Experience working with different data sources, formats and integration patterns.
• Ability to work closely with Data Scientists, Analysts and other technical stakeholders and understand their data requirements.
• Strong problem-solving skills and a pragmatic approach to data engineering.
Relevant experience
You should have hands-on experience in one or more of the following areas:
•
Data Platform & Lakehouse Engineering – building scalable platforms for analytics, reporting, ML and AI workloads.
• Data Integration & Transformation – integrating structured and unstructured data from multiple source systems into reliable, reusable data pipelines.
• Data Quality & Governance – implementing processes and frameworks for data quality, monitoring, lineage and governance.
Tech stack
Databricks, Apache Spark, PySpark, SQL, Delta Lake, Cloud (Azure / AWS / GCP)
Why join?
You’ll be part of a large-scale data & AI transformation, building the data foundations that power analytics, machine learning and AI use cases while working with modern Databricks and lakehouse architecture.
We are looking for a Data Engineer with strong Databricks experience to join a data & AI project focused on building a modern, scalable data platform supporting analytics, machine learning and AI use cases.
What you’ll do
• Design, build and maintain scalable data pipelines using Databricks and Apache Spark.
• Integrate data from multiple sources and build reliable, reusable data flows.
• Develop and optimize data processing solutions for large volumes of data.
• Work closely with Data Scientists, Analysts and business stakeholders to deliver data products supporting ML, BI and analytics.
• Contribute to the development of a unified data platform and high-quality data layer.
• Ensure data pipelines are reliable, scalable, performant and easy to maintain.
• Optimize data processing and infrastructure with a focus on performance, scalability and cost efficiency.
• Implement and maintain data quality, monitoring and data engineering best practices.
• Support the development and evolution of modern lakehouse and cloud data architectures.
Your experience
•
Hands-on experience with Databricks is required.
• Strong experience in Data Engineering and building production-grade data pipelines.
• Strong SQL and PySpark / Apache Spark skills.
• Experience working with large datasets and distributed data processing.
• Good understanding of modern data platform, lakehouse and data architecture concepts.
• Experience with cloud environments such as Azure, AWS or GCP.
• Experience with Delta Lake and data orchestration tools is a strong advantage.
• Experience working with different data sources, formats and integration patterns.
• Ability to work closely with Data Scientists, Analysts and other technical stakeholders and understand their data requirements.
• Strong problem-solving skills and a pragmatic approach to data engineering.
Relevant experience
You should have hands-on experience in one or more of the following areas:
•
Data Platform & Lakehouse Engineering – building scalable platforms for analytics, reporting, ML and AI workloads.
• Data Integration & Transformation – integrating structured and unstructured data from multiple source systems into reliable, reusable data pipelines.
• Data Quality & Governance – implementing processes and frameworks for data quality, monitoring, lineage and governance.
Tech stack
Databricks, Apache Spark, PySpark, SQL, Delta Lake, Cloud (Azure / AWS / GCP)
Why join?
You’ll be part of a large-scale data & AI transformation, building the data foundations that power analytics, machine learning and AI use cases while working with modern Databricks and lakehouse architecture.
🔍 Dekoder Ogłoszenia
🔴
building a modern, scalable data platform supporting analytics, machine learning and AI use cases
Projekt może być na wczesnym etapie rozwoju, wymagać budowania od podstaw i nie mieć jeszcze ugruntowanych procesów.
🟡
Work closely with Data Scientists, Analysts and business stakeholders
Może oznaczać konieczność częstych spotkań i tłumaczenia złożonych zagadnień technicznych na język biznesowy.
🔴
Contribute to the development of a unified data platform
Może oznaczać, że platforma jest w fazie tworzenia i wymagać będzie znaczącego wkładu w jej kształtowanie, a nie tylko utrzymanie istniejących rozwiązań.
🟡
Optimize data processing and infrastructure with a focus on performance, scalability and cost efficiency
Oprócz optymalizacji wydajności, będziesz odpowiedzialny za kontrolę kosztów infrastruktury chmurowej.
🔴
Support the development and evolution of modern lakehouse and cloud data architectures
Może oznaczać, że będziesz pracować z nowymi technologiami, które nie są jeszcze w pełni dojrzałe lub stabilne.