Data Governance & Metadata Engineer
⚲ Kraków, Podgórze Duchackie
Do uzgodnienia
Wymagania
- Java
- Python
- pandas
- Linux
- REST APIs
- SQL
- Databricks
- Unity Catalog
- Microsoft Purview
- Git
- Delta Lake
- Azure services
- Collibra
Opis stanowiska
Nasze wymagania:
Hands‑on experience with metadata platforms and governance tooling.
Strong understanding of data modelling, metadata architectures, lineage, and catalog concepts.
Proficiency in Java and Python (mandatory pandas expertise).
Solid Linux skills.
Experience with REST APIs, SQL, and metadata extraction.
Practical experience with Databricks and Unity Catalog.
Familiarity with Microsoft Purview concepts (scanning, classifications, lineage).
Knowledge of data governance and data quality frameworks.
Strong documentation and communication skills.
Experience with Git and versioning workflows.
Mile widziane:
Advanced hands‑on experience with Databricks (Spark, Delta Lake, Unity Catalog, jobs).
Experience integrating Purview and Collibra in hybrid governance setups.
Azure services and orchestration tools experience.
Metadata scanning, MDM, or lineage tooling experience.
Collibra workflow development.
Understanding of data security, classification, and access control models.
Familiarity with industry lineage standards and open metadata approaches.
Zakres obowiązków:
Metadata Engineering & Automation
• Model and document data structures, metadata relationships, and end to end lineage across data platforms (including lakehouse architectures).
• Configure and maintain metadata workflows, connectors, and governance assets.
• Build automations and integrations in Java and Python (strong pandas expertise required).
• Integrate technical metadata from Databricks (Spark, Delta Lake, Unity Catalog) and Microsoft Purview into governance tools.
• Operate and troubleshoot metadata services on Linux (logs, services, deployments).
Data Governance
• Maintain data domains, dictionaries, glossaries, classification models, and stewardship structures.
• Define and enforce metadata standards and data quality rules across analytical platforms.
• Support impact analysis, remediation, and governance alignment for data products and lakehouse use cases.
Collibra / Purview
• Maintain metadata, lineage, stewardship models, and workflows in Collibra.
• Integrate Collibra with technical metadata sources (e.g. Databricks, Unity Catalog, Microsoft Purview, SQL engines) via APIs, scanners, or pipelines.
• Align governance models between Collibra and Purview (glossaries, classifications, lineage where applicable).
• Provide onboarding, training, and high quality documentation.
Cross Functional Work
• Collaborate with data engineers, platform teams, and product owners to ensure consistent governance standards in Databricks, Unity Catalog, and Purview.
• Assess risks, impacts, and compliance aspects in data related projects.
• Translate technical platform concepts (Spark, lakehouse, catalogs, semantic layers) into clear governance artefacts.
Oferujemy:
Work closely with inspiring, supportive and engaged colleagues from more than 80 different countries.
Practice your talents in a highly professional international environment.
Join a learning and development environment with an emphasis on knowledge sharing and training.
Competitive salary and comprehensive benefits.
Hands‑on experience with metadata platforms and governance tooling.
Strong understanding of data modelling, metadata architectures, lineage, and catalog concepts.
Proficiency in Java and Python (mandatory pandas expertise).
Solid Linux skills.
Experience with REST APIs, SQL, and metadata extraction.
Practical experience with Databricks and Unity Catalog.
Familiarity with Microsoft Purview concepts (scanning, classifications, lineage).
Knowledge of data governance and data quality frameworks.
Strong documentation and communication skills.
Experience with Git and versioning workflows.
Mile widziane:
Advanced hands‑on experience with Databricks (Spark, Delta Lake, Unity Catalog, jobs).
Experience integrating Purview and Collibra in hybrid governance setups.
Azure services and orchestration tools experience.
Metadata scanning, MDM, or lineage tooling experience.
Collibra workflow development.
Understanding of data security, classification, and access control models.
Familiarity with industry lineage standards and open metadata approaches.
Zakres obowiązków:
Metadata Engineering & Automation
• Model and document data structures, metadata relationships, and end to end lineage across data platforms (including lakehouse architectures).
• Configure and maintain metadata workflows, connectors, and governance assets.
• Build automations and integrations in Java and Python (strong pandas expertise required).
• Integrate technical metadata from Databricks (Spark, Delta Lake, Unity Catalog) and Microsoft Purview into governance tools.
• Operate and troubleshoot metadata services on Linux (logs, services, deployments).
Data Governance
• Maintain data domains, dictionaries, glossaries, classification models, and stewardship structures.
• Define and enforce metadata standards and data quality rules across analytical platforms.
• Support impact analysis, remediation, and governance alignment for data products and lakehouse use cases.
Collibra / Purview
• Maintain metadata, lineage, stewardship models, and workflows in Collibra.
• Integrate Collibra with technical metadata sources (e.g. Databricks, Unity Catalog, Microsoft Purview, SQL engines) via APIs, scanners, or pipelines.
• Align governance models between Collibra and Purview (glossaries, classifications, lineage where applicable).
• Provide onboarding, training, and high quality documentation.
Cross Functional Work
• Collaborate with data engineers, platform teams, and product owners to ensure consistent governance standards in Databricks, Unity Catalog, and Purview.
• Assess risks, impacts, and compliance aspects in data related projects.
• Translate technical platform concepts (Spark, lakehouse, catalogs, semantic layers) into clear governance artefacts.
Oferujemy:
Work closely with inspiring, supportive and engaged colleagues from more than 80 different countries.
Practice your talents in a highly professional international environment.
Join a learning and development environment with an emphasis on knowledge sharing and training.
Competitive salary and comprehensive benefits.
🔍 Dekoder Ogłoszenia
🔴
Hands‑on experience with metadata platforms and governance tooling.
Oczekuje się, że kandydat będzie samodzielnie i praktycznie obsługiwał narzędzia do zarządzania metadanymi i ładu danych, a nie tylko teoretycznie je znał.
🔴
Proficiency in Java and Python (mandatory pandas expertise).
Umiejętność programowania w tych językach jest kluczowa, a znajomość biblioteki pandas jest absolutnie wymagana i będzie sprawdzana.
🔴
Model and document data structures, metadata relationships, and end to end lineage across data platforms (including lakehouse architectures).
Praca może wymagać tworzenia szczegółowej dokumentacji i modeli, co może być czasochłonne i wymagać dużej precyzji.
🔴
Build automations and integrations in Java and Python (strong pandas expertise required).
Kandydat będzie odpowiedzialny za tworzenie skryptów automatyzujących procesy, co wymaga nie tylko znajomości języków, ale też umiejętności rozwiązywania problemów.
🟡
Experience integrating Purview and Collibra in hybrid governance setups.
Wymaga to znajomości specyficznych narzędzi i umiejętności ich integracji w złożonych, hybrydowych środowiskach.