IT Operations Engineer (Kafka Platform Support)
⚲ Warszawa
18 480 - 21 840 PLN netto (B2B)
Wymagania
- Apache Kafka
Opis stanowiska
In Cyclad we work with top international IT companies in order to boost their potential in delivering outstanding, cutting-edge technologies that shape the world of the future. We are seeking an experienced IT Operations Engineer (Kafka Platform Support) to join a global platform team. In this role, you will be responsible for the operation and support of a Kafka-based messaging platform in production, ensuring its availability, stability, and performance.
Project information:
• Office location: Poland
• Work mode: 100% Remote
• Office location: Remote from Poland
• Budget: 110- 130 PLN net/ h- B2B
• Project length: Long-term
• Only candidates with citizenship in the European Union and residence in Poland
• Start date: ASAP (depending on candidate’s availability)
Project scope:
• Operate, monitor, and maintain a Kafka-based messaging platform in a production environment
• Ensure platform availability, stability, and performance in line with operational SLAs
• Monitor system health using logs, metrics, and alerting tools
• Perform routine operational checks and maintenance activities
• Handle incidents and service requests via ticketing systems and support channels
• Troubleshoot issues across Kafka components (brokers, producers, consumers, integrations)
• Analyze logs, metrics, and system behavior to identify root causes of incidents
• Execute operational procedures based on runbooks and standard operating procedures (SOPs)
• Perform configuration changes (topics, access controls, settings) following established processes
• Maintain and continuously improve operational documentation and runbooks
• Act as a primary support contact for internal users of the Kafka platform
• Provide technical support via collaboration tools (e.g., Slack, Teams)
• Assist users with troubleshooting and best practices
• Translate user-reported issues into actionable insights for technical teams
• Collaborate closely with engineering and platform teams to resolve incidents
• Participate in incident reviews and post-mortems
• Identify recurring operational issues and suggest improvements or automation opportunities
• Contribute to improving platform usability and support processes
•
Competence demands:
• Minimum 3–6 years of experience in IT operations, production support, or platform support roles
• Hands-on experience with Apache Kafka or similar event streaming platforms
• Strong understanding of distributed systems (partitioning, replication, scaling)
• Strong troubleshooting skills in production IT environments
• Experience with monitoring, logging, and alerting tools (e.g., Grafana, Prometheus)
• Knowledge of Git and version control practices
• Familiarity with GitLab CI/CD and working with existing pipelines
• Experience working with incident management processes and support tools
• Experience working with runbooks, SOPs, and structured operational environments
• Strong communication skills and ability to explain technical issues clearly
• Experience working with internal customers and cross-functional teams
• Customer-focused mindset with a proactive approach to support
• Fluency in English (written and spoken)
Nice to have:
• Experience with AWS or other cloud platforms
• Familiarity with Kubernetes and containerized environments
• Experience with monitoring tools such as Grafana and Prometheus
• Understanding of automation in IT operations and platform support environments
We offer:
• Remote working model
• Full-time job agreement based on b2b
• Private medical care with dental care (covering 70% of costs)
• Multisport card (also for an accompanying person)
• Life insurance
Recruitment process:
• Introductory call
• Technical interview
• Final decision
Project information:
• Office location: Poland
• Work mode: 100% Remote
• Office location: Remote from Poland
• Budget: 110- 130 PLN net/ h- B2B
• Project length: Long-term
• Only candidates with citizenship in the European Union and residence in Poland
• Start date: ASAP (depending on candidate’s availability)
Project scope:
• Operate, monitor, and maintain a Kafka-based messaging platform in a production environment
• Ensure platform availability, stability, and performance in line with operational SLAs
• Monitor system health using logs, metrics, and alerting tools
• Perform routine operational checks and maintenance activities
• Handle incidents and service requests via ticketing systems and support channels
• Troubleshoot issues across Kafka components (brokers, producers, consumers, integrations)
• Analyze logs, metrics, and system behavior to identify root causes of incidents
• Execute operational procedures based on runbooks and standard operating procedures (SOPs)
• Perform configuration changes (topics, access controls, settings) following established processes
• Maintain and continuously improve operational documentation and runbooks
• Act as a primary support contact for internal users of the Kafka platform
• Provide technical support via collaboration tools (e.g., Slack, Teams)
• Assist users with troubleshooting and best practices
• Translate user-reported issues into actionable insights for technical teams
• Collaborate closely with engineering and platform teams to resolve incidents
• Participate in incident reviews and post-mortems
• Identify recurring operational issues and suggest improvements or automation opportunities
• Contribute to improving platform usability and support processes
•
Competence demands:
• Minimum 3–6 years of experience in IT operations, production support, or platform support roles
• Hands-on experience with Apache Kafka or similar event streaming platforms
• Strong understanding of distributed systems (partitioning, replication, scaling)
• Strong troubleshooting skills in production IT environments
• Experience with monitoring, logging, and alerting tools (e.g., Grafana, Prometheus)
• Knowledge of Git and version control practices
• Familiarity with GitLab CI/CD and working with existing pipelines
• Experience working with incident management processes and support tools
• Experience working with runbooks, SOPs, and structured operational environments
• Strong communication skills and ability to explain technical issues clearly
• Experience working with internal customers and cross-functional teams
• Customer-focused mindset with a proactive approach to support
• Fluency in English (written and spoken)
Nice to have:
• Experience with AWS or other cloud platforms
• Familiarity with Kubernetes and containerized environments
• Experience with monitoring tools such as Grafana and Prometheus
• Understanding of automation in IT operations and platform support environments
We offer:
• Remote working model
• Full-time job agreement based on b2b
• Private medical care with dental care (covering 70% of costs)
• Multisport card (also for an accompanying person)
• Life insurance
Recruitment process:
• Introductory call
• Technical interview
• Final decision
🔍 Dekoder Ogłoszenia
🟡
boost their potential in delivering outstanding, cutting-edge technologies that shape the world of the future
To ogólne, marketingowe stwierdzenie, które nie precyzuje konkretnych technologii ani roli kandydata.
🔴
global platform team
Może oznaczać pracę z ludźmi z różnych stref czasowych, co może wpływać na elastyczność godzin pracy lub komunikację.
🔴
ensure its availability, stability, and performance
Podstawowy obowiązek inżyniera operacyjnego, ale w praktyce może oznaczać presję na utrzymanie ciągłości działania za wszelką cenę.
🔴
in line with operational SLAs
SLA (Service Level Agreement) może narzucać bardzo rygorystyczne terminy reakcji i rozwiązywania problemów, co może prowadzić do stresu i pracy pod presją.
🔴
ASAP (depending on candidate’s availability)
Oznacza, że projekt chce zacząć jak najszybciej, co może sugerować pilną potrzebę obsadzenia stanowiska i potencjalnie presję na szybkie wdrożenie.