JustJoin.IT Praca zdalna Mid

IT Operations Engineer (Kafka Platform Support)

Cyclad

⚲ Warszawa

18 480 - 21 840 PLN netto (B2B)

Wymagania

  • Apache Kafka

Opis stanowiska

In Cyclad we work with top international IT companies in order to boost their potential in delivering outstanding, cutting-edge technologies that shape the world of the future. We are seeking an experienced IT Operations Engineer (Kafka Platform Support) to join a global platform team. In this role, you will be responsible for the operation and support of a Kafka-based messaging platform in production, ensuring its availability, stability, and performance. 

Project information: 
• Office location: Poland 
• Work mode: 100% Remote 
• Office location: Remote from Poland 
• Budget: 110- 130 PLN  net/ h- B2B 
• Project length: Long-term  
• Only candidates with citizenship in the European Union and residence in Poland 
• Start date: ASAP (depending on candidate’s availability) 
 
 
Project scope: 
• Operate, monitor, and maintain a Kafka-based messaging platform in a production environment  
• Ensure platform availability, stability, and performance in line with operational SLAs  
• Monitor system health using logs, metrics, and alerting tools  
• Perform routine operational checks and maintenance activities  
• Handle incidents and service requests via ticketing systems and support channels  
• Troubleshoot issues across Kafka components (brokers, producers, consumers, integrations)  
• Analyze logs, metrics, and system behavior to identify root causes of incidents  
• Execute operational procedures based on runbooks and standard operating procedures (SOPs)  
• Perform configuration changes (topics, access controls, settings) following established processes  
• Maintain and continuously improve operational documentation and runbooks  
• Act as a primary support contact for internal users of the Kafka platform  
• Provide technical support via collaboration tools (e.g., Slack, Teams)  
• Assist users with troubleshooting and best practices  
• Translate user-reported issues into actionable insights for technical teams  
• Collaborate closely with engineering and platform teams to resolve incidents  
• Participate in incident reviews and post-mortems  
• Identify recurring operational issues and suggest improvements or automation opportunities  
• Contribute to improving platform usability and support processes 

Competence demands: 
• Minimum 3–6 years of experience in IT operations, production support, or platform support roles  
• Hands-on experience with Apache Kafka or similar event streaming platforms  
• Strong understanding of distributed systems (partitioning, replication, scaling)  
• Strong troubleshooting skills in production IT environments  
• Experience with monitoring, logging, and alerting tools (e.g., Grafana, Prometheus)  
• Knowledge of Git and version control practices  
• Familiarity with GitLab CI/CD and working with existing pipelines  
• Experience working with incident management processes and support tools  
• Experience working with runbooks, SOPs, and structured operational environments 
• Strong communication skills and ability to explain technical issues clearly  
• Experience working with internal customers and cross-functional teams  
• Customer-focused mindset with a proactive approach to support  
• Fluency in English (written and spoken) 
Nice to have: 
• Experience with AWS or other cloud platforms  
• Familiarity with Kubernetes and containerized environments  
• Experience with monitoring tools such as Grafana and Prometheus  
• Understanding of automation in IT operations and platform support environments 
We offer: 
• Remote working model 
• Full-time job agreement based on b2b 
• Private medical care with dental care (covering 70% of costs)  
• Multisport card (also for an accompanying person) 
• Life insurance 
Recruitment process:  
• Introductory call  
• Technical interview  
• Final decision

🔍 Dekoder Ogłoszenia

🟡
boost their potential in delivering outstanding, cutting-edge technologies that shape the world of the future
To ogólne, marketingowe stwierdzenie, które nie precyzuje konkretnych technologii ani roli kandydata.
🔴
global platform team
Może oznaczać pracę z ludźmi z różnych stref czasowych, co może wpływać na elastyczność godzin pracy lub komunikację.
🔴
ensure its availability, stability, and performance
Podstawowy obowiązek inżyniera operacyjnego, ale w praktyce może oznaczać presję na utrzymanie ciągłości działania za wszelką cenę.
🔴
in line with operational SLAs
SLA (Service Level Agreement) może narzucać bardzo rygorystyczne terminy reakcji i rozwiązywania problemów, co może prowadzić do stresu i pracy pod presją.
🔴
ASAP (depending on candidate’s availability)
Oznacza, że projekt chce zacząć jak najszybciej, co może sugerować pilną potrzebę obsadzenia stanowiska i potencjalnie presję na szybkie wdrożenie.