Senior Infrastructure / DevOps Engineer
⚲ Lisbon
Do uzgodnienia
Wymagania
- Documentation
- Provisioning
- Database Administration (DBA)
- SAN (Storage Area Network)
- Operations
- Linux
- Network
- Security
- Python
- PostgreSQL
Opis stanowiska
We are looking for a Senior Infrastructure / DevOps Engineer with strong hands-on experience in on-premises infrastructure and automation, who is evolving or consolidating their skills in cloud environments. This role is critical to the stability, reliability and continuous improvement of our technology platforms, including mission-critical production systems.
This position combines infrastructure administration, DevOps practices, automation and continuous operations, with mandatory participation in a x7 on-call rotation (L2 support).
Key Responsibilities
On-Premises Infrastructure (Core Focus)
• Operate, maintain and evolve on-premises infrastructure, ensuring stability, performance and high availability.
• Administer virtualization environments based on Proxmox (KVM), including clusters, capacity planning and HA.
• Manage storage systems (SAN/NAS), including performance, redundancy, backup and recovery.
• Perform Linux systems administration, including hardening, patching, security and best practices.
• Work with networking concepts such as IP networking, VLANs, basic routing and firewalls.
DevOps & Automation
• Automate infrastructure provisioning and configuration using Ansible and Infrastructure as Code principles.
• Build, maintain and improve CI/CD pipelines.
• Collaborate closely with development teams to improve deployment, reliability and troubleshooting processes.
• Reduce manual operational tasks through automation and standardization.
Cloud Operations (Growth & Consolidation Phase)
• Operate and support public cloud environments (AWS, Azure), primarily using IaaS services.
• Work in hybrid environments (on-prem + cloud), ensuring integration, security and operational consistency.
• Apply cloud best practices related to cost management, security and reliability.
• Operate and support Kubernetes workloads, with a focus on day-to-day operations, networking, storage and troubleshooting.
Reliability, Operations & On-Call
• Ensure high levels of availability, reliability and operational excellence.
• Participate actively in a 24x7 on-call rotation (L2 support) for production systems.
• Respond to incidents, perform root cause analysis and define corrective and preventive actions.
• Implement and maintain solutions (e.g., Prometheus, Grafana) for monitoring, logging and metrics.
Collaboration & Documentation
• Work closely with engineering and product teams in a multidisciplinary environment.
• Create and maintain clear and practical documentation (runbooks, operational guides, handbooks).
• Support onboarding and promote operational and infrastructure best practices.
Key Requirements
• 7+ years of experience in infrastructure, systems, DevOps or similar roles.
• Strong, hands-on experience with on-premises environments.
• Solid experience with Proxmox (KVM).
• Strong knowledge of Linux systems administration.
• Proven experience using Ansible in production environments.
• Good understanding of networking fundamentals (IP, routing, firewalls, VLANs).
• Strong understanding of storage concepts and high-availability architectures.
• Experience with CI/CD tools (GitLab CI, Jenkins or equivalent).
• Experience with monitoring and observability stacks.
• Scripting skills (Bash, Python or similar).
• Experience with database administration (e.g. PostgreSQL, Oracle).
• Hands-on experience with container technologies such as Docker.
• Functional knowledge of Kubernetes, with an operational focus.
• Experience with public cloud platforms (AWS and/or Azure).
Profile & Mindset
• Strong operational mindset and sense of ownership.
• Comfortable working with critical production systems.
• Able to work autonomously and make technical decisions.
• Clear interest in growing cloud skills, grounded in strong traditional infrastructure experience.
• Practical, resilient and solution-oriented approach.
• Strong communication and collaboration skills.
• Professional proficiency in English (written and spoken).
This position combines infrastructure administration, DevOps practices, automation and continuous operations, with mandatory participation in a x7 on-call rotation (L2 support).
Key Responsibilities
On-Premises Infrastructure (Core Focus)
• Operate, maintain and evolve on-premises infrastructure, ensuring stability, performance and high availability.
• Administer virtualization environments based on Proxmox (KVM), including clusters, capacity planning and HA.
• Manage storage systems (SAN/NAS), including performance, redundancy, backup and recovery.
• Perform Linux systems administration, including hardening, patching, security and best practices.
• Work with networking concepts such as IP networking, VLANs, basic routing and firewalls.
DevOps & Automation
• Automate infrastructure provisioning and configuration using Ansible and Infrastructure as Code principles.
• Build, maintain and improve CI/CD pipelines.
• Collaborate closely with development teams to improve deployment, reliability and troubleshooting processes.
• Reduce manual operational tasks through automation and standardization.
Cloud Operations (Growth & Consolidation Phase)
• Operate and support public cloud environments (AWS, Azure), primarily using IaaS services.
• Work in hybrid environments (on-prem + cloud), ensuring integration, security and operational consistency.
• Apply cloud best practices related to cost management, security and reliability.
• Operate and support Kubernetes workloads, with a focus on day-to-day operations, networking, storage and troubleshooting.
Reliability, Operations & On-Call
• Ensure high levels of availability, reliability and operational excellence.
• Participate actively in a 24x7 on-call rotation (L2 support) for production systems.
• Respond to incidents, perform root cause analysis and define corrective and preventive actions.
• Implement and maintain solutions (e.g., Prometheus, Grafana) for monitoring, logging and metrics.
Collaboration & Documentation
• Work closely with engineering and product teams in a multidisciplinary environment.
• Create and maintain clear and practical documentation (runbooks, operational guides, handbooks).
• Support onboarding and promote operational and infrastructure best practices.
Key Requirements
• 7+ years of experience in infrastructure, systems, DevOps or similar roles.
• Strong, hands-on experience with on-premises environments.
• Solid experience with Proxmox (KVM).
• Strong knowledge of Linux systems administration.
• Proven experience using Ansible in production environments.
• Good understanding of networking fundamentals (IP, routing, firewalls, VLANs).
• Strong understanding of storage concepts and high-availability architectures.
• Experience with CI/CD tools (GitLab CI, Jenkins or equivalent).
• Experience with monitoring and observability stacks.
• Scripting skills (Bash, Python or similar).
• Experience with database administration (e.g. PostgreSQL, Oracle).
• Hands-on experience with container technologies such as Docker.
• Functional knowledge of Kubernetes, with an operational focus.
• Experience with public cloud platforms (AWS and/or Azure).
Profile & Mindset
• Strong operational mindset and sense of ownership.
• Comfortable working with critical production systems.
• Able to work autonomously and make technical decisions.
• Clear interest in growing cloud skills, grounded in strong traditional infrastructure experience.
• Practical, resilient and solution-oriented approach.
• Strong communication and collaboration skills.
• Professional proficiency in English (written and spoken).
🔍 Dekoder Ogłoszenia
🔴
evolving or consolidating their skills in cloud environments
Oczekuje się, że kandydat będzie uczył się chmury od podstaw lub będzie miał minimalne doświadczenie, a nie będzie ekspertem.
🔴
critical to the stability, reliability and continuous improvement of our technology platforms
Oznacza to, że będziesz odpowiedzialny za rozwiązywanie problemów, gdy coś pójdzie nie tak, a niekoniecznie za proaktywne wprowadzanie innowacji.
🔴
mandatory participation in a x7 on-call rotation (L2 support)
Będziesz musiał być dostępny 24/7, aby rozwiązywać problemy na drugim poziomie wsparcia, co może być bardzo obciążające.
🔴
On-Premises Infrastructure (Core Focus)
Pomimo wzmianki o chmurze, główny nacisk i odpowiedzialność spoczywa na utrzymaniu i rozwijaniu istniejącej infrastruktury lokalnej.
🟡
Collaborate closely with development teams
Może oznaczać, że będziesz musiał często tłumaczyć złożone problemy infrastrukturalne zespołom deweloperskim lub pomagać w ich rozwiązywaniu.