JustJoin.IT Praca zdalna Senior

Network Reliability Engineer - Senior

Margo

⚲ Warsaw

33 600 - 42 000 PLN netto (B2B)

Wymagania

  • Networking
  • GPU
  • Python
  • GitLab
  • MariaDB
  • Ansible
  • Linux (Debian)
  • Grafana
  • Prometheus
  • Ubuntu

Opis stanowiska

Growth is driving Client to strengthen its SRE team to support and scale its production environments.
Your mission will be to build and maintain reliable, observable, and secure infrastructure in order to ensure optimal service availability for our customers around the world.
#HPC #AI #GPU #CLUSTERS

!!! Client is in CALIFORNIA, USA !!!
Working Hours - More or less like strart at 18.00 CEST
Long term Project of minimum a year

YOUR DAILY ROUTINE:
- Build a large AI infrastructure with monitoring, diagnosis, and remediation of production incidents- Troubleshoot high-impact production issues in collaboration with other engineering teams
- Participate in an on-call rotation to handle incidents and ensure service continuity
- Implement and maintain observability solutions to monitor AI infrastructure and application health
- Contribute to AI infrastructure lifecycle management across different environments and countries
- Promote and apply best practices in terms of stability, resiliency, scalability, and security
- Maintain clear technical documentation for tools and procedures
- Contribute to system and tool evolution based on production feedback
- Collaborate closely with development teams to ensure infrastructure readiness- Participate in team rituals and knowledge-sharing initiatives

SOFTSKILLS :
- Proactive and solution-oriented mindset
- Passion for automation and continuous improvement
- Strong collaboration and communication skills
- Ability to work independently and in a team
- Willingness to mentor and share knowledge

HARDSKILLS :
- Experience with Go or Python
- Strong scripting skills (Bash, Python)
- Hands-on experience with Linux systems (Ubuntu/Debian)
- Preferred hands-on experience with GPU & HPC infrastructure
- Knowledge of networking (LAN/VLAN TCP/IP, DNS, BGP, load-balancing, IPv6, etc.)
- Boot systems like PXE
- Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.)
- Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.)
- Experience managing relational databases (MariaDB)
- Understanding of CI/CD pipelines (GitLab)
- Comfortable with English (written and spoken)

🔍 Dekoder Ogłoszenia

🔴
Growth is driving Client to strengthen its SRE team to support and scale its production environments.
Firma się rozwija, co może oznaczać zwiększone obciążenie pracą i potrzebę szybkiego reagowania na problemy.
🔴
Working Hours - More or less like strart at 18.00 CEST
Praca będzie odbywać się w godzinach wieczornych i nocnych czasu polskiego, co może być uciążliwe.
🟡
Long term Project of minimum a year
Projekt jest długoterminowy, ale nie gwarantuje zatrudnienia po jego zakończeniu.
🔴
Troubleshoot high-impact production issues in collaboration with other engineering teams
Będziesz musiał rozwiązywać poważne problemy produkcyjne, co może wiązać się ze stresem i presją czasu.
🔴
Participate in an on-call rotation to handle incidents and ensure service continuity
Będziesz musiał być dostępny poza standardowymi godzinami pracy, aby reagować na awarie.