Senior SRE
⚲ Kraków
30 408 - 33 936 PLN (B2B)
Wymagania
- GCP
- Communication skills
- Terraform
- Kubernetes
- Docker
- Go
- AI
- Monitoring
- Google Cloud Platform
- SRE
- VPC (nice to have)
- Storage (nice to have)
- Security (nice to have)
Opis stanowiska
O projekcie:
We are seeking a Senior Cloud SRE (Site Reliability Engineer) that will be primarily responsible for the performance, uptime, and growth of various OpenX systems and services on GCP. Much of your software development focuses on optimizing cloud-native systems, orchestrating cloud infrastructure and eliminating manual work through automation.
Excellent communication skills are crucial in this position so you could successfully interact with globally distributed OpenX teams operating in a 24x7 manner.
We are operating at a scale:
100% Cloud-based (GCP) platform
Over 380 billion Ad requests every day
Over 120 000 CPU daily
Over 140 TB RAM daily
Over 50 PB of data per week
Over 1200 production deployments a month
Wymagania:
What you need to have to be successful: - At least 5 years of GCP experience - Bachelor’s degree in Computer Science, related technical field involving systems engineering, or equivalent practical experience - Experience with Kubernetes/Docker/Containers - Experience with CI/CD - Experience in one of the following: Java, Python, Go, or other - Experience with current AI tools, especially Claude Code and Codex - Good English skills Desirable Qualifications: - Expertise in designing, analyzing and troubleshooting large-scale distributed systems - Good understanding of public cloud services and tasks, such as: VPC; load balancing; relational and non-relational datastores (e.g., Google Cloud SQL, Memorystore, AWS RDS); storage (e.g., GCS, AWS S3); monitoring (e.g., GCP Stackdriver, AWS CloudWatch, Prometheus); serverless computing (e.g., GCF, AWS Lambda); and auto-scaling - Ability to debug and optimize code and automate routine tasks - Experience with security
Codzienne zadania:
- Design, write and deliver software to implement and support large web-scale, highly-performant, highly-available infrastructure on GCP (e.g. Terraform)
- Monitor infrastructure, respond to incidents, correct and improve systems to prevent incidents, and plan capacity
- Support system deployments and product releases
- Tune large-scale clusters for optimal performance and efficiency
- Work closely with engineering, project management, and operational peers to develop innovative technical tools and solutions
- Participation in on-call rotation
We are seeking a Senior Cloud SRE (Site Reliability Engineer) that will be primarily responsible for the performance, uptime, and growth of various OpenX systems and services on GCP. Much of your software development focuses on optimizing cloud-native systems, orchestrating cloud infrastructure and eliminating manual work through automation.
Excellent communication skills are crucial in this position so you could successfully interact with globally distributed OpenX teams operating in a 24x7 manner.
We are operating at a scale:
100% Cloud-based (GCP) platform
Over 380 billion Ad requests every day
Over 120 000 CPU daily
Over 140 TB RAM daily
Over 50 PB of data per week
Over 1200 production deployments a month
Wymagania:
What you need to have to be successful: - At least 5 years of GCP experience - Bachelor’s degree in Computer Science, related technical field involving systems engineering, or equivalent practical experience - Experience with Kubernetes/Docker/Containers - Experience with CI/CD - Experience in one of the following: Java, Python, Go, or other - Experience with current AI tools, especially Claude Code and Codex - Good English skills Desirable Qualifications: - Expertise in designing, analyzing and troubleshooting large-scale distributed systems - Good understanding of public cloud services and tasks, such as: VPC; load balancing; relational and non-relational datastores (e.g., Google Cloud SQL, Memorystore, AWS RDS); storage (e.g., GCS, AWS S3); monitoring (e.g., GCP Stackdriver, AWS CloudWatch, Prometheus); serverless computing (e.g., GCF, AWS Lambda); and auto-scaling - Ability to debug and optimize code and automate routine tasks - Experience with security
Codzienne zadania:
- Design, write and deliver software to implement and support large web-scale, highly-performant, highly-available infrastructure on GCP (e.g. Terraform)
- Monitor infrastructure, respond to incidents, correct and improve systems to prevent incidents, and plan capacity
- Support system deployments and product releases
- Tune large-scale clusters for optimal performance and efficiency
- Work closely with engineering, project management, and operational peers to develop innovative technical tools and solutions
- Participation in on-call rotation
🔍 Dekoder Ogłoszenia
🔴
Much of your software development focuses on optimizing cloud-native systems, orchestrating cloud infrastructure and eliminating manual work through automation.
Oczekuje się, że będziesz pisać dużo kodu do automatyzacji i zarządzania infrastrukturą, a nie tylko monitorować i reagować.
🔴
Excellent communication skills are crucial in this position so you could successfully interact with globally distributed OpenX teams operating in a 24x7 manner.
Będziesz musiał komunikować się z zespołami na całym świecie, co może oznaczać pracę w niestandardowych godzinach lub ciągłą dostępność.
🟡
Over 380 billion Ad requests every day
Praca z ogromną skalą danych i ruchu, co wymaga bardzo wydajnych i stabilnych rozwiązań.
🟡
Experience with current AI tools, especially Claude Code and Codex
Oczekuje się aktywnego wykorzystania narzędzi AI do wspomagania pracy, co może być nowością dla niektórych kandydatów.
🟢
Bachelor’s degree in Computer Science, related technical field involving systems engineering, or equivalent practical experience
Formalne wykształcenie jest preferowane, ale doświadczenie może je zastąpić, co daje pewną elastyczność.