Site Reliability Engineer (Kubernetes Platform)
Lees verder om te ontdekken wat u nodig heeft om te slagen in deze functie, inclusief vaardigheden, kwalificaties en ervaring.
About the Team
The Kubernetes Platform team builds and operates an internal Kubernetes platform underpinning hundreds of engineering teams. The team manages a large fleet of EKS and on-premises clusters across multiple regions and focuses on treating operations as a software problem. When something is repetitive, it is automated. When something breaks, lessons are learned and improvements are made.
The Role
This is a hands-on operational Site Reliability Engineering role focused on keeping the Kubernetes platform healthy, up-to-date, and well-supported.
You will spend most of your time on cluster maintenance, component upgrades, and helping engineering teams successfully run their workloads on the platform. In addition, you may act as a consultant to product teams on Kubernetes best practices and reliability topics.
Day-to-Day Responsibilities
Cluster Operations & Maintenance
- Plan and execute Kubernetes version upgrades across EKS and on-premises clusters, coordinating with internal teams to minimise disruption.
- Perform routine maintenance, including add-on upgrades, storage and networking configuration, and upgrades for monitoring, security, and other platform tooling.
- Monitor cluster health across the fleet and proactively address degradation signals before they become incidents.
Internal Customer Support
- Act as the first point of contact for engineering teams running workloads on the platform.
- Triage issues, diagnose failures, and guide teams towards resolution.
- Help teams understand platform capabilities, quota management, and best practices for running reliable workloads.
- Evaluate quota requests and usage requirements against platform capacity.
- Contribute to runbooks and FAQs to reduce recurring support requests.
Toil Reduction & Automation
- Identify repetitive manual tasks and reduce them through scripting and automation.
- Flag and address technical debt that increases operational risk or slows down delivery.
- Partner with the wider platform team on tooling improvements that reduce operational burden across the fleet.
Required Skills
- StrongKubernetes operational experience, including:
- Node management
- Kubernetes upgrades
- Workload debugging
- Cluster health management
- Experience with managed Kubernetes platforms (EKS or equivalent) and/or on-premises Kubernetes environments.
- Experience withobservability tooling such as:
- Prometheus
- VictoriaMetrics
- Grafana
- Alerting pipelines
- Strong written and verbal communication skills.
- Ability to work independently, prioritise effectively, and meet commitments.
Nice to Have
- Terraform for infrastructure provisioning.
- Puppet or similar configuration management tools.
- AWS experience.
- Experience supporting internal developer platforms or infrastructure teams.
What We're Looking For
- Think Customer First
- Own It
- Succeed Together
- Learn Forever
- Do The Right Thing
What This Role Is Not
- A pure software development role. The focus is operational excellence, platform reliability, and customer support rather than feature development.
- A solo contributor role. Collaboration, escalation, pairing, and knowledge sharing are essential.
- A reactive-only role. xnkfpon Proactive maintenance, automation, and continuous improvement are equally important as incident response.
Contract Details
- Location: Amsterdam
- Start Date: 01 August 2026
- End Date: 31 January 2027
- Duration: 6 months
- Hours: 40 per week
- Working Model: Hybrid (Amsterdam-based)