Talent.com
Nebul
Site Reliability Engineer – AI Cloud PlatformNebul • Leiden, South Holland, Netherlands
Site Reliability Engineer – AI Cloud Platform

Site Reliability Engineer – AI Cloud Platform

Nebul • Leiden, South Holland, Netherlands
5 dagen geleden
Functieomschrijving

About Nebul

At Nebul, we’re building Europe’s sovereign AI cloud — trusted, secure, and purpose-built for the next generation of intelligent infrastructure.

Our platform combines Kubernetes, NVIDIA GPU infrastructure, cloud-native services and software written in Go and Python. As the platform grows, we need to increase our internal reliability capability and reduce the number of operational issues that depend on a small group of senior engineers.


What You’ll Be Doing

As a Site Reliability Engineer, you’ll work approximately 70% on site reliability and operational troubleshooting and 30% on broader DevOps and platform engineering.

Your primary responsibility will be to keep Nebul’s AI cloud platform stable, observable and operationally scalable. You’ll take ownership of incidents, investigate complex issues and ensure that senior engineers are not pulled into every troubleshooting session.

You’ll work across Kubernetes, NVIDIA GPU infrastructure, services written in Go and Python, networking and cloud-native platform components.

The role is not only about responding to incidents. You’ll identify recurring problems, automate operational work and improve the platform so that the same issues do not continue to return.


Key Responsibilities

  • Monitor and improve the reliability, availability and performance of Nebul’s AI cloud platform.
  • Troubleshoot incidents across Kubernetes, NVIDIA GPU infrastructure, Linux, networking and platform services.
  • Take ownership of technical troubleshooting sessions and coordinate issues through to resolution.
  • Investigate issues affecting services written in Go and Python.
  • Perform root-cause analysis and translate findings into structural platform improvements.
  • Build and improve monitoring, metrics, logging, tracing and alerting.
  • Ensure alerts are relevant, actionable and connected to clear operational procedures.
  • Create runbooks, escalation procedures and troubleshooting documentation.
  • Automate repetitive operational tasks using Go, Python or scripting.
  • Support Kubernetes cluster operations, deployments, upgrades and platform changes.
  • Improve the production readiness of new services and infrastructure components.
  • Work with engineering teams to improve resilience, observability and failure handling.
  • Identify reliability risks before they result in platform or customer impact.
  • Support operational improvements across GPU workloads and NVIDIA-based infrastructure.
  • Reduce the operational dependency on senior platform engineers.
  • Contribute to broader DevOps work when additional capacity is needed within the team.


What Your Day Will Not Look Like

  • Acting as a first-line support engineer who only closes tickets.
  • Escalating every complex issue another senior engineer.
  • Spending all your time manually operating Kubernetes.
  • Resolving incidents without addressing their underlying causes.
  • Building isolated automation that is not integrated into the platform.


What You Bring

  • Strong experience as a Site Reliability Engineer, DevOps Engineer or Platform Engineer.
  • Hands-on production experience with Kubernetes.
  • Strong Linux and infrastructure troubleshooting skills.
  • Experience investigating issues across applications, infrastructure, networking and cloud platforms.
  • Experience with services written in Go or Python.
  • Practical experience with monitoring, logging, metrics and alerting.
  • Experience responding to production incidents and performing root-cause analysis.
  • The ability to independently lead complex troubleshooting sessions.
  • Experience automating operational work using Python, Go or scripting.
  • A solid understanding of cloud-native and distributed systems.
  • A calm, analytical and structured approach to incidents.
  • An ownership mindset and the ability to move from reactive troubleshooting to lasting improvements.


Bonus Points If You Have

  • Experience with NVIDIA GPU infrastructure.
  • Experience supporting AI, machine-learning or high-performance computing workloads.
  • Knowledge of Kubernetes GPU scheduling and resource management.
  • Experience with multi-tenant cloud environments.
  • Familiarity with Go-based cloud or platform services.
  • Experience with Infrastructure as Code and automated platform deployment.
  • Knowledge of distributed storage, networking or database troubleshooting.
  • Experience defining service-level indicators, objectives and operational reliability targets.
  • Experience working in sovereign, regulated or security-sensitive cloud environments.


Eligibility & Application Information

We welcome non-native Dutch speakers to apply. However, to be eligible, you must:

  • Have a valid work permit in the Netherlands. ( wo do offer sponsorship if needeed)
  • Reside in the Netherlands and be able to travel to the office in Leiden (near The Hague).
  • Be fluent in English. Dutch is not required.


Ready to make Europe’s sovereign AI cloud more reliable and operationally scalable?

Apply now through Frank Poll and help Nebul build a cloud platform that engineering teams and customers can depend on.


Maak een vacature-alert aan voor deze zoekopdracht

Site Reliability Engineer – AI Cloud Platform • Leiden, South Holland, Netherlands

Vergelijkbare banen

Site Reliability Engineer

eTeamden haag, zuid holland, Netherlands

Team The Hague, South Holland, Netherlands /ph3Site Reliability Engineer /h3peTeam The Hague, South Holland, Netherlands /ppGet AI-powered advice on this job and more exclusive features.Direct mess... Laat meer zien

 • Gesponsord

Reliability Engineer

Maintenance BanenLeiden, The Netherlands

Vergeet niet uw cv te controleren voordat u solliciteert.Zorg er ook voor dat u alle vereisten met betrekking tot deze functie doorleest.Dit 24/7 draaiende productiebedrijf is wereldmarktleider in ... Laat meer zien

 • Gesponsord

Senior Full-Stack AI Engineer (LLMs & Agentic Systems)

Nebulleiden, zuid holland, Netherlands

At Nebul, we’re building Europe’s sovereign AI cloud — trusted, secure, and purpose-built for the next generation of intelligent infrastructure.On top of this foundation, we’re building AI-native a... Laat meer zien

 • Gesponsord

Forward Deployed Engineer

Freedayrotterdam, zuid holland, Netherlands

At Freeday, we build digital employees that automate repetitive work so people can focus on what really matters.Our AI platform helps companies streamline processes, connect systems and bring intel... Laat meer zien

 • Gesponsord

Site Reliability Engineer

ETeamDen Haag, The Netherlands

Team The Hague, South Holland, Netherlands.Hieronder staat alles wat u moet weten over wat deze vacature inhoudt, en wat er van sollicitanten wordt verwacht.Team The Hague, South Holland, Netherlan... Laat meer zien

 • Gesponsord

Platform Leader: Reliability & AI-Ready SaaS Scale

Last Mile Solutionsrotterdam, zuid holland, Netherlands

Last Mile Solutions is seeking a Head of Platforms in Rotterdam to own and evolve the core platform, infrastructure and reliability.You will lead Platform, Database and DevSecOps teams to deliver a... Laat meer zien

 • Gesponsord

SRE (Site Reliability Engineering)

Huawei EuropeRijswijk, Zuid-Holland, Netherlands

Huawei is a leading global provider of information and communications technology (ICT) infrastructure and smart devices.With integrated solutions across four key domains – telecom networks, IT, sma... Laat meer zien

Senior Backend & Platform Engineer, AI Cloud & Edge

Clockworksrotterdam, zuid holland, Netherlands

A leading AI specialist in Rotterdam is seeking an experienced Back-end Platform Engineer.You will design and build robust back-end systems for our proprietary AI platform, focusing on cloud archi... Laat meer zien

 • Gesponsord

Senior Cloud Architect: Azure, AWS & AI-Ready Cloud

CIMSOLUTIONSRemote, zuid holland, Netherlands
Op afstand

CIMSOLUTIONS zoekt een Cloud Architect om te helpen bij de ontwikkeling van innovatieve cloudproducten.In deze rol ben je verantwoordelijk voor het ontwerpen van de architectuur, het vertalen van r... Laat meer zien

 • Gesponsord

Cloud Platform Engineer

Human Source Grouprotterdam, zuid holland, Netherlands

Ben jij een cloud platform engineer die het liefst werkt aan design en beheer van een volwassen infrastructuur? Als cloud platform engineer word je hier onderdeel van een team van 14 ervaren platfo... Laat meer zien

 • Gesponsord

Cloud Engineer: AWS, Kubernetes & IaC Automation

The Preferred Supplierden haag, zuid holland, Netherlands

Een technologiepartner voor juridische en fiscale professionals zoekt een Cloud Engineer (medior/senior) voor het optimaliseren van hun AWS-omgeving.Je ontwerpt en beheert de cloud infrastructuur e... Laat meer zien

 • Gesponsord

Cloud Engineer

PostNLden haag, zuid holland, Netherlands

PostNL The Hague, South Holland, Netherlands /ph3strongJoin or sign in to find your next job /strong /h3pJoin to apply for the strongCloud Engineer /strong role at strongPostNL /strong /ppPostNL Th... Laat meer zien

 • Gesponsord

Site Reliability Engineer

Online Payment Platformdelft, zuid holland, Netherlands

Site Reliability Engineer (SRE) /h3 pbWaarom jij in ons technology team past: /b /p ol liJe kent dat gevoel - die ‘yeah!’ als jouw code live gaat en direct op miljoenen gebruikers impact heeft.Je z... Laat meer zien

 • Gesponsord

Senior AI Platform & Agents Engineer (Remote)

Wolters Kluwerworkfromhome, zuid holland, Netherlands
Op afstand

A global information services company in the Netherlands is seeking a Full Stack Engineer to build the GenAI platform for critical decision-making in healthcare and compliance sectors.The ideal can... Laat meer zien

 • Gesponsord

Support & Site Reliability Engineer

OntzorgdRotterdam, The Netherlands

Bij Ontzorgd ontwikkelen we slimme software die GGZ-professionals helpt om betere zorg te leveren met meer aandacht.We bouwen aan een AI-platform waarmee psychologen en psychiaters sneller, efficië... Laat meer zien

 • Gesponsord

Fintech Site Reliability Engineer for Scale & Uptime

Rotterdam Innovation Citydelft, zuid holland, Netherlands

An innovative fintech company in Delft is seeking a Site Reliability Engineer to enhance platform reliability and performance.You will build monitoring systems, automate incident responses, and ens... Laat meer zien

 • Gesponsord

Senior AI Platform Engineer – GenAI & Agents (Remote)

Wolters Kluwerworkfromhome, zuid holland, Netherlands
Op afstand

A leading global information services company is seeking a Full Stack Engineer for its AI Platform Agents team to build the GenAI platform that powers critical decisions in various industries.The ... Laat meer zien

 • Gesponsord

SRE (Site Reliability Engineering)

HuaweiRijswijk, The Netherlands

Huawei is a leading global provider of information and communications technology (ICT) infrastructure and smart devices.With integrated solutions across four key domains – telecom networks, IT, sma... Laat meer zien

 • Gesponsord

Senior Cloud Architect: AWS, Kubernetes & CI/CD Leader

The Preferred Supplierden haag, zuid holland, Netherlands

Een softwareleverancier voor de juridische sector in Den Haag is op zoek naar een ervaren Cloud Engineer of Architect.In deze rol ben je verantwoordelijk voor het opstellen en implementeren van de ... Laat meer zien

 • Gesponsord

Cloud DevOps Engineer — AWS, CI/CD & Automation

CoinMarketCapzoetermeer, zuid holland, Netherlands

A leading cryptocurrency analytics platform based in the Netherlands is seeking a skilled Cloud Operations Engineer to optimize cloud architectures and automate operations.The ideal candidate will ... Laat meer zien