Jobiglo

No results.

Senior Site Reliability Engineer – MAAS & GPU Infrastructure

Pragmatike · Erevan

New Remote
Remote Senior 🇬🇧 English
Linux MAAS Kubernetes Ansible Bash Python Terraform OpenTofu Git Prometheus Grafana Alertmanager VictoriaMetrics VictoriaLogs IPMI Redfish BMCs RAID Proxmox KVM libvirt OpenStack VMware GPU infrastructure VLAN L2/L3 routing VPN firewalls DNS

Job description

About the role

Pragmatike is looking for a Senior Site Reliability Engineer to own and scale a distributed infrastructure platform that includes bare‑metal GPU nodes, Kubernetes, virtualization, networking and multi‑site environments. You will work hands‑on from low‑level hardware provisioning through high‑level automation and observability.

Key responsibilities

  • Operate and maintain large‑scale Debian/Ubuntu Linux infrastructure on both bare‑metal and virtualised environments.
  • Own MAAS‑based bare‑metal provisioning, including PXE, cloud‑init, node lifecycle and API/CLI automation.
  • Run production Kubernetes clusters – upgrades, node pools, networking, storage, security hardening and troubleshooting.
  • Design and manage multi‑site networking (VLANs, L2/L3 routing, VPNs, firewalls, DNS).
  • Automate provisioning and operations with Ansible, Bash/Python, OpenTofu/Terraform and Git‑based workflows.
  • Maintain observability platforms such as Prometheus, Grafana, Alertmanager and VictoriaMetrics/VictoriaLogs.
  • Lead incident response, on‑call coverage, post‑mortem analysis and reliability improvements.
  • Manage virtualization platforms (Proxmox, KVM/libvirt, OpenStack, VMware) and GPU passthrough where required.

Required profile

  • 5+ years of hands‑on SRE, infrastructure or platform engineering experience.
  • Deep expertise in Linux administration (Debian/Ubuntu) and MAAS bare‑metal provisioning.
  • Proven production experience with Kubernetes, including networking, storage and upgrades.
  • Strong network engineering skills across VLANs, L2/L3 routing, bonding, VPNs, firewalls and DNS.
  • Extensive automation experience with Ansible, Bash or Python and Terraform/OpenTofu.
  • Experience with observability tools (Prometheus, Grafana) and on‑call incident management.

Required skills

  • Linux (Debian/Ubuntu)
  • MAAS
  • Kubernetes
  • Ansible
  • Bash
  • Python
  • Terraform / OpenTofu
  • Git
  • Prometheus
  • Grafana
  • Alertmanager
  • VictoriaMetrics / VictoriaLogs
  • IPMI / Redfish
  • BMCs
  • RAID
  • Proxmox
  • KVM / libvirt
  • OpenStack
  • VMware
  • GPU infrastructure
  • VLAN, L2/L3 routing, VPN, firewalls, DNS

What we offer

  • 100% remote work with flexible hours (EU timezones)
  • High‑impact role with technical ownership and autonomy
  • Opportunity to shape architecture of a growing cloud platform
  • International, engineering‑driven team focused on automation and reliability

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Pragmatike.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches for Հայաստան.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 5 ժամ առաջ

Expires 1 ամիսից

3 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Pragmatike

Erevan