B2B telecom operator. Seven-person SRE team owning the full technical foundation — private OpenStack cloud across 3 datacenters, servers, storage arrays, 13 Kubernetes clusters and the delivery platform. Availability target 99%.
- Operate the group's entire technical foundation, hardware to applications: private OpenStack cloud across 3 datacenters (Kolla, bare metal via Bifrost/Ironic), Pure Storage arrays and 13 Kubernetes clusters totalling 250 nodes and 200 applications.
- Designed and maintain the platform Helm chart that deploys every application in the group, plus the data layer: 72 RabbitMQ clusters, 127 PostgreSQL databases under CNPG, 60 Redis/Valkey instances, a 3 TB / 18-node Cassandra cluster and MySQL.
- Migrated 110 applications and servers from internal IT to the private OpenStack/Kubernetes cloud with no interruption on business-critical scope.
- Initiated and shipped the on-premise LLM agent platform (vLLM, llm-d, GAIE, KServe, HolmesGPT) with RBAC and a Milvus/RagFlow knowledge base: automated alert triage, incident qualification time down ~80%.
- Migrated all 13 clusters from Kubespray to Talos — build time 1 hour → 5 minutes — and wrote and rehearsed the disaster-recovery plan: etcd restore, full cluster rebuild, backups, RTO and RPO per scope.
- Standardised GitOps (Argo CD, Helm, Kustomize), policy-as-code (Kyverno) and elasticity (KEDA on RabbitMQ queues); automated access (Boundary), identity (Keycloak/OIDC via Terraform) and upgrades (Renovate).