Site Reliability · AI Platform · Paris, remote

Fewer incidents. A smaller cloud bill.

Ten years in infrastructure, eight of them running Kubernetes in production. 50 clusters, 1,000+ nodes, from bare metal to multi-cloud. What I leave behind: a platform your team runs without me.

Selected work

Platforms I have owned, and what changed

Inherent · Unyc 2025 — present · B2B telecom

The whole technical foundation, hardware to developer self-service: a private OpenStack cloud, 13 Kubernetes clusters and 250 nodes, 72 RabbitMQ clusters and 127 PostgreSQL databases. 110 applications moved onto it with no interruption on business-critical scope.

Cluster build 1 hour → 5 minutes. Incident qualification time down 80% with an on-premise LLM agent.

Kering 2024 — 2025 · Luxury group

Multi-cloud AWS and Alibaba platform: a majority of workloads on ECS/Fargate plus 6 Kubernetes clusters, about 200 nodes and 500+ microservices under continuous deployment for the group's houses, sized to absorb private sales and Fashion Week. Co-owned the worldwide customer and loyalty database — Couchbase, 2 TB, XDCR across three regions.

Capacity matched to real load with Karpenter and KEDA, not provisioned for peaks.

Ministère de la Transition écologique 2021 — 2024 · Central government

The shared internal GitLab platform deployed on Kubernetes for 2,500 developers and 350 projects, with its runner fleet industrialised on Kubernetes and reusable Terraform modules adopted across the ministry.

Job queue time down 70%. Environment provisioning 1 day → 2 hours.

Independent clients 2020 — present · ~10 engagements

Short public-cloud engagements alongside the long missions: audits of architecture, reliability, security posture and cost, delivered as a prioritised remediation plan — and the builds that follow. One GCP estate rebuilt from scratch as code after a security breach.

Public cloud

AWS and GCP, in production

Most of the public-cloud work happened on these two. The rest of the stack is below.

AWS Since 2018 · Solutions Architect – Associate

EKS and ECS under continuous deployment for 500+ microservices, five Terraform / Terragrunt / Terramate foundation stacks across accounts, Karpenter and KEDA for elasticity, IAM and IRSA for access, with RDS, Aurora, ElastiCache, SNS, SQS and CloudWatch underneath. Crossplane self-service on top, so a product team ships application and infrastructure in one merge request.

40 microservices off Elastic Beanstalk in a year. Delivery pipeline 20 minutes → 3.

GCP Since 2019 · GKE, Pub/Sub

Managed Kubernetes on GKE and event pipelines on Pub/Sub for independent clients, run alongside the AWS work — including one client's entire estate rebuilt from nothing as infrastructure as code after a security breach: assessment, reconstruction, hardened access and network exposure.

A compromised estate back in production, entirely as code.

Stack

What I work with

Tools are evidence, not the argument. This is what is in production today.

Containers & orchestration

Kubernetes — EKS, ACK, GKE, RKE2, Talos, Kubespray, OpenShift, Magnum · Docker · Helm · Kustomize · Karpenter · KEDA · Kyverno · Crossplane

Cloud & infrastructure

AWS · OpenStack · GCP · Alibaba Cloud · OVHcloud · Scaleway · Cloudflare · Pure Storage · Linux (Debian, Ubuntu, RHEL)

Infrastructure as code

Terraform · OpenTofu · Terragrunt · Terramate · Spacelift · Ansible · Helm · OpenStack HEAT

CI/CD & GitOps

GitLab CI · GitHub Actions · Azure DevOps · Argo CD · Flux · Renovate · Jenkins · blue/green deployment

Observability

Prometheus · Grafana · Loki · Mimir · Tempo · Alloy · Faro · OpenTelemetry · Alertmanager · Datadog · Sentry · Fluent Bit

AI / LLMOps

vLLM · KServe · HolmesGPT · Kagent · LiteLLM · Milvus · RagFlow · MCP · AWS Bedrock · Hugging Face

Security & secrets

HashiCorp Vault · Boundary · Consul · Keycloak (OIDC) · SOPS · Trivy · Falco · IAM/IRSA · policy-as-code

Data & storage

Couchbase · PostgreSQL (CNPG) · Cassandra · MySQL · MongoDB Atlas · Redis / Valkey · RabbitMQ · ClickHouse · Ceph · MinIO

Practices

SLI / SLO / error budgets · DORA metrics · blameless post-mortems · on-call · toil reduction · developer self-service · FinOps · DevSecOps

Blog

Latest articles

All posts
Open to work — Q4 2026

Long-term platform work, or a two-week audit that tells you the truth.

Paris, fully remote. Seed-stage through regulated enterprise. Contract or permanent.

Start a conversation