Founder · Platform & Observability Engineer

Mehrshad Lotfi.

I help teams run Kubernetes platforms that are observable, reliable and cost-efficient.

Grafana Champion and Kubestronaut with a research background in cloud networking. I design and operate the Grafana LGTM stack, Prometheus and OpenTelemetry on Kubernetes, and I turn monitoring into code your team can own.

Frankfurt am Main · remote or on-site within 250 kmAvailable for new projects
Grafana Champion 2024KubestronautUpwork Top Rated Plus
Book a call
Mehrshad Lotfi
KubestronautGrafana ChampionCertified Cloud Native Platform Engineering Associate
Founder & Platform EngineerOptiop GmbH

$150k+

yearly cloud savings delivered for one client

~20%

lower cloud spend after moving off CloudWatch

15+

cloud-native certifications

100%

Job Success Score on Upwork

Companies I've worked with.

From exchanges and banks to research and fast-growing SaaS teams


Deutsche Börse
Software / DevOps engineer
dwpbank
Monitoring & alerting
MPI-SWS
Max Planck Institute for Software Systems
REDCap Cloud logo
Healthcare
Cequens logo
Cloud communications
Patstown logo
Technology partner
Scope ICT logo
ICT solutions

Selected projects.

What I built, and what it changed for the teams I worked with


Healthcare SaaS · via Optiop2025

$150k+

saved per year

From CloudWatch to the Grafana LGTM stack

Replaced CloudWatch-based logging, metrics and tracing with a self-hosted stack on Kubernetes: Loki for logs, Tempo for traces and Cortex for long-term metrics, all visualised in Grafana and instrumented with OpenTelemetry.

  • Cut the client's cloud spend by around 20%, more than $150k per year
  • One place to correlate logs, metrics and traces during incidents
  • Rolled out with Helm and ArgoCD so every change is reviewed in Git
GrafanaLokiTempoCortexOpenTelemetryAWSArgoCD
E-commerce & healthcare · via Optiop2025 – today

Dedicated monitoring platform with SSO

Built dedicated monitoring clusters that run independently of the workloads they observe, deployed with Helm and ArgoCD. Dashboards and alert rules live in Git as code, and access is handled through single sign-on.

  • Alerts and dashboards as code, versioned and reviewed like application code
  • SSO with Azure Entra ID and Keycloak, mapped to team roles
  • Monitoring stays available even when production clusters are degraded
KubernetesHelmArgoCDPrometheusGrafanaKeycloakEntra ID
Platform teams · via Optiop2025

Cost and delivery visibility with OpenCost & DORA

Introduced OpenCost to break Kubernetes spend down by team and workload, and DORA metrics to measure how fast and how safely teams ship.

  • Per-namespace and per-team cost dashboards in Grafana
  • Deployment frequency, lead time, change failure rate and MTTR tracked automatically
  • Shared numbers for engineering and management discussions
OpenCostPrometheusGrafanaKubernetesDORA
Deutsche WertpapierService Bank (dwpbank)

Monitoring and alerting for banking microservices

Worked as a DevOps engineer improving monitoring and alerting across microservices, databases and Kafka, with a focus on actionable alerts and faster root-cause analysis.

  • Consistent dashboards for services, databases and Kafka
  • Alert rules tuned to reduce noise and catch real incidents earlier
PrometheusGrafanaKafkaMicroservicesDatabases
Deutsche Börse02/2023 – 07/2023

High-performance client for Cloud Stream market data

Built a high-performance client for Deutsche Börse's Cloud Stream market data service and supported the migration of workloads from AWS to Google Cloud.

  • Low-latency consumption of real-time market data
  • Hands-on support for the AWS to Google Cloud migration
Google CloudAWSStreamingPerformance engineering
Max Planck Institute for Software Systems10/2019 – 01/2023

CPU-efficient network stacks for cloud virtualization

Research on making cloud network virtualization cheaper in CPU cycles, working close to the hardware with virtual switches, NICs and offloading.

  • Research on CPU-efficient networking for multi-tenant clouds
  • Deep understanding of how networking behaves under load, which now shapes how I monitor it
Open vSwitchNICsPCIeFPGAC/C++

Freelance highlights

A selection of completed projects on Upwork

Real-time alerts for a crisis management application

Real-time alerting and notifications built on Grafana dashboards and Loki log queries.

GrafanaLokiAlerting

Grafana Cloud monitoring for network appliances

Monitoring of network appliances with Grafana Cloud, from data collection to dashboards.

Grafana CloudNetworking

Velocity dashboard

Dashboard build project in Grafana, delivered end to end and rated five stars by the client.

GrafanaDashboards

AWS infrastructure & application monitoring

End-to-end monitoring for AWS infrastructure and the applications running on it.

AWSGrafanaPrometheus

Experience.


  1. 2025 – today

    Founder & Platform Engineer

    Optiop GmbH

    Observability, platform engineering and Kubernetes consulting for teams in healthcare and e-commerce.

  2. 08/2023 – today

    Freelance DevOps & Site Reliability Engineer

    Self-employed

    Observability and Kubernetes projects for clients in Germany and worldwide, including via Upwork.

  3. 02/2023 – 07/2023

    DevOps Engineer

    Deutsche Börse

    Cloud Stream market data client and AWS to Google Cloud migration.

  4. 10/2019 – 01/2023

    Research Associate

    Max Planck Institute for Software Systems

    CPU-efficient network stacks for cloud network virtualization.

Education

  • M.Sc. Computer Science

    Saarland University & Max Planck Institute for Software Systems

  • B.Sc.

    Sharif University of Technology

Languages

English (business fluent)German (business fluent)Persian (native)

Certifications.

Verified credentials across Kubernetes, observability and platform engineering


Let's make your platform observable.

Tell me about your stack and your goals. In a free first call we look at where monitoring, reliability or cloud costs can improve.

mehrshad.lotfi@optiop.org