Sample portfolio · fictional candidate · built by Fieldcraft ← All portfoliosView the CVGet yours
~/lena-kowalski

Berlin, Germany · open to remote

Lena Kowalski

// Senior Site Reliability Engineer

Six years building and operating distributed systems at scale. At Zalando I own the reliability posture for 40 microservices processing 2.8M orders a day — and keep them faster, cheaper to run, and up 99.97% of the time.

  • 6 yrsdistributed systems at scale
  • 420★k8s-cost-exporter
  • 3 certsGCP · CKA · AWS

Terminal introduction: whoami — Lena Kowalski, Senior Site Reliability Engineer. Now: reliability for 40 microservices at Zalando SE, 2.8M orders a day, 99.97% uptime across 24 months. Career: Zalando SE (2022, Senior SRE), N26 GmbH (2020, Software Engineer, Platform), Delivery Hero (2018, Software Engineer). Status: open to remote, based in Berlin.

01// reliability dashboard

Reliability, by the numbers

Headline results from my current role as Senior SRE at Zalando SE — the service estate I own, and what changed while I owned it.

sre-overview scope: zalando-se / 40 services Feb 2022 → nowsnapshot

Availability

24 months

Maintained across 24 months for 40 microservices.

p99 API latency

ms

880210ms

Service-mesh routing redesign + circuit-breaking policies.

MTTR

−62%

4216min

Over 12 months, after new on-call runbooks and incident response.

Infrastructure cost

saved

€340K/ year

Cut by leading the migration of 18 services from VM-based deployments to GKE.

Throughput

/ day

2.8Morders / day

02// case studies

The systems work behind the numbers

Problem, approach, outcome — how the headline figures were actually earned.

case_01 · Zalando SE · latency

Cutting p99 latency across a 40-service mesh

Problem
p99 API latency across a mesh of 40 microservices — handling 2.8M orders a day — stood at 880ms.
Approach
Redesigned the service mesh routing and introduced circuit-breaking policies between services.
Outcome
p99 latency down to 210ms.
  • service mesh
  • circuit breaking
  • p99 latency
incoming traffic 2.8M orders/day SERVICE MESH routing redesigned routing layer closed closed open service service failing ×40 microservices in the mesh p99 880ms → 210ms
closed breaker · traffic flowsopen breaker · failing dependency cut off Simplified schematic of the approach, not a production topology.

case_02 · Zalando SE · platform

Moving 18 services from VMs to GKE

Problem
18 services were still running on VM-based deployments.
Approach
Led the migration of all 18 services onto Google Kubernetes Engine (GKE).
Outcome
Infrastructure cost cut by €340K a year.
  • GKE
  • Kubernetes
  • cost
BEFORE VM-based deployments VM VM 18 services migration I led AFTER Google Kubernetes Engine GKE −€340K / year infra cost
1 hexagon = 1 migrated service Simplified schematic; VM count is illustrative.

case_03 · Zalando SE · incident response

Cutting MTTR from 42 to 16 minutes

Problem
Mean time to recovery stood at 42 minutes.
Approach
Designed on-call runbooks and an incident response framework.
Outcome
MTTR down to 16 minutes over 12 months — a 62% cut.

case_04 · N26 GmbH · delivery

Releases in hours, not days, for 12 teams

Problem
A release cycle of 3 days for 12 backend teams.
Approach
Built an internal deployment pipeline in Go.
Outcome
Release cycle down to 4 hours.

More from the log

200+services

Automated TLS certificate rotation, eliminating 4 outages a year caused by expired certificates.

N26 GmbH · 2020–22

8critical services

Instrumented with OpenTelemetry — surfaced 3 latency regressions before they reached production.

N26 GmbH · 2020–22

120Kevents / sec

Order-tracking service I developed in Go.

Delivery Hero · 2018–20

0downtime

Zero downtime on that order-tracking service since its launch in Nov 2019.

Delivery Hero · 2018–20

−67%query time

PostgreSQL query time cut through index optimisation and query-plan analysis on a 2.4TB dataset.

Delivery Hero · 2018–20

3/3mentees certified

Mentored 3 junior SREs; all 3 passed Google Professional Cloud Architect within 6 months of joining the team.

Zalando SE · 2022–now

03// open source & writing

In the open

Tools I've built and contributed to, and what I've written up for other engineers.

k8s-cost-exporter

Prometheus exporter for Kubernetes cost attribution.

prometheuskubernetescost 420 GitHub stars

go-slo-toolkit

contributor

SLO tracking library for Go services.

goslo 2 merged PRs

Zero-downtime database migrations with Go and Postgres

  • Go
  • PostgreSQL
  • migrations

18Kreads

04// stack

The stack, layer by layer

From the code I write down to the clouds it runs on.

  1. languagesGoPython
  2. data & streamingKafkaPostgreSQLRedis
  3. observabilityPrometheusOpenTelemetry
  4. deliveryHelmArgoCDGitHub Actions
  5. orchestration & IaCKubernetesTerraform
  6. cloudGCPAWS

Certifications

  • 2023Google Professional Cloud Architect
  • 2022Certified Kubernetes Administrator (CKA)
  • 2021AWS Solutions Architect Associate

AI-assisted development

35%faster shipping

Using GitHub Copilot and Cursor since adopting them in 2024.

05// experience

Six years, three Berlin companies

From building services, to building the platform, to owning reliability.

  1. Feb 2022 – Present
    now

    Senior Site Reliability Engineer

    Zalando SE · Berlin

    • Owns reliability for 40 microservices processing 2.8M orders/day; 99.97% uptime across 24 months.
    • p99 API latency 880ms → 210ms by redesigning service mesh routing and introducing circuit breaking.
    • Led the migration of 18 services from VMs to GKE, cutting infrastructure cost by €340K/year.
    • Designed on-call runbooks and incident response; MTTR 42 → 16 min over 12 months.
    • Mentored 3 junior SREs — all passed Google Professional Cloud Architect within 6 months.
  2. Aug 2020 – Jan 2022

    Software Engineer — Platform

    N26 GmbH · Berlin

    • Built an internal deployment pipeline in Go: release cycle 3 days → 4 hours for 12 backend teams.
    • Instrumented 8 critical services with OpenTelemetry; caught 3 latency regressions before production.
    • Automated certificate rotation across 200+ services, eliminating 4 expired-TLS outages a year.
  3. Sep 2018 – Jul 2020

    Software Engineer

    Delivery Hero · Berlin

    • Developed an order-tracking service in Go handling 120K events/sec; zero downtime since launch in Nov 2019.
    • Cut PostgreSQL query time by 67% with index optimisation and query-plan analysis on a 2.4TB dataset.
  4. Oct 2014 – Sep 2018

    B.Sc. Computer Science

    Technische Universität Berlin

    Grade 1.4 (First Class equivalent).

06// contact

Let's talk reliability

Based in Berlin, Germany, and open to remote roles.

Lena is a fictional candidate — her contact details are shown as plain text for illustration only.

email
lena.kowalski@email.com
location
Berlin, Germany · open to remote
github
github.com/lenakowalski
linkedin
linkedin.com/in/lenakowalski