Metrics & dashboards
Prometheus and Grafana, with dashboards built around services and user journeys — not just CPU graphs.
Too many teams learn about outages from their customers. I set up metrics, logs and alerting with Prometheus and Grafana that show what users actually experience — and prepare the incident process for the day something breaks anyway.
When to call me
These are the situations where teams usually bring me in.
What I deliver
Prometheus and Grafana, with dashboards built around services and user journeys — not just CPU graphs.
Logs from every service in one place, searchable when it matters.
Service-level objectives and alerts on the symptoms users feel: fewer alerts, each one meaningful.
On-call rules, runbooks and blameless post-mortems that turn incidents into improvements.
Rate limits, caching and capacity planning, so a traffic spike doesn’t become an outage.
How it works
Review the current monitoring, alerting and recent incidents.
Add the metrics and logs that answer real questions during an incident.
Define SLOs, cut alert noise and write runbooks.
Walk through failure scenarios, so the first real incident isn’t the first rehearsal.
FAQ
Clusters on Google Cloud (GKE) or bare metal — designed, migrated, upgraded and hardened. Secure by default, boring to operate.
Fast, reliable pipelines in GitLab CI or GitHub Actions — from commit to production with automated checks and safe rollbacks.
An independent review of your infrastructure: what to fix first, what to automate and where you are overpaying — in the cloud or on hardware.
Contact
Tell me what you’re building and where it hurts — I’ll get back to you with next steps.