Skip to content

Migration playbook: Datadog → Prometheus and Grafana

Datadog to Prometheus and Grafana migration

Replace a growing Datadog bill with OpenTelemetry, Prometheus and Grafana, keeping the dashboards and alerts your on-call team relies on.

Migration drivers

Why teams make this move

Driver 01

Unpredictable billing

Invoices that jump with custom metric cardinality and log volume.

Driver 02

Proprietary agents

Instrumentation tied to vendor agents instead of vendor-neutral OpenTelemetry.

Driver 03

Retention cost

Paying a premium to keep traces and logs for more than a few weeks.

Execution sequence

How the migration runs

Each phase ends with a check you can verify — data parity, error rates, latency — and the rollback path is agreed before any traffic moves.

  1. 01Phase

    OpenTelemetry Collector

    Deploying the OpenTelemetry Collector to receive traces, metrics and logs in OTLP.

  2. 02Phase

    Metrics backend

    Setting up Prometheus or VictoriaMetrics with long-term storage on object storage.

  3. 03Phase

    Dashboards and alert rules

    Translating Datadog dashboards and monitors into Grafana dashboards and Alertmanager rules.

  4. 04Phase

    Datadog agent removal

    Removing the Datadog agents once alert coverage has been checked.

Risk prevention

Pitfalls that derail this migration

Risk 01

High-cardinality labels

Exporting user IDs or transaction hashes as Prometheus labels, which exhausts memory.

Risk 02

Copied alert thresholds

Porting Datadog anomaly monitors to PromQL without retuning thresholds and evaluation intervals.

Risk 03

Log shipping bandwidth

Sending uncompressed logs over public networks instead of aggregating them locally first.

Before and after

What we measure

We take a baseline before any change and report the same numbers after cutover, from your own tools. They are the evidence of whether the migration worked — not figures promised in advance.

Monthly cost
Datadog invoice vs the new stack's infrastructure and licenses
Alert coverage
Datadog monitors with a working equivalent before agents go
Retention
Days of metrics, logs and traces kept, by data type

Questions

Frequently asked migration questions

No. OpenTelemetry with Grafana Tempo or Jaeger gives you end-to-end traces and waterfall views. The UI differs, and some Datadog-specific features need a replacement or a decision to drop them.

Yes. Grafana Alerting and Prometheus Alertmanager integrate with PagerDuty, Opsgenie, Slack and custom webhooks.

Rehearse the cutover before the real one

Tell us about your data volume, traffic and timeline. An engineer will reply within one business day to set up a call about the migration plan and its rollback path.