Skip to content

Datadog bills driven up by custom metrics and logs

Datadog bill reduction and ingestion governance

Find what drives your Datadog bill, then cut it at the source — high-cardinality tags, noisy logs and over-sampled traces — while keeping the signals on-call depends on.

Symptoms

Signs your platform has this problem

If several of these sound familiar, the plan below is where we would start.

01

Observability costing more than compute

Invoices spiking from log volume and custom metric tags.

02

High-cardinality tags

User IDs or transaction hashes exported as metric tags, creating thousands of billable custom metrics.

03

Noisy logs

Debug and health-check logs sent to Datadog without exclusion filters.

Remediation plan

How we fix it, step by step

Each phase ends with a measurement, so you can see what changed before the next one starts.

01

Usage and cardinality audit

Using Datadog's usage pages to find the biggest custom metric and log cost drivers.

02

Agent-level log filters

Dropping health checks and debug noise in the Datadog Agent, before ingestion.

03

Metric tag allow-lists

Pruning high-cardinality tags and keeping only the dimensions you query.

04

Trace sampling

Keeping every error trace while sampling successful requests at a lower rate.

Technical checklist

Remediation checklist

What we check before a change goes to production:

  • Audit high-cardinality metric tags and configure tag allow-lists
  • Add Agent-level log exclusion filters for health checks and debug noise
  • Keep all error traces and sample successful requests at a lower rate
  • Set up usage monitors and budget alerts

What we measure

We take a baseline first and report the same measurements after each change, from your own monitoring — evidence, not promised results.

Monthly bill
Datadog invoice by product (logs, APM, metrics), before and after
Custom metrics
Billable custom metric count
Alert coverage
Monitors and SLOs still firing correctly after the changes

Related service

Cloud & DevOps

Cloud cost optimization, Kubernetes platforms, and CI/CD that make deploys boring — savings and reliability measured in your dashboards, not our deck.

Explore Cloud & DevOps

Questions

Questions about this remediation

Usually log volume surges, user IDs exported as custom metric tags, or full APM trace sampling on high-traffic endpoints.

No. We filter out repetitive health-check and debug noise while keeping every warning, error and exception log.

Want an engineer to look at this with you?

Send us the symptoms and any metrics you have. We'll reply within one business day, set up a call and agree what to measure before anything changes.