Datadog on Kubernetes: A Practical Setup That Pays for Itself
Datadog can see almost everything in your platform. The problem is rarely missing data - it’s too much disconnected data, noisy alerts, and a bill that grows faster than your traffic. A few decisions made early make the difference between a tool your engineers trust and an expensive wall of dashboards.
This is the setup we recommend for running Datadog on Kubernetes.
Install the Agent the Kubernetes way
Deploy the Datadog Agent as a DaemonSet (one per node) together with the Cluster Agent, which collects cluster-level data and takes load off the Kubernetes API server. Either the Datadog Operator or the official Helm chart will do this; pick one and keep its configuration in Git.
A minimal set of Helm values:
datadog:
apiKeyExistingSecret: datadog-secret # never commit the API key
site: datadoghq.eu # EU region, if data residency matters
clusterName: prod-eu-west-2
logs:
enabled: true
containerCollectAll: true
apm:
portEnabled: true
clusterAgent:
enabled: true
For UK and EU businesses, choosing the EU site (datadoghq.eu) keeps
your telemetry in the EU, which is often simpler for data-protection reviews.
The site is fixed per Datadog organisation, so decide before you start.
Unified service tagging: the step everyone skips
Datadog becomes genuinely useful when you can jump from a spike on a
dashboard to the traces behind it, and from a slow trace to its logs. That
only works if metrics, traces and logs share the same three tags: env,
service and version. Datadog calls this unified service tagging.
On Kubernetes, set them as labels on the Deployment and its pod template, and pass them to the tracer as environment variables:
apiVersion: apps/v1
kind: Deployment
metadata:
name: checkout
labels:
tags.datadoghq.com/env: prod
tags.datadoghq.com/service: checkout
tags.datadoghq.com/version: "1.4.2"
spec:
template:
metadata:
labels:
tags.datadoghq.com/env: prod
tags.datadoghq.com/service: checkout
tags.datadoghq.com/version: "1.4.2"
spec:
containers:
- name: checkout
env:
- name: DD_ENV
valueFrom:
fieldRef:
fieldPath: metadata.labels['tags.datadoghq.com/env']
- name: DD_SERVICE
valueFrom:
fieldRef:
fieldPath: metadata.labels['tags.datadoghq.com/service']
- name: DD_VERSION
valueFrom:
fieldRef:
fieldPath: metadata.labels['tags.datadoghq.com/version']
Set version from your CI pipeline (the image tag or Git SHA). You then get
deployment tracking for free: Datadog can compare error rates and latency
between versions, so a bad release shows up within minutes.
Traces and logs that link up
With the Agent’s admission controller enabled, Datadog can inject the tracing library and the environment variables above into your pods automatically - or you can add the tracer to each service yourself. Either way, turn on log injection so every log line carries the trace ID. From a slow request in APM you can then open exactly the logs it produced, instead of searching by timestamp.
If you run a service mesh, enable the Datadog Istio integration too: the Envoy proxies emit request metrics for every service, which gives you consistent golden signals even for services you haven’t instrumented yet (see our Istio post).
Alert on symptoms, not causes
The fastest way to lose trust in monitoring is paging people for things users never notice - CPU at 85%, a single pod restart. Instead:
- Define SLOs for your key user journeys - for example, “99.9% of checkout requests succeed in under 500 ms over 30 days”. Datadog supports both metric-based and monitor-based SLOs.
- Alert on error-budget burn rate. A burn-rate alert pages when you’re consuming error budget fast enough to miss the SLO, catching real incidents quickly while ignoring short, harmless blips.
- Keep cause-based monitors as warnings, not pages. Disk filling up is worth a ticket; it’s not worth waking someone at 3am unless it’s about to hurt users.
- Every page needs an owner and a runbook link. Use the monitor message to say what the alert means and what to check first.
Keep the bill under control
Datadog pricing tracks what you send and keep, so costs can quietly climb. The biggest levers:
- Custom metric cardinality. Each unique combination of tag values is a
separate custom metric. A tag like
user_idorrequest_idon a metric can multiply its cost thousands of times. Keep high-cardinality values in traces and logs, not metric tags. - Log indexing. Ingesting logs and indexing them are billed separately. Use exclusion filters to drop noisy, low-value logs (health checks, debug output) from indexes, and send everything to a cheap archive in S3 if you need it for compliance.
- Trace sampling. You rarely need 100% of traces. Use ingestion controls to sample high-volume, healthy traffic while keeping errors and slow requests.
- Review usage monthly. The usage and cost pages show which services and tags drive spend - most savings come from a few noisy sources.
Where to start
If Datadog is already running but not delivering, the usual order is: fix tagging, define two or three SLOs for the journeys that matter, replace noisy monitors with burn-rate alerts, then tackle cost. Each step makes the next one easier.
We help teams set up Datadog from scratch and rescue installations that have become noisy or expensive - see our observability service or get in touch.