Datadog

The leading cloud monitoring and observability SaaS platform — provides infrastructure metrics, APM, log management, and alerting in a single unified interface.

Datadog is the dominant commercial observability platform, used by over 27,000 organizations to monitor infrastructure, application performance, logs, security, and user experience in a single unified interface. Unlike the self-hosted Prometheus/Grafana stack, Datadog is a fully managed SaaS product — teams instrument their services with the Datadog agent or APM libraries and get dashboards, alerts, and distributed tracing out of the box. It is deployed at Airbnb, DoorDash, Samsung, and most technology companies above a certain scale. Datadog experience is listed on a high proportion of DevOps, SRE, and platform engineering job postings.

Typical time to job-readiness: ~2 weeks.

Learning Datadog

Beginner

Start a free Datadog trial and install the agent on a server or container. Explore the Infrastructure and Metrics pages — understand how tags work (they are Datadog's core organizational primitive). Set up a basic alert on CPU or memory and route it to Slack or email.

Intermediate

APM (Application Performance Monitoring): instrument a service with a Datadog tracing library, read flame graphs and trace waterfalls to identify latency bottlenecks, and use Log Management to correlate logs with traces. Learn Datadog's query language for building custom dashboards with composite metrics.

Advanced

Synthetic monitoring for uptime testing, custom metrics with DogStatsD, SLO tracking, Datadog Monitors with composite alert logic, and cost optimization (controlling metric and log ingestion volume). SRE interviews often include a scenario: 'latency spiked on a service — walk me through how you'd diagnose it' — be ready to walk through Datadog APM, traces, and logs in that workflow.

Key concepts

  • Infrastructure monitoring: host metrics (CPU, memory, disk, network) from the Datadog agent
  • APM (Application Performance Monitoring): distributed tracing that shows latency across services and endpoints
  • Logs: centralized log collection, parsing, and search — correlated with traces via trace ID injection
  • Tags: the core organizational primitive in Datadog; tag by env, service, team, and version for filtering
  • Monitors and Alerts: threshold or anomaly-based alerts that route to PagerDuty, Slack, or email
  • The four golden signals: latency, traffic, errors, and saturation — the standard starting point for any service dashboard

Common interview topics

  • How would you instrument a new microservice for observability in Datadog
  • Walk me through how you'd investigate a latency spike using Datadog
  • What is APM and how does distributed tracing work
  • How do you set up meaningful alerts that don't cause alert fatigue
  • What are the four golden signals and how do you monitor each one

Browse Datadog jobs