Software company Atlassian has published an account of its migration from gostatsd to OpenTelemetry, replacing the internal workings of a metrics platform that receives data from about 100,000 hosts across 14 regions. The existing service had a 99.95% SLO, so the team had to change the platform without disrupting the metrics used by production alerts.
In a post on the CNCF blog, Iris Grace Endozo, Farzad Vazirnia and Albert Kerr explain that gostatsd had worked reliably for years, but its UDP-only design could not handle traces or logs. More services were also producing OpenTelemetry data. Continuing with gostatsd would have required Atlassian to reproduce features already being developed by the OpenTelemetry Collector community.
The team avoided asking thousands of services to move immediately from StatsD clients to the OpenTelemetry SDK. It retained the existing StatsD-over-UDP interface and rebuilt the pipeline behind it. The authors say this changed the work from an organisation-wide application migration into a migration owned by the platform team.
"We kept the interface and rebuilt everything behind it, which turned an org-wide migration into a platform-team migration."
Iris Grace Endozo, Farzad Vazirnia and Albert Kerr, Atlassian

Atlassian introduced purpose-built Collector distributions for collection, ingest, aggregation and forwarding. This reflects the Collector's pipeline model, in which receivers accept data, processors modify it and exporters send it to one or more destinations. Dividing the platform into four stages allowed the engineers to change one part without replacing the others.
For collection, Atlassian replaced the gostatsd sidecar with the Collector distribution already used by its tracing team. Applications continued sending StatsD packets to the same address, while an OTLP receiver allowed newer services to send OpenTelemetry metrics. Folding metrics into the tracing sidecar saved an average of 3.9% CPU for each of Atlassian's most expensive Micros services. The team estimates that it cut sidecar cost by about 30% across the fleet. An OpenTelemetry Lambda extension preserves the same interface for serverless workloads, where a sidecar cannot run.
The ingest tier presented a different problem because metric aggregation is stateful. Every point for a time series must reach the same aggregator. Atlassian's old nomad proxy hashed each service and environment to one shard, which concentrated the largest services on a few busy replicas. The new system uses the OpenTelemetry contrib load-balancing exporter to hash by stream ID. Individual time series remain together, while one large service can be spread across the pool. Atlassian reports a more even CPU distribution, fewer hot shards and better off-peak scaling.
Aggregation produced the largest reduction in data volume. The platform receives about 4.8 billion data points each minute and stores about 220 million, a reduction of roughly 96%. Since the upstream components did not aggregate delta metrics in the way Atlassian required, the company developed and open-sourced its own aggregation processor. The revised tier uses about half as much CPU for the same traffic.
"Small tests and benchmarks were not enough; continuous profiling in prod is what actually told us where to optimise."
Iris Grace Endozo, Farzad Vazirnia and Albert Kerr, Atlassian
The migration guidance focuses on operational practice. Atlassian began with development and staging workloads, chose early adopters that had something to gain, and then increased rollouts through 1%, 10%, 50% and 100% stages. The team also recommends preserving familiar operational workflows while old and new systems run together, and profiling under production load rather than relying on small benchmarks. Atlassian says gostatsd aggregation and nomad account for about 38% of CPU requests in its metrics clusters, with nomad alone using about 13% of total resources. Removing them will complete the end-to-end OpenTelemetry pipeline. The next planned step is to move application instrumentation from StatsD, DogStatsD and vendor clients to OpenTelemetry SDKs.
In the final stage, Atlassian replaced a bespoke forwarding service with a stateless Collector distribution called metrics-gateway. Upstream exporters now fan out data to destinations including SignalFx and S3, while the Collector supplies retry, queueing and backpressure behaviour. The authors say adding another destination is now a configuration change rather than a separate integration project.
The approach differs slightly from Airbnb's recent OpenTelemetry metrics migration. Airbnb used dual emission from a shared metrics library and selected VictoriaMetrics' vmagent for streaming aggregation, eventually processing more than 100 million samples per second. Atlassian instead preserved the StatsD endpoint while moving the pipeline beneath it. Both projects separated application migration from the replacement of central infrastructure.
Earlier InfoQ coverage of Skyscanner's observability overhaul described another route. Skyscanner standardised instrumentation and transport on OpenTelemetry, then moved more than 300 microservices by updating a common library. Atlassian's compatibility layer addresses an environment where immediate re-instrumentation would have been a longer and riskier programme.
In a translated conference report, GREE monitoring infrastructure lead Iwahori called Atlassian's approach "well-thought-out". Iwahori particularly valued the stream-ID routing and custom delta aggregation processor, but raised a practical question about scaling a stateful aggregation tier. The report says Atlassian's answer indicated that this part relied on manual configuration.