Grab has migrated the storage backend of its Counter Service from a wide-column database to Aerospike, redesigning its data model and introducing a staged traffic migration to avoid downtime. The service supports Grab’s anti-fraud platform, handling tens of thousands of queries per second and about one billion requests per day. Grab reports roughly 50% lower production p99 read latency, a reduction in on-disk data from about 3 TB to 1 TB, and 45% to 50% lower cost per node.
Counter Service answers time-windowed queries such as recent ride requests or failed payment attempts. Its original design stored 15-minute, hourly, and daily counts as separate rows. Each incoming event triggered three parallel reads, an in-memory increment, and a batch write, resulting in four network round trips. The existing wide column model used clustering columns to efficiently retrieve rows within a time range, a pattern documented by Apache Cassandra.
Rather than directly replacing the storage implementation, Grab first separated storage access from the Rust service’s business logic. A storage facade exposed legacy, Aerospike, and mock implementations, while configuration controlled single-backend, shadow, and split-traffic modes. Shadow reads were gradually increased from 5% to 20%, 50%, and 100%, with parity measured through existing metrics before live traffic was shifted to Aerospike.
Daniel Lim, an engineering leader at Grab, described the approach on LinkedIn as
Combining dual read-write paths for traffic shadowing and data parity validation with gradual traffic migration.

Grab’s Counter Service migration architecture (Source: Grab Blog Post)
The larger architectural change was the data model. Grab initially evaluated retaining a row per time bucket, but Aerospike maintains approximately 64 bytes of primary index metadata per record. At the record counts required by the service, that overhead made record cardinality a significant capacity consideration.
Grab instead consolidated all buckets for a counter and granularity into a single record containing a timestamp-ordered map. Writes use Aerospike’s atomic map increment operation, while stale entries are explicitly removed. The redesign reduced record count by more than an order of magnitude, according to Grab, and contributed to the lower index and disk footprint.
The read path also changed. The original backend returned paginated rows filtered by time range on the server. Aerospike returns records containing maps, which Grab filters on the client. Its batch API groups multiple record operations into requests to database nodes, reducing the number of network requests for the service’s multi-key access pattern. Aerospike documents that batching can combine multiple record commands into a request to each relevant node.
During rollout, Grab also encountered issues with the Aerospike Rust client, including its initially synchronous API and DNS handling during cluster replacement. The team subsequently adopted the asynchronous client and worked with upstream maintainers on the DNS issue before continuing the rollout.
Grab reports that the migration completed without downtime or data integrity issues. The resulting architecture combines a storage abstraction, shadow validation, gradual traffic shifting, and a data model designed around Aerospike’s record and batch access characteristics.