- Source
- Huawei Cloud DDS 4.0 (replica set and cluster)
- Target
- Percona Server for MongoDB 8.0 on Huawei Cloud CCE
- Transport
- Kafka-compatible broker (Redpanda, 3 brokers)
- Approach
- Full load + change data capture
- Largest single run
- 100,000,000 documents with active writes
- Production scope
- 19 databases in steady CDC
Client: a platform team running production MongoDB on Huawei Cloud. The client and its database names are not named in this write-up.
Problem
The team ran its MongoDB workloads on Huawei Cloud DDS 4.0. They wanted to move to Percona Server for MongoDB 8.0 on Kubernetes, under their own control, without a long write freeze. An in-place upgrade path does not exist between a managed DDS instance and a self-managed cluster, and a 4.0 to 8.0 jump would otherwise mean several major-version upgrades in a row.
Environment
- Source: Huawei Cloud DDS 4.0, both cluster and replica-set deployments.
- Target: Percona Server for MongoDB 8.0 on Huawei Cloud CCE: a sharded cluster plus two isolated replica-set targets, so that specific workloads stay separated.
- Transport: a three-broker Redpanda cluster (Kafka API) on the same CCE cluster.
- Observability: Prometheus, Prometheus Adapter and Percona Monitoring and Management (PMM).
- Delivery: GitLab CI with a runner on CCE, container image scanning and SBOM generation.
Challenges
- Change streams on DDS. DDS rejected change streams that request full-document lookups, with
ReadConcernMajorityNotEnabled. The standard CDC approach could not be used as-is. - Large documents. Some collections held documents of around 15 MiB, above default Kafka producer batch limits.
- Live writes. The source kept accepting writes throughout, so online count comparisons can legitimately differ for a moment.
- Major-version jump. Indexes, validators and collection options had to be recreated and compared on a version four majors newer.
Solution
We used MYC's in-house migration engine, the same engine behind our MongoDB migration service:
- Segmented, parallel full load into per-collection Kafka topics, with CDC armed before the snapshot so no write is missed.
- Idempotent apply by document
_id, so replays and retries are safe. - BSON preserved end to end. Documents never pass through JSON on the data path.
- A delta change-stream mode for DDS: when full-document lookup is not available, the engine applies the ordered update descriptions instead. A polling mode exists as a last resort.
- Metadata discovery and synchronization for indexes, validators and collection options, followed by a validation job that gates cutover.
- Workers that autoscale on Kafka consumer lag through the Horizontal Pod Autoscaler and Prometheus Adapter.
Architecture
Huawei DDS 4.0 Kafka-compatible broker Percona Server for MongoDB 8.0
(replica set / cluster) (Redpanda, 3 brokers) on Huawei Cloud CCE
┌──────────────────────┐ ┌───────────────────────────┐ ┌────────────────────────────┐
│ source databases │─────►│ migration.<db>.<coll> │─────►│ sharded cluster │
│ │ full │ cdc.<db>.<coll> │apply │ + isolated replica-set │
│ change streams │ load │ migration-dlq │ by │ targets │
│ (delta fallback) │─────►│ key = document _id │ _id │ │
└──────────────────────┘ CDC └───────────────────────────┘ └────────────────────────────┘
▲ ▲ ▲ ▲
│ fullsync / cdc consumer indexer · validator
└──────────── controller (state machine, checkpoints, metadata phases) ───────────┘
workers autoscale on Kafka consumer lag (HPA + Prometheus Adapter)
Implementation
- Preflight inside the cluster. Mongo and Kafka preflight jobs checked connectivity, TLS, privileges and destination writes. Each production database returned 14 passes, 1 warning (the expected 4 to 8 major-version review) and no failures.
- Proof runs. Snapshot-only and CDC smoke migrations proved the delta fallback, including convergence of updates, deletes and inserts made after CDC was armed.
- Scale test. A 100-million-document run with active CDC churn reached completion with data and metadata validation passing.
- Large-document fix. The producer, consumer, topic and preflight message-size limits were made configurable (20 MiB default). Two affected collections resumed from checkpoints and later validated with matching counts and checksums.
- Production waves. 19 production databases completed their full load and were held in steady CDC with zero collection failures and zero aggregate Kafka lag at the verification point in July 2026.
Results
- 100,000,000 documents migrated in a single validated run.
- 19 production databases kept in sync through CDC, with zero aggregate lag at verification.
- Proof migrations from both DDS cluster and DDS replica-set sources completed with data and metadata validation passing and a successful cutover.
- A 4.0 to 8.0 major-version jump in one migration, with the source left untouched as a fallback.
A note on validation: for three databases, online count checks differed by a handful of documents because the source was still taking writes between the two count calls. Those reports were not marked as passed. The cutover procedure is to freeze writes, drain lag to zero and re-run validation before switching traffic.
Technologies
Huawei Cloud DDS, Huawei Cloud CCE, Percona Server for MongoDB 8.0, Percona Operator, Redpanda (Kafka API), Kubernetes HPA, Prometheus, Prometheus Adapter, PMM, GitLab CI, Trivy.
Services used
MongoDB migration service · Cloud migration · Huawei Cloud · Kubernetes