Distributed RDBMSs – a view from the bridge

I have started appreciating the recent crop of distributed RDBMSs that engineers have developed specifically to respond to the technical needs of China and Russia. A few months of running Clickhouse in my homelab as a bit of a “lakehouse” for my economic history analytics environment has reaffirmed my confidence in the gaggle of performance-obsessed Russians with time and money. They created a technically formidable power tool.

Most of these fall into the category of “cool power tools that I have not had a reasonably occasion to use because Alibaba’s problems are not my problems.” The pieces of software that I understand in most detail are the ones that I use in a personal capacity or have been forced to learn for work.

For the purposes of this post, by RDBMS, I mean a system with the following capabilities.

  • CRUD operations
  • SQL
  • Joins
  • Indexes
  • Transaction Isolation control/declarative concurrency semantics
  • Transactional processing/OLTP

We will save analytic engines, document databases, key-value stores, etc. for another day.

A Common Progression

We start with the single-node Postgres/MySQL/RDBMS deployment. This is wonderful in the homelab or for situations where we can tolerate a few hours of downtime. A homelab really never needs anything more complicated than a single-node Postgres instance. A single instance can even serve a small business even having hundreds to a few thousand users with sufficient tuning and optimization

Then we may add some caches and proxies and asynchronous replication to handle larger read loads. This has been a very well understood/“solved” operational pattern probably since the 1980s but certainly by the early 2000s.

But eventually the day will come when either no single node can handle all of the data and/or no single node can handle all of the write-load. This is where the fun begins because tradeoffs reflecting the operational requirements/constraints must be made. While companies like Oracle, Microsoft, Sybase have dominated this space with closed-source proprietary solutions since the 1980s, atleast in America and Europe, the 1990s and early 2000s saw the fully open source MySQL and Postgres really come into their own. The fact that they were well-engineered enough for such a large amount of use cases while remaining free for 30 years should baffle and amaze everyone. And don’t even get me started on Sqlite!

But once we require a real distributed system, we also require formidable amount of engineering effort to iron out the interactions between transactions/isolation levels, consensus, more complex network toloplogies, leader election/failover, sharding, replication, etc.

Many of these baked for 10+ years within the walls of some of the largest corporations in the world. Especially since 2015, many companies have open sourced these (or atleast “public-sourced” these) largely for “clout” and recruiting.

GPT 5.6 Summary: Comparative guide to YDB, TiDB, OceanBase, PolarDB-X, openGauss, AliSQL, SQL Server, Oracle, CockroachDB, and Comdb2

Snapshot date: July 15, 2026.

A warning about the “enterprisiness” scale: these levels measure architectural and operational fit, not simply database quality. A mature single-writer Oracle or SQL Server installation can be safer for a bank than a poorly operated distributed cluster. Conversely, distributed SQL becomes compelling when one machine, one write region, or manual sharding is the actual constraint.

Also, CAP applies meaningfully only during a network partition. Consensus databases generally choose consistency over availability for the affected shard/range: a minority partition stops accepting writes. Traditional databases with asynchronous replicas may remain locally available, but can lose acknowledged transactions during forced failover.

Executive summary

Database Origin; development/open source Fundamental architecture SQL affinity Isolation Battle-tested / velocity Sweet spot Enterprise range
YDB Yandex; started 2014; OSS 2022 Shared-nothing; automatic range sharding; custom distributed storage; per-tablet consensus; distributed serializable transactions YQL; emerging PostgreSQL compatibility Serializable default; snapshot/read-only and relaxed-read modes Very strong inside Yandex; high velocity Huge, high-throughput, strongly consistent operational systems where schema can follow the primary key 2–4, best 3–4
TiDB PingCAP founders Max Liu, Dylan Cui, Edward Huang; started/open from 2015 Stateless SQL over TiKV; key ranges replicated with Raft; PD control plane; TiFlash columnar replicas High MySQL compatibility Snapshot isolation advertised as RR; RC; optimistic/pessimistic transactions Strong large-scale production record; very high velocity Replacing large sharded-MySQL estates; HTAP; hundreds of TB to PB 2–4, best 3–4
OceanBase Alibaba/Ant Group; started 2010; OSS 2021 Shared-nothing equal nodes; tablets/log streams; Multi-Paxos; GTS; distributed SQL/2PC; LSM row/column storage MySQL and Oracle modes RC; snapshot-style RR/Serializable, but not strict serializability Extremely battle-tested in Ant/Alipay; high velocity Financial-grade distributed OLTP, especially in Chinese/APAC ecosystems 1–4, best 3–4
PolarDB-X Alibaba Cloud; TDDL/DRDS lineage from early 2010s, product 2019/2.0 in 2021; OSS 2021 Stateless compute; sharded DNs with Paxos; GMS/TSO; 2PC; global indexes; CDC and columnar nodes High MySQL compatibility RC and RR Strong Alibaba/Double 11 lineage; high velocity Distributed MySQL with global indexes and cloud-native operations, especially on Alibaba Cloud 2–4, best 3–4
openGauss Huawei, PostgreSQL-derived; product lineage predates OSS; OSS 2020 Primarily single-node or primary/standby; optional resource-pooling/distributed solutions; WAL replication PostgreSQL lineage plus compatibility modes, but not drop-in PostgreSQL Primarily RC/RR semantics; verify exact compatibility mode/version Strong Chinese enterprise/banking adoption; medium-high velocity Enterprise PostgreSQL-like replacement on Huawei/Kunpeng and Chinese domestic stacks 1–4, core best 2–3
AliSQL Alibaba MySQL branch; development lineage around 2010; public in 2016, renewed OSS release in Dec. 2025 Conventional MySQL/InnoDB primary-replica; current branch adds DuckDB engine and HNSW vectors Very high MySQL compatibility MySQL/InnoDB RU, RC, RR, Serializable Immensely tested as Alibaba/RDS lineage; public project history uneven, currently active High-performance conventional MySQL, especially Alibaba/RDS-derived environments 1–3; 4 only with extensive surrounding architecture
SQL Server Microsoft, Sybase, Ashton-Tate; first release 1989; proprietary Single writable instance; AG log replication; FCI shared storage; no transparent native sharding T-SQL, proprietary RU, RC, RR, Serializable, Snapshot, RCSI Exceptionally battle-tested; high commercial velocity Regional/national enterprise OLTP, BI, Microsoft estates 1–4, best 2–4
Oracle Database Oracle/RSI; development 1977–79; proprietary Single instance; RAC shared-everything; Data Guard; optional shared-nothing Oracle Sharding Oracle SQL/PL/SQL RC, Serializable snapshot-style, Read Only The most institutionally battle-tested; high commercial velocity The hardest mission-critical bookkeeping, complex SQL, packaged enterprise apps 1–4, best 2–4
CockroachDB Spencer Kimball, Peter Mattis, Ben Darnell; started OSS in 2014; company 2015 Symmetric SQL nodes; KV ranges; Raft; MVCC; HLC; automatic sharding and Parallel Commits High PostgreSQL wire/syntax affinity, incomplete extension compatibility Serializable default; RC Strong and growing production record; very high velocity Multi-region/global OLTP requiring SQL, active-active-like access, and strong consistency 2–4, best 3–4
Comdb2 Bloomberg; started 2004; OSS around 2016 Full-copy replicated clusters; any node executes SQL, writes coordinated through one master; Berkeley DB storage Custom/SQLite-derived SQL, low PG/MySQL compatibility Default/RC, Snapshot, Serializable; optional linearizable configuration Very battle-tested inside Bloomberg; moderate public velocity Many modest-to-large highly available databases in a tightly controlled environment 1–3; 4 mainly with Bloomberg-level expertise

Architectural families

These products are easier to understand as four families.

Native shared-nothing distributed SQL

Clients
   │
SQL coordinators / any-node SQL
   │
Distributed transaction layer
   │
Shard A       Shard B       Shard C
Raft/Paxos    Raft/Paxos    Raft/Paxos
replicas      replicas      replicas

YDB, TiDB, OceanBase, PolarDB-X, and CockroachDB primarily belong here. Data is automatically divided into many independently replicated consensus groups. They scale writes by placing leaders for different shards on different machines.

Traditional primary plus replicas

                    ┌── read replica
Application ──► primary ── read replica
                    └── DR replica

AliSQL, SQL Server Availability Groups, openGauss primary/standby, and Oracle Data Guard primarily use this model. It is simpler, gives excellent single-shard transactions, and often offers better latency per transaction, but writable capacity is ultimately bounded by one primary unless the application is manually sharded.

Shared-everything cluster

Instance A ─┐
Instance B ─┼── shared database files/storage
Instance C ─┘
       cache-coherency protocol

Oracle RAC is the major example. Multiple instances can modify one database, with Cache Fusion maintaining block-cache coherence. This scales connections and selected workloads and provides excellent availability, but is not equivalent to shared-nothing linear write scaling.

Replicated whole-database cluster

SQL on node A ─┐
SQL on node B ─┼── one elected write master
SQL on node C ─┘       │
                synchronous physical logs

Comdb2 fits here. It distributes query execution and availability, but each database is normally replicated as a whole rather than automatically divided into independently writable shards.


RDBMS-by-RDBMS Overview

1. YDB

YQL / PostgreSQL endpoint
          │
   distributed query layer
          │
  range-sharded table tablets
          │
custom distributed block/blob storage
replication or erasure coding across failure domains

YDB was created by Yandex, with the first commit dated January 10, 2014, and was released under Apache 2.0 in 2022. It grew from Yandex’s internal storage systems and now underpins storage and processing for a broad range of Yandex services. The engineering team identifies Andrey Fomichev as founder/CTO of YDB. (ydb.tech )

YDB is a shared-nothing, actor-based distributed database. Row tables are range-sharded by primary key; shards are implemented as stateful “tablets” that can move between machines. The system automatically splits, moves, and balances shards. Tablets persist into YDB’s own distributed storage subsystem rather than RocksDB or a conventional filesystem database. Replication and failure-domain placement occur below the table layer, while cross-shard transactions are coordinated above it. YDB’s distributed transaction work is influenced by Calvin-style deterministic transaction ideas, though it is not simply a textbook Calvin implementation. (ydb.tech )

The default read-write transaction level is Serializable, including transactions spanning multiple tables and shards. YDB also provides snapshot/read-only and less-consistent modes for workloads that can trade freshness or isolation for speed. Cross-shard transactions are real ACID transactions, but naturally cost more than single-shard operations; YDB’s documentation gives a broad 20–500 ms range for distributed transactions versus much lower latency for point operations. Its CAP position is therefore CP per affected tablet: a tablet without a viable replica quorum stops making progress rather than accepting divergent writes. (ydb.tech )

YQL is recognizably SQL, but YDB is not as transparent a migration target as TiDB for MySQL or CockroachDB for PostgreSQL. PostgreSQL compatibility has been expanding, but PostgreSQL extensions, exact planner behavior, system catalogs, procedural ecosystem, and operational tooling should not be assumed compatible. The most important “index” remains a well-designed primary key that places related operations in the same key range. Secondary, vector, and full-text indexes exist, but a poor distribution key can turn ordinary operations into expensive cross-shard work.

Sweet spot: very large operational databases, event deduplication, profiles, counters, metadata, streaming-plus-table transactions, and high-throughput services with predictable key-based access. It is especially attractive when you need serializable transactions at enormous scale and are willing to design around distributed-system locality. It is less attractive for a PostgreSQL application expecting arbitrary extensions, ad hoc relational access, or a large Western vendor ecosystem.

Battle-tested: extremely strong inside Yandex, with public material describing million-RPS workloads and clusters scaling to thousands of servers and hundreds of petabytes. External adoption and support breadth are much smaller than Oracle, SQL Server, TiDB, or CockroachDB. (ydb.tech )


2. TiDB

MySQL clients
     │
stateless TiDB SQL nodes
     │
 PD: metadata, placement, timestamps
     │
TiKV regions ── Multi-Raft replicas ── RocksDB
     │
optional TiFlash columnar replicas

TiDB was started in 2015 by PingCAP founders Max Liu, Dylan Cui, and Edward Huang as an open-source project. Version 1.0 was declared production-ready in October 2017. It has one of the healthiest development communities in this comparison; TiKV is also a CNCF graduated project. (pingcap.com )

The SQL layer consists of stateless TiDB servers speaking the MySQL protocol. SQL rows and indexes are encoded into keys in TiKV, an LSM-based transactional KV store using RocksDB. TiKV divides the keyspace into relatively small Regions, each independently replicated using Raft. The Placement Driver, or PD, stores cluster topology and Region metadata, performs scheduling, and issues transaction timestamps. TiFlash maintains columnar replicas for analytical workloads, allowing TiDB to combine OLTP and near-real-time analytics. (docs.pingcap.com )

Transactions use a Percolator-family timestamp and two-phase-commit model, with optimizations such as one-phase commit and asynchronous commit. TiDB supports optimistic and pessimistic transactions, with pessimistic mode the default. Its advertised REPEATABLE READ is actually snapshot isolation, and therefore differs from both ANSI Repeatable Read and MySQL/InnoDB Repeatable Read. Read Committed is also available, but TiDB should not be treated as providing CockroachDB-style strict serializable isolation. This is its most significant ACID tradeoff: atomicity, durability, and distributed consistency are strong, but default isolation permits snapshot-isolation phenomena such as write skew unless applications lock or structure transactions appropriately. (docs.pingcap.com )

TiDB’s strongest indexing techniques are clustered primary keys, carefully ordered composite secondary indexes, covering indexes, partition pruning, and—when scans are unavoidable—TiFlash columnar replicas. Secondary indexes are globally distributed KV records, so an index lookup followed by a base-row lookup may cross Regions. Avoid random or monotonically hot keys unless the schema and partitioning strategy account for them. Placement policies can control replicas at database, table, partition, region, zone, rack, and host levels. (docs.pingcap.com )

TiDB is highly MySQL-compatible at protocol, syntax, and driver level, but it is not MySQL internally. Differences exist in isolation semantics, unsupported features, foreign-key/version history, optimizer behavior, stored routines, DDL, transaction limits, and operational tooling. Migration is usually much easier than moving a MySQL application to YDB or CockroachDB, but compatibility testing remains mandatory.

Sweet spot: consolidating dozens or hundreds of sharded MySQL databases; high-volume SaaS, fintech, logistics, billing, ledgers that do not require strict serializability, and HTAP systems needing fresh analytics without a separate ETL copy. It is best within one region or a small number of nearby regions/AZs. Cross-continent replica placement and follower/stale reads are possible, but globally distributed writes pay timestamp, Raft, and 2PC latency; CockroachDB generally has stronger first-class global-locality abstractions.

Battle-tested: WeBank reports more than 80 clusters, 1.3 PB across nearly 1,000 servers, with its largest cluster above 200 TB and 237,000 QPS. This is credible national-bank-scale deployment evidence, although vendor case-study numbers should not be read as independent benchmarks. (pingcap.com )


3. OceanBase

        any OBServer
SQL + optimizer + transaction engine
        │
tablets grouped into log streams
        │
 Multi-Paxos replicated logs
        │
LSM MemTables + SSTables
row / column / mixed storage

OceanBase began development at Alibaba in 2010 and later became central to Ant Group and Alipay. It was officially open-sourced on June 1, 2021. Its architecture evolved from an early single-write/multi-read system to a fully distributed system around the V1 lineage and then to V4’s integrated standalone/distributed design. (oceanbase.github.io )

OceanBase uses a shared-nothing equal-node architecture. Every OBServer contains SQL, transaction, replication, and storage engines. Tables are split into tablets; tablets are assigned to log streams, which are replicated with Multi-Paxos. A Global Timestamp Service supplies transaction and read versions, and distributed transactions spanning log streams use an optimized two-phase commit protocol. Multi-tenancy is unusually deeply integrated: each tenant acts much like an independent logical database with its own placement and resource configuration. (en.oceanbase.com )

Storage is LSM-tree based: writes initially enter MemTables and are later compacted into SSTables. Recent releases support integrated row and column storage, vectorized execution, and HTAP. The LSM design provides strong write throughput and compression but introduces compaction management, write amplification, and possible read amplification. Primary-key design, local/global indexes, covering composite indexes, partition pruning, and keeping related rows within the same tablet/log-stream locality are the main tools for minimizing physical reads and distributed execution. (en.oceanbase.com )

OceanBase exposes both MySQL-compatible and Oracle-compatible modes. Community Edition historically emphasized MySQL mode, supporting most MySQL 5.7 syntax and portions of MySQL 8.0, but architectural and catalog differences remain. Isolation internally is principally Read Committed and snapshot-style Serializable: MySQL READ UNCOMMITTED maps to RC, while RR maps to the stronger snapshot level. Despite its name, OceanBase’s Serializable level is explicitly documented as not strictly serializable and can permit write skew unless explicit locks such as SELECT FOR UPDATE are used. (en.oceanbase.com )

OceanBase is a CP system for strongly consistent operations. A transaction commits after its logs are persisted by a Paxos majority. Strong and weak replica reads are both available. Cross-city and cross-region topologies are a core feature, but—as with every consensus system—placing voting replicas across continents adds WAN round trips to commits. Good deployments place a majority near the write region, use locality for tenant/table leaders, and use remote replicas for DR or appropriate read workloads.

Sweet spot: financial ledgers, payments, telecom, high-volume e-commerce, and Oracle/MySQL replacement where a customer needs automatic sharding and financial-grade HA. OceanBase’s most impressive components are its mature Paxos/log-stream implementation, integrated multi-tenancy, LSM storage, and ability to use one engine for standalone through very large distributed deployments.

Battle-tested: among the strongest distributed databases here. It has operated Ant/Alipay core workloads through Double 11 events for more than a decade. One public Alipay historical-data deployment describes more than 20 clusters, with the largest transaction-payment cluster group reaching 15 PB. (oceanbase.github.io )


4. PolarDB-X

MySQL clients
     │
stateless Compute Nodes
SQL / optimizer / MPP / 2PC / global indexes
     │
GMS: metadata + TSO
     │
Data Nodes ── Paxos replicas
AliSQL / X-Engine / MVCC
     │
CDC binlog + optional columnar nodes

PolarDB-X combines several Alibaba lineages: TDDL and DRDS for distributed SQL and sharding, AliSQL/X-DB for storage, and PolarDB cloud-native storage technology. PolarDB-X became a distinct product around 2019; PolarDB-X 2.0 arrived in 2021 and was released as open source that year. (alibabacloud.com )

Compute Nodes are stateless MySQL-speaking SQL engines. The Global Meta Service stores schemas, statistics, accounts, placement metadata, and provides a Timestamp Oracle. Data Nodes own shards and use MVCC plus Paxos replication. Distributed transactions use TSO ordering and two-phase commit. The CDC component emits MySQL-compatible binlogs, while columnar nodes maintain analytical representations. Overall, it is logically shared-nothing across Data Nodes, although cloud editions may use shared-storage technology inside a Data Node’s primary/read-replica group. (alibabacloud.com )

The standout indexing capability is the strongly consistent global secondary index. In old-school MySQL sharding, secondary indexes that do not contain the shard key usually require scatter/gather, duplicate lookup tables, or application-maintained indexes. PolarDB-X represents global indexes as distributed tables and updates them atomically with base data. This is powerful, but every extra global index adds distributed writes, storage, and 2PC work. Good partition-key selection, covering global indexes, partition pruning, pushdown joins/aggregations, and keeping high-frequency transactions shard-local remain essential.

PolarDB-X supports MySQL protocol, drivers, common syntax, and MySQL-style binlogs. It supports Read Committed and Repeatable Read. Compatibility is high but not perfect: distributed architecture imposes limits on identifiers, tables, partitions, DDL, cross-partition features, and certain SQL behaviors. (alibabacloud.com )

Its CAP stance is CP for Paxos-replicated Data Nodes and for strongly consistent distributed transactions. Cross-region deployment is viable, but performance depends heavily on where Paxos majorities, TSO, and transaction participants live. It is strongest in Alibaba Cloud or DBStack environments where its control plane, Kubernetes deployment, monitoring, backup, and migration ecosystem are available.

Sweet spot: a large MySQL application exceeding one primary, especially when global secondary indexes, online repartitioning, binlog compatibility, and mixed OLTP/OLAP matter. It is less compelling for a modest workload where ordinary MySQL would be simpler or outside an organization comfortable with Alibaba’s ecosystem.

Battle-tested: its DRDS/X-DB lineage has supported Alibaba’s Double 11 traffic and production systems across finance, logistics, energy, public service, and e-commerce. Public material claims PB to hundreds-of-PB capability, but those figures are product capacity claims, not necessarily one independently documented customer cluster. (alibabacloud.com )


5. openGauss

PostgreSQL-like client
        │
single openGauss primary
Astore heap / Ustore undo / optional column storage
        │
WAL or consensus-assisted replication
        │
synchronous/asynchronous/cascaded standbys

openGauss is a Huawei-originated PostgreSQL-derived database. Huawei announced the project in 2019 and released its source on June 30, 2020. The kernel retains substantial PostgreSQL ancestry but has diverged through NUMA, ARM/Kunpeng, security, storage, AI-assisted operations, and enterprise-HA work. LTS releases are planned every two years, with preview releases every six months. (opengauss.org )

The core openGauss architecture is closer to enterprise PostgreSQL than to CockroachDB or TiDB. A primary owns the writable database and sends WAL to synchronous, asynchronous, or cascaded standbys. It supports conventional row-store heap/Astore, the update-optimized undo-based Ustore engine, and column-oriented capabilities in relevant editions/configurations. Additional distributed or resource-pooling solutions exist, but buyers must distinguish the openGauss kernel from Huawei GaussDB, partner distributions such as MogDB, and ShardingSphere-based solutions. “openGauss supports this” may mean different product layers.

Indexing inherits much of the PostgreSQL family’s strength: B-tree, hash, GIN for multivalue/full-text-like searches, GiST-related capabilities in appropriate builds, partial indexes, expression indexes, partitioned indexes, and online index creation. The most valuable patterns are selective composite B-trees, partial indexes excluding cold or irrelevant rows, covering/index-only access where supported, GIN for arrays or full text, partition pruning, and matching indexes to the physical storage engine. (opengauss.org )

SQL compatibility is nuanced. It is PostgreSQL-derived and much closer to PostgreSQL than MySQL, but it should not be assumed binary-compatible with arbitrary PostgreSQL extensions or operational tooling. Compatibility modes provide Oracle- and other-dialect conveniences, but these do not turn it into Oracle or MySQL. Isolation behavior and exact support vary by version and compatibility mode; RC and repeatable/snapshot semantics are the common operational choices. Do not assume PostgreSQL’s modern SSI implementation merely because the syntax says SERIALIZABLE.

CAP depends on replication mode. A single node is outside the distributed CAP framing. Synchronous/quorum configurations prefer consistency and may stop or fail over when safe progress is impossible; asynchronous remote standbys trade potential RPO for distance and availability. It is a much more conventional architecture than native distributed SQL, which is often an advantage operationally.

Sweet spot: replacing PostgreSQL or Oracle in Chinese enterprise stacks, running on Kunpeng/ARM, or deploying an enterprise primary/standby database with strong domestic vendor support. The openGauss ecosystem is particularly relevant where technology sovereignty and Chinese certification/procurement are requirements.

Battle-tested: strong in Chinese banking. The Postal Savings Bank of China’s distributed core, built with openGauss and GaussDB components, was described as serving 637 million customers with design capacity for two billion daily transactions and 67,000 transactions per second. That validates the broader ecosystem, although it does not mean a single vanilla openGauss node handled the entire load. (opengauss.org )


6. AliSQL

MySQL application
      │
one writable AliSQL/InnoDB primary
      │
binlog replication / managed HA
      ├── read replicas
      └── DR replica

Current OSS branch:
InnoDB OLTP + DuckDB analytical engine + HNSW vector index

AliSQL is Alibaba’s enterprise MySQL branch. Its development lineage emerged from Alibaba’s large-scale MySQL work around 2010. The repository was made public in 2016; the current project describes a renewed open-source phase beginning in December 2025, with an AliSQL 8.0.44 release in January 2026. Thus, 2016 is the original open-source year, while 2025 marks the current revival/republication cycle. (alibabacloud.com )

Architecturally, AliSQL itself is not a transparent shared-nothing distributed database. It is a MySQL branch centered on InnoDB, conventional local transactions, redo/undo, buffer pools, B+ tree indexes, binlogs, and primary-replica topology. Alibaba historically added performance, thread-pool, hotspot, replication, backup, security, and operational improvements needed for its internal fleet and ApsaraDB RDS. Distributed scale came from surrounding technologies such as TDDL/DRDS or later PolarDB-X—not from AliSQL automatically sharding one database.

The renewed branch adds a DuckDB-backed analytical storage engine and native vector columns/indexes using HNSW. For conventional OLTP, the important indexes remain clustered InnoDB primary keys, selective composite secondary indexes, covering indexes, prefix indexes where appropriate, full-text/spatial indexes, and minimizing secondary indexes on write-heavy tables. For vector workloads, HNSW is powerful for approximate nearest-neighbor search, but carries memory, build-time, update, and recall tradeoffs. (github.com )

SQL and transactions are essentially MySQL/InnoDB territory: Read Uncommitted, Read Committed, Repeatable Read, and Serializable, with Repeatable Read normally default. ACID tradeoffs arise primarily from the replication topology, not the local engine. Asynchronous replicas can lose recent commits on forced failover; semi-synchronous or consensus-based external systems reduce RPO but add latency.

Sweet spot: organizations wanting a conventional MySQL architecture with Alibaba-derived optimizations, or experimenting with combined MySQL OLTP, embedded DuckDB analytics, and vector search. It is not the right answer if the primary requirement is transparent multi-writer horizontal scale or globally distributed ACID transactions; PolarDB-X or another distributed SQL product is the relevant comparison.

Battle-tested: the underlying Alibaba/RDS lineage is extraordinarily tested—Alibaba describes it as powering millions of databases. The caveat is that public-source development has been uneven and the 2025–26 feature-rich repository is newer than the internal production lineage. Public development velocity is currently high, but its long-term external governance remains less proven than MySQL, PostgreSQL, TiDB, or CockroachDB. (github.com )


7. Microsoft SQL Server

T-SQL clients
      │
single read/write SQL Server instance
rowstore / columnstore / memory-optimized tables
      │ transaction log
      ├── synchronous AG secondary
      ├── readable secondary
      └── asynchronous remote DR AG

Microsoft SQL Server 1.0 shipped in 1989 as a joint effort by Microsoft, Ashton-Tate, and Sybase. Modern SQL Server is Microsoft’s proprietary engine and is completely separate from the old shared Sybase code lineage. It is one of the world’s most mature enterprise RDBMS products. (learn.microsoft.com )

SQL Server is fundamentally a single writable database instance. Always On Availability Groups stream transaction-log records to up to eight secondary replicas per AG, which may be synchronous or asynchronous and optionally readable. Failover Cluster Instances instead use shared storage. Distributed Availability Groups join separate AGs across sites, but there remains only one globally writable copy. SQL Server does not transparently break an ordinary table into consensus-replicated writable shards. (learn.microsoft.com )

Its indexing toolbox is one of the strongest here: clustered and nonclustered B+ trees, included columns for covering queries, filtered indexes, computed-column indexes, partition-aligned indexes, columnstore indexes, full-text and spatial indexes, and hash/range indexes for memory-optimized tables. For minimizing row reads, a well-designed clustered key plus selective covering filtered nonclustered indexes is often decisive. Columnstore is exceptional for scans and aggregations but not a replacement for OLTP indexes. (learn.microsoft.com )

SQL Server supports Read Uncommitted, Read Committed, Repeatable Read, Serializable, transaction-level Snapshot, and Read Committed Snapshot Isolation. Locking is traditionally central, but row versioning and newer optimized-locking features significantly reduce reader/writer blocking. SQL Server’s SERIALIZABLE uses key-range locking and is genuinely strong, although it can be costly under contention. (learn.microsoft.com )

For cross-DC deployments, synchronous commit works best at metro or low-WAN latency; asynchronous AGs are generally used over long distances. Readable secondaries scale reads but do not scale writes. CAP is topology-dependent: synchronous HA is consistency-oriented and may stop safe writes/failover without quorum; asynchronous DR accepts a nonzero loss window.

Sweet spot: corporate applications, ERP, finance, BI/reporting, Microsoft/.NET estates, data volumes fitting a powerful scale-up primary, and organizations valuing integrated tooling, commercial support, and mature DBA practices over transparent horizontal write scaling.

Battle-tested: top-tier. It is appropriate through your highest category when properly engineered, although at truly global write scale it usually requires application sharding, multiple independent databases, Azure-specific database services, or a different distributed database.


8. Oracle Database

Single instance
     or
Oracle RAC instances A/B/C
       │ Cache Fusion
       │
shared database storage
       │
Data Guard standby(s) for remote DR

Optional Oracle Sharding:
independent shared-nothing Oracle databases by shard key

Oracle was founded as Software Development Laboratories in 1977; Relational Software, Inc. introduced the first commercially available SQL implementation in 1979. Oracle Database is proprietary and represents nearly five decades of continuous database engineering. (docs.oracle.com )

A normal Oracle database is a single writable instance. Oracle RAC runs multiple active instances against the same shared database files and uses Cache Fusion to transfer current blocks between instance buffer caches. RAC is shared-everything, not shared-nothing. It provides exceptional node-level availability and can scale many workloads, but hot blocks, global-cache traffic, and shared-storage limits remain. Data Guard maintains physical or logical standby databases for HA/DR, while Active Data Guard permits read offload. (docs.oracle.com )

Oracle Sharding is a separate shared-nothing architecture. Data is horizontally partitioned across independent Oracle databases using consistent-hash, range, list, or composite sharding. It offers major scale and fault isolation, but applications need a clear sharding key, and cross-shard joins, indexes, constraints, and transactions are inherently more expensive or restricted. Oracle Sharding is not “turn on automatic Cockroach-style sharding for any legacy schema.” (docs.oracle.com )

Oracle’s index repertoire is arguably the broadest: B-tree, bitmap, function-based, reverse-key, partitioned local/global, index-organized tables, domain indexes, spatial/text indexes, invisible indexes, and advanced optimizer access paths such as skip scans. B-trees and function-based covering strategies dominate OLTP; bitmap indexes excel in low-concurrency warehouses; local partitioned indexes simplify lifecycle operations; index-organized tables can eliminate a second table lookup for primary-key-heavy schemas. (docs.oracle.com )

Oracle supports Read Committed by default, Read Only transactions, and snapshot-style Serializable. Oracle Serializable can raise ORA-08177 on conflicting updates, but historically is not strict serializability in the same sense as modern PostgreSQL SSI or CockroachDB; explicit locking and schema constraints remain important for certain invariants. Local ACID semantics are extremely mature. Distributed ACID is available through two-phase commit and sharding features, but latency and operational complexity increase rapidly with participants.

Sweet spot: core banking, airline reservations, telecom billing, ERP, government systems, complex stored procedures, enormous schemas, mixed vendor packages, and workloads where mature recovery, diagnostics, optimizer behavior, support, and institutional knowledge matter more than license cost.

Battle-tested: the highest in this list, alongside SQL Server in broad enterprise history. Oracle RAC plus Active Data Guard remains a standout for extreme HA; Oracle Sharding is the scale-out option. Its disadvantages are expense, complexity, specialized staffing, feature/licensing complexity, shared-storage requirements for RAC, and less cloud/vendor portability.


9. CockroachDB

PostgreSQL clients
       │
any symmetric CockroachDB node
SQL + distributed optimizer
       │
transaction coordinator / HLC / Parallel Commits
       │
ordered KV ranges
       │
Raft replicas on Pebble LSM storage

CockroachDB was started as an open-source project in 2014 by former Google engineers Spencer Kimball, Peter Mattis, and Ben Darnell. Cockroach Labs was formed in 2015. It was explicitly designed as an open-source approximation of Google Spanner’s operational model using commodity clocks and infrastructure. (cockroachlabs.com )

Every node can accept SQL and acts as a gateway. SQL is translated into ordered KV operations. Tables and indexes are divided into contiguous ranges, which automatically split and merge. Each range is independently replicated through Raft, normally with three or more voting replicas. One replica holds a lease and coordinates consistent reads and writes. Storage uses the Pebble LSM engine, while Hybrid Logical Clocks provide timestamp ordering. (cockroachlabs.com )

CockroachDB’s transaction layer is one of its most impressive components. It provides distributed ACID transactions across arbitrary rows, ranges, and tables using MVCC, write intents, timestamp pushing/refreshing, transaction pipelining, and Parallel Commits. Serializable is the default and is genuinely serializable, but contention can produce transaction-retry errors that applications must handle. Read Committed is now also available and reduces retry burden at the cost of permitting ordinary RC anomalies. (cockroachlabs.com )

Indexes are themselves distributed KV keyspaces. Important types and patterns include primary and secondary indexes, composite and covering STORING indexes, partial indexes, expression indexes, inverted indexes for JSON/arrays, hash-sharded indexes to mitigate sequential-key hotspots, and regional/partitioning strategies. A secondary-index lookup may contact one range for the index and another for the base row, so covering indexes can remove a network hop as well as a disk lookup.

CockroachDB has high PostgreSQL wire and SQL compatibility, including drivers and many common ORM patterns. It is not PostgreSQL internally: arbitrary C extensions, many superuser behaviors, physical replication tools, some procedural features, planner details, and PostgreSQL-specific operational assumptions are unavailable or different.

Its global topology support is arguably the best in this list. Multi-region SQL can mark tables as regional, regional-by-row, or global, controlling replica and lease placement. Non-voting replicas and follower reads provide local read scale. Strong global writes still obey physics: transactions spanning continents pay WAN latency. The goal is to keep most transactions local while retaining one logical database. CockroachDB explicitly characterizes itself as CP; an affected range stops without quorum. (cockroachlabs.com )

Sweet spot: global SaaS, control planes, account and entitlement databases, inventories, orders, metadata, and financial applications where active use from several regions and strong consistency matter more than minimum single-region latency. It is less ideal for hotspot-heavy counter workloads, enormous bulk scans without an analytical companion, or applications unable to retry serializable transactions.

Battle-tested: strong and improving, but younger than Oracle/SQL Server and with fewer public hyperscale deployments than Alibaba’s internal systems. Its architectural polish for global SQL is nevertheless a standout.


10. Comdb2

Client library chooses nearest healthy node
         │
SQL executes on a replicant
         │
writes/offloaded changes sent to elected master
         │
master synchronously ships Berkeley DB physical logs
         │
all coherent nodes hold a complete database copy

Comdb2 was started at Bloomberg in 2004 to replace an older internal database and simplify keeping databases synchronized. It was open-sourced around its 2016 VLDB publication period. A dedicated Bloomberg team continues development, and the system stores a substantial portion of Bloomberg’s data. (github.com )

A Comdb2 “cluster” is a set of machines receiving the same physical replication stream. One master is elected, while replicants execute user SQL. Read/write transactions may be submitted through any node; modifications are gathered and sent to the master for application and replication. By default, replication is synchronous, and lagging nodes are marked incoherent and removed from request service. This is HA and distributed query admission, but not transparent horizontal data sharding: every cluster node normally has the full database. (bloomberg.github.io )

The engine combines a SQLite-derived SQL layer with Bloomberg’s transaction and networking components and Berkeley DB B-tree storage. Indexes are primarily fixed-size-field B-tree structures, with unique, partial, and covering-related schema strategies. Partial indexes are especially valuable for excluding irrelevant rows, reducing both index size and lookup cost. Compared with PostgreSQL, Oracle, or SQL Server, the index and SQL ecosystems are narrower. (bloomberg.github.io )

Comdb2 offers its default/Read Committed-like level, Snapshot, and Serializable isolation. Serializable is optimistic and may reject a transaction at commit. With HASql, durable LSNs, synchronized clocks, and Serializable enabled, it can provide linearizable consistency for a supported subset of SQL. Asynchronous physical replicants are available for remote DR but are outside the source cluster’s synchronous consistency boundary. (bloomberg.github.io )

Long-distance synchronous clusters incur high latency because acknowledged writes normally propagate throughout the coherent cluster. A practical topology uses a low-latency synchronous cluster in one metro/region and asynchronous physical replication to remote sites. Its CAP behavior is consistency-oriented under durable/linearizable settings, but exact guarantees depend on configuration.

Sweet spot: many independently deployed relational databases, each needing simple any-node access, automatic master election, synchronous replicas, and strong internal reliability. It is compelling inside Bloomberg because the client libraries, service discovery, operational conventions, and expert team already exist. For a new external project, its small ecosystem and limited commercial support are substantial drawbacks.


11. Primary/replica systems versus native distributed SQL

Characteristic Traditional PostgreSQL/MySQL/AliSQL/SQL Server/openGauss/Oracle Data Guard Native distributed SQL
Simple transaction latency Usually lower Higher because of consensus/routing
Operational concepts Familiar and relatively simple More nodes, ranges, quorum, rebalancing, hotspots
Write scale Bounded by one primary unless manually sharded Scales across shard/range leaders
Read scale Excellent with replicas, possibly stale Distributed reads; follower/stale reads often available
Strict local transactions Mature and efficient Available, but cross-shard transactions cost more
Failover Seconds to minutes, topology-dependent Per-shard automatic leader election
Global writes Usually one write region Multiple entry regions, but data locality still determines latency
SQL compatibility Native Usually incomplete emulation of MySQL/PostgreSQL
Index cost Local storage and write amplification Storage/write amplification plus possible cross-node work
Best choice Data fits one primary; team values simplicity Primary is the scale/availability bottleneck or global topology is mandatory

The main mistake is adopting distributed SQL merely because it sounds “enterprise.” If a 2–8 TB database with 10,000 writes/sec fits comfortably on one well-engineered primary, conventional PostgreSQL/MySQL/SQL Server/Oracle is usually simpler and cheaper. Distributed SQL earns its complexity when you need write scaling, automatic sharding, region-aware placement, online growth beyond one server, or consensus-based availability without manual failover architecture.


Enterprise-level placement

Level 1: Homelab or fewer than 1,000 noncritical users

Standouts:

  1. AliSQL if you want MySQL behavior and to explore its new vector/DuckDB work.
  2. openGauss if PostgreSQL-like administration or Huawei/Kunpeng technology interests you.
  3. SQL Server Developer/Express for .NET and Windows development.
  4. Oracle Free/Developer editions for learning Oracle.
  5. CockroachDB single-node for learning distributed SQL semantics, not because one node provides distributed HA.

YDB, TiDB, OceanBase, and PolarDB-X can run small or standalone configurations, but they are generally more machinery than this category needs. OceanBase V4’s integrated standalone mode is the most explicitly designed to bridge small and distributed deployments.

Level 2: Medium regional business requiring availability

Standouts:

  1. SQL Server Always On — excellent tooling, support, readable replicas, and conventional operations.
  2. Oracle plus Data Guard — strongest high-end option if cost and staffing are acceptable.
  3. openGauss primary/standby — strong in its primary ecosystem.
  4. AliSQL/MySQL with managed HA — simple and cost-effective.
  5. CockroachDB or OceanBase if zero-downtime scaling and multi-AZ consensus are already requirements.

At this level, conventional primary-plus-replica architecture is often preferable. Most businesses do not yet need automatic sharding.

Level 3: Large national business requiring sharding, consensus, or distributed writes

Standouts by workload:

  • MySQL migration and HTAP: TiDB
  • Financial distributed OLTP: OceanBase
  • MySQL plus strong global secondary indexes: PolarDB-X
  • Serializable, key-oriented extreme scale: YDB
  • PostgreSQL-facing distributed SQL: CockroachDB
  • Complex enterprise applications that still fit scale-up: Oracle RAC or SQL Server AG
  • Chinese domestic enterprise stack: openGauss/GaussDB ecosystem

This is the category where native distributed SQL begins to justify itself. TiDB is usually easiest for a sharded-MySQL estate; CockroachDB has stronger serializability and geographic abstractions; OceanBase has deeper financial/Alibaba production history; YDB offers exceptional scale but a smaller ecosystem.

Level 4: International bank, airline, government, intelligence, critical bookkeeping

There are two different winning groups.

Institutional/mature enterprise winners

  1. Oracle Database: RAC + Active Data Guard, optionally Sharding
  2. SQL Server Always On/Distributed AG
  3. OceanBase Enterprise, particularly in APAC financial environments
  4. TiDB Enterprise, where MySQL compatibility and scale outweigh the lack of strict serializability

Oracle is the strongest general answer when the priorities are decades of operational evidence, complex transactions, vendor accountability, recovery tooling, security certification, and global availability of expert staff. SQL Server is similarly strong in Microsoft-centered organizations.

Native global-distribution winners

  1. CockroachDB — best general-purpose global SQL/locality model in this list
  2. OceanBase — strongest long-running financial distributed-OLTP pedigree
  3. YDB — exceptional scale and serializable semantics for engineering-heavy organizations
  4. TiDB — excellent scale and operational maturity, but default snapshot isolation is a meaningful distinction for bookkeeping invariants
  5. PolarDB-X — compelling within Alibaba Cloud and MySQL-centric environments

For military, intelligence, or sovereign-government use, technical architecture is only one factor. Vendor jurisdiction, supply-chain assurance, source-code review, export controls, support location, personnel clearances, cryptographic certifications, and procurement rules may disqualify a technically capable product. In a United States classified or highly regulated environment, Oracle, SQL Server, and approved domestic/open-source stacks are generally more realistic procurement choices than Yandex- or China-origin products, regardless of their technical merit.


Final recommendations by sweet spot

Requirement Best candidates
Simplest reliable regional OLTP SQL Server, Oracle, AliSQL/MySQL, openGauss
Replace a large sharded-MySQL fleet TiDB, PolarDB-X, OceanBase MySQL mode
Strict serializable distributed SQL CockroachDB, YDB
Financial-grade distributed OLTP with decade-plus hyperscale history OceanBase
Best global/multi-region application-facing model CockroachDB
Extreme Yandex-style key-oriented throughput YDB
Strong distributed global secondary indexes in MySQL ecosystem PolarDB-X
HTAP with mature row and column paths TiDB, OceanBase, PolarDB-X
Conventional enterprise DB with richest SQL/features Oracle, SQL Server
Chinese sovereign technology stack openGauss/GaussDB, OceanBase, PolarDB-X
Bloomberg-style replicated database clusters Comdb2
MySQL plus embedded vector/analytical experimentation AliSQL
Lowest operational risk when one primary is sufficient Oracle, SQL Server, ordinary MySQL/AliSQL or PostgreSQL/openGauss

Overall ranking by concern

  • Most proven traditional enterprise database: Oracle
  • Best Microsoft-enterprise option: SQL Server
  • Best globally distributed SQL architecture: CockroachDB
  • Best distributed financial-production pedigree: OceanBase
  • Best MySQL-compatible scale-out generalist: TiDB
  • Best MySQL distributed-index architecture: PolarDB-X
  • Most technically distinctive extreme-scale platform: YDB
  • Best Chinese PostgreSQL-derived enterprise kernel: openGauss
  • Best conventional Alibaba-derived MySQL branch: AliSQL
  • Most impressive specialized internal system: Comdb2

The practical shortlist for most new evaluations would therefore be:

  • Oracle or SQL Server when one writable database plus HA/DR is enough.
  • TiDB when replacing sharded MySQL.
  • CockroachDB when multi-region serializable SQL is the defining requirement.
  • OceanBase for very large financial OLTP, especially with strong regional vendor support.
  • PolarDB-X for Alibaba Cloud and MySQL/global-index-heavy applications.
  • YDB when extreme scale and serializable key-oriented access justify adopting its distinct ecosystem.