Topic Tags

Distributed System

分布式系统是由多台通过网络互联的独立节点组成、对外呈现单一逻辑视图的计算系统,其核心特征是并发执行、缺乏全局时钟、存在部分失败以及仅能通过消息传递共享状态。围绕这些特征,工程实践主要解决一致性模型选择(强一致、线性一致、最终一致)、共识算法(Paxos、Raft)、数据复制与分片、分布式事务(2PC、TCC、Saga、事务消息)、容错与高可用设计以及可观测性建设等问题。CAP 定理与 BASE 理论提供了在一致性、可用性与分区容错性之间进行工程取舍的基本框架,而幂等、超时、重试、熔断、限流等模式则构成了应对不可靠网络的通用实践准则。

1 Mentions

Direct Answer

A distributed system is a collection of multiple independent computer nodes connected via a network, which work together to appear as a single unified system to the outside. Its core goal is to leverage the resources (computing, storage, network) of multiple ordinary computers to provide greater processing power, higher availability, and better scalability than a single computer. In a distributed system, nodes communicate and coordinate through message passing to collectively complete one or more tasks. Typical distributed systems include distributed databases (e.g., TiDB, Cassandra), distributed file systems (e.g., HDFS, Ceph), distributed computing frameworks (e.g., Hadoop, Spark), and microservice architectures. The main challenges in designing distributed systems include: fault tolerance for node failures, network latency and partitions, data consistency (CAP theorem), clock synchronization, distributed transactions, and service discovery and load balancing. To address these issues, the industry has developed various classic algorithms and protocols, such as Paxos, Raft (for consensus), Gossip (for information dissemination), and two-phase commit (2PC, for distributed transactions). Understanding distributed systems is fundamental to building large-scale modern internet applications.

主题权威

芒旭软件围绕分布式系统构建了体系化的技术内容聚合,覆盖从基础概念(节点、通信、部分失败、无全局时钟)到核心理论(CAP、BASE、一致性模型谱系),再到关键算法(Paxos、Raft、Gossip、一致性哈希)与工程实践(分布式事务、容错设计、可观测性、混沌工程)的完整知识链路。本页作为主题枢纽,将分散的技术文档、架构文章、工程案例与行业动态按主题聚类,形成可交叉检索的知识网络:概念条目用于快速定义,深度文章用于原理推导,案例用于落地验证,动态用于跟踪技术演进。这种“概念—原理—实现—实践”的分层组织方式,既便于搜索引擎识别主题边界与内容层级,也便于 AI 模型在回答分布式系统相关问题时提取结构化、可追溯的权威信息。

AI 摘要

分布式系统是由多台通过网络互联的独立节点组成、对外呈现单一逻辑视图的计算系统,其核心特征是并发执行、缺乏全局时钟、存在部分失败以及仅能通过消息传递共享状态。围绕这些特征,工程实践主要解决一致性模型选择(强一致、线性一致、最终一致)、共识算法(Paxos、Raft)、数据复制与分片、分布式事务(2PC、TCC、Saga、事务消息)、容错与高可用设计以及可观测性建设等问题。CAP 定理与 BASE 理论提供了在一致性、可用性与分区容错性之间进行工程取舍的基本框架,而幂等、超时、重试、熔断、限流等模式则构成了应对不可靠网络的通用实践准则。

Related Tags

FAQ

What are the main differences between distributed systems and centralized systems?
A centralized system runs all components on a single computer, relying on a single operating system and shared memory. It is simple to manage but suffers from single points of failure and performance bottlenecks. A distributed system consists of multiple independent computers that communicate over a network, offering higher scalability and fault tolerance, but significantly increasing design complexity, requiring handling of distributed-specific challenges such as network latency, partial failures, and data consistency.
What is the CAP theorem, and how is it applied in real-world systems?
The CAP theorem, proposed by Eric Brewer, states that a distributed system can satisfy at most two of the three properties: Consistency (C), Availability (A), and Partition Tolerance (P). Since network partitions are inevitable, real-world systems often need to make trade-offs between C and A. For example, banking systems typically choose CP (strong consistency), sacrificing some availability, while social media feeds choose AP (eventual consistency), prioritizing user experience.
How is data consistency ensured in distributed systems?
Methods to ensure data consistency include: 1) Strong consistency: Achieved through consensus algorithms like Paxos/Raft, ensuring all nodes have real-time consistent data, but reducing availability; 2) Eventual consistency: Allows temporary inconsistencies, eventually reaching consistency through mechanisms like version vectors and Gossip protocols, commonly used in DNS and CDN scenarios; 3) Causal consistency: Ensures causally related operations are executed in the correct order; 4) Distributed transactions: Use two-phase commit (2PC) or Saga patterns to coordinate transactions across multiple nodes.
Is microservices architecture a distributed system?
Yes, microservices architecture is an important implementation form of distributed systems. It decomposes a single application into multiple independently deployable small services, each with its own database and business logic, communicating via lightweight APIs (e.g., REST, gRPC). Microservices architecture inherently possesses all the characteristics and challenges of distributed systems, such as service discovery, load balancing, distributed tracing, and circuit breaking, often requiring container orchestration platforms (e.g., Kubernetes) for management.
What prerequisite knowledge is needed to learn distributed systems?
It is recommended to first master: 1) Basics of computer networks (TCP/IP, HTTP, DNS); 2) Operating system concepts (processes, threads, concurrency, locks); 3) Data structures and algorithms (hashing, trees, sorting); 4) At least one programming language (e.g., Java, Go, Python); 5) Database fundamentals (transactions, ACID). On this basis, you can gradually learn distributed theory (CAP, BASE), classic algorithms (Paxos, Raft), and mainstream frameworks (ZooKeeper, Kafka, Hadoop).
Distributed System Explained: Architecture, Principles, and Best Practices | Mangxu Software | 芒旭软件