PostgreSQL-11.3-主从流复制+手动主备切换

时间 2019-11-15

标签 postgresql 11.3 主从复制手动切换栏目 Postgre SQL 繁體版

原文原文链接

1 摘要

使用PostgreSQL 11.3 建立两个节点：node1 和 node2；配置主从流复制，而后作手动切换（failover）。为了配置过程简单，两个节点在同一台物理机器上。node

首先创建主从同步流复制，起初，node1做为主节点（Primary），node2做为从节点（Standby）。sql

接下来，模拟主节点(node1)没法工做，把从节点(node2)手动切换成主节点。而后把原来的主节点(node1)做为从节点加入系统。此时，node2是主 (Primary), node1是从(Standby)。数据库

最后，模拟主节点(node2)没法工做，把从节点(node1)手动切换成主节点。而后把node2做为从节点加入系统。此时，node1是主 (Primary), node2是从(Standby)。bash

2 建立node1和node2，配置主从流复制

上级目录是:app

/Users/tom

两个子目录，testdb113,testdb113b, 分别属于两个节点, node1, node2:dom

testdb113  --> node1
 testdb113b --> node2

2.1 建立主节点(Primary) node1

mkdir testdb113
initdb -D ./testdb113
pg_ctl -D ./testdb113 start

2.1.1 node1: 建立数据库帐号，用于流复制

建立帐号 'replicauser'，该帐号存在于node1和node2 。

CREATE ROLE replicauser WITH REPLICATION LOGIN ENCRYPTED PASSWORD 'abcdtest';

2.1.2 node1: 建立归档目录

mkdir ./testdb113/archive

2.1.3 node1: 在配置文件中(postgresql.conf)设定参数

#for the primary 
wal_level = replica 
synchronous_commit = remote_apply 
archive_mode = on 
archive_command = 'cp %p /Users/tom/testdb113/archive/%f'
max_wal_senders =  10 
wal_keep_segments = 10 
synchronous_standby_names = 'pgslave001'

#for the standby
hot_standby = on

2.1.4 node1: 编辑文件pg_hba.conf，设置合法地址:

add following to then end of pg_hba.conf

host     replication    replicauser          127.0.0.1/32            md5
host     replication    replicauser          ::1/128                 md5

2.1.5 node1: 预先编辑文件:recovery.done

注意: 该文件后来才会用到。当node1成为从节点时候，须要把recovery.done重命名为recovery.conf

standby_mode = 'on'
recovery_target_timeline = 'latest'
primary_conninfo = 'host=localhost port=6432 user=replicauser password=abcdtest application_name=pgslave001'
restore_command = 'cp /Users/tom/testdb113b/archive/%f %p'
trigger_file = '/tmp/postgresql.trigger.5432'

2.2 建立和配置从节点(node2)

mkdir testdb113b
pg_basebackup  -D ./testdb113b
chmod -R 700 ./testdb113b

2.2.1 node2: 编辑文件postgresql.conf，设定参数

这个文件来自node1(因为使用了pg_basebackup)，只须要更改其中的个别参数。

#for the primary 
port = 6432 
wal_level = replica 
synchronous_commit = remote_apply 
archive_mode = on 
archive_command = 'cp %p /Users/tom/testdb113b/archive/%f'
max_wal_senders =  2 
wal_keep_segments = 10 
synchronous_standby_names = 'pgslave001'


#for the standby
hot_standby = on

2.2.2 node2: 创建归档目录

mkdir ./testdb113b/archive

2.2.3 node2: 检查/编辑文件pg_hba.conf

这个文件来自node1，不须要编辑

2.2.4 node2: 建立/编辑文件recovery.conf

PG启动时，若是文件recovery.conf存在，则工做在recovery模式，看成从节点。若是当前节点升级成主节点，文件recovery.conf将会被重命名为recovery.done。

standby_mode = 'on'
recovery_target_timeline = 'latest'
primary_conninfo = 'host=localhost port=5432 user=replicauser password=abcdtest application_name=pgslave001'
restore_command = 'cp /Users/tom/testdb113/archive/%f %p'
trigger_file = '/tmp/postgresql.trigger.6432'

2.3 简单测试

2.3.1 从新启动主节点(node1):

pg_ctl -D ./testdb113 restart

2.3.2 启动从节点(node2)

pg_ctl -D ./testdb113b start

2.3.3 在node1上作数据更改

node1支持数据的写和读

psql -d postgres

postgres=# create table test ( id int, name varchar(100));
CREATE TABLE
postgres=# insert into test values(1,'1');
INSERT 0 1

2.3.4 节点node2读数据

node2是从节点(Standby)，只能读，不能写。

psql -d postgres -p 6432

postgres=# select * from test;
 id | name 
----+------
  1 | 1
(1 row)

postgres=# insert into test values(1,'1');
ERROR:  cannot execute INSERT in a read-only transaction
postgres=#

3 模拟主节点node1不能工做时，把从节点node2提高到主节点

3.1 中止主节node1

node1中止后，只有node2工做，仅仅支持数据读。

pg_ctl -D ./testdb113 stop

3.1.1 尝试链接到节点node1时失败

由于node1中止了。

Ruis-MacBook-Air:tom$ psql -d postgres
psql: could not connect to server: No such file or directory
	Is the server running locally and accepting
	connections on Unix domain socket "/tmp/.s.PGSQL.5432"?

3.1.2 尝试链接到node2，成功；插入数据到node2，失败

node2在工做，可是只读

Ruis-MacBook-Air:~ tom$ psql -d postgres -p 6432
psql (11.3)
Type "help" for help.

postgres=# insert into test values(1,'1');
ERROR:  cannot execute INSERT in a read-only transaction
postgres=#

3.2 node2: 把node2提高到主

提高后，node2成为了主(Primary)，既能读，也能写。

touch /tmp/postgresql.trigger.6432

3.3 node2: 插入数据到node2，阻塞等待

node2是主节点(Primary)，可以成功插入数据。可是因为设置了同步流复制，node2必须等待某个从节点把数据更改应用到从节点，而后返回相应的LSN到主节点。

postgres=# insert into test values(2,'2');

3.4 node1: 建立文件recovery.conf

使用以前创建的文件recovery.done做为模版，快速建立recovery.conf。

mv recovery.done recovery.conf

3.5 node1: 启动node1，做为从节点(Standby)

启动后，系统中又有两个节点：node2是主，node1是从。

pg_ctl -D ./testdb113 start

3.6 插入数据到node2(见3.3)成功返回

因为节点node1做为从节点加入系统，应用主节点的数据更改，而后返回相应WAL的LSN到主节点，使主节点等到了回应，从而事务处理得以继续。

postgres=# insert into test values(2,'2');
INSERT 0 1
postgres=#

3.7 分别在node1和node2读取数据

因为两个节点都只是数据读，并且流复制正常工做，全部读取到相同的结果

postgres=# select * from test;
 id | name 
----+------
  1 | 1
  2 | 2
(2 rows)

4 模拟node2没法工做时，提高node1成为主(Primary)

4.1 中止node2

pg_ctl -D ./testdb113b stop

4.1.1 尝试访问node2，失败

由于node2已经中止了.

Ruis-MacBook-Air:~ tom$ psql -d postgres -p 6432
psql: could not connect to server: No such file or directory
	Is the server running locally and accepting
	connections on Unix domain socket "/tmp/.s.PGSQL.5432"?

4.1.2 node1: 链接到node1，成功；插入数据到node1，失败

node1只读

Ruis-MacBook-Air:~ tom$ psql -d postgres
psql (11.3)
Type "help" for help.

postgres=# insert into test values(3,'3');
ERROR:  cannot execute INSERT in a read-only transaction
postgres=#

4.2 node1: 提高node1，成为主(Primary)

touch /tmp/postgresql.trigger.5432

4.3 node1: 插入数据到node1，阻塞等待

node1成功插入数据，可是要等待从节点的 remote-apply.

postgres=# insert into test values(3,'3');

4.4 node2: 建立文件recovery.conf

mv recovery.done recovery.conf

4.5 node2: 启动node2，做为从节点(Standby)

pg_ctl -D ./testdb113b start

4.6 node1: 插入数据(见4.3)成功返回。

postgres=# insert into test values(3,'3');

INSERT 0 1

4.7 分别在node1和node2读取数据

因为两个节点都只是数据读，并且流复制正常工做，全部读取到相同的结果

postgres=# select * from test;
 id | name 
----+------
  1 | 1
  2 | 2
  3 | 3
(3 rows)

5 自动化解决方案

当主节点不能工做时，须要把一个从节点提高成为主节点。若是是手工方式提高，可使用trigger文件，也可使用命令'pg_ctl promote' 。socket

目前也有不少种自动化监控和切换的方案，在网上都能搜索到。或者，能够本身动手写一个自动化监控和切换的方案。post