背景:把平台从「单机单进程」改造成「1 组 Manager + N 台 Worker + 共享归属状态」,
硬约束 = 全程兼容单例模式(deployMode 默认 local;生产切换前 47 一行未动)。
主要改动
1) 数据模型 v7(SQLite 与 PG 两方言同步):新增 dsh_hosts 注册表 +
dsh_instances.{host_id,epoch,heartbeat_at,lease_until};claimInstance 原子抢占
(UPDATE … WHERE host_id IS NULL OR lease_until < now)+ pinInstanceHost 钉住归属。
2) 租约与 fencing:src/supervisor/lease.ts(acquire/renew/release + stillHolder 判据 +
ttl > 2×renew 硬校验);心跳里续租,失权即向 worker 下发更高 epoch(self-fencing)。
⚠️ release 只清租约(lease_until),**保留 host_id** —— host_id 是「用户数据在哪台」的锚点。
3) Worker agent(src/worker/agent.ts,子命令 dshs worker):实例生命周期 + 文件面 /fs/*
+ 幂等键(operationId)+ 鉴权(timingSafeEqual);Worker 不写控制面数据
(apiKey/uid 由 Manager 随 launch 投递,R5 收窄)。
4) 远端 Spawner + LeasedSpawner:按 host 路由(**粘性优先**:有历史归属且那台 up 就留在原地,
否则按容量选最空的)+ 容量准入 + deployMode=cluster 装配(systemd drop-in,可回滚)。
5) bwrap 修正:**所有挂载点的中间目录统一前置 + 去重 + 由外到内**(「就近创建」会在嵌套前缀下
遮掉已绑挂载点 ⇒ bwrap: Can't chdir);且**只能用 --tmpfs**,用 --perms 会让 47 的
bwrap 0.4.0 直接拒启动(沙箱全挂)。
6) 跨机隧道 src/worker/tunnel.ts:SSH ControlMaster + 动态 -R 转发;**自愈由 agent 本地
20s 定时器驱动**(不能只放 /healthz —— 心跳本身经隧道进来,断了就没人触发它)。
7) 文件面按归属路由(RemoteUserFs):实例与文件必须落在同一台机器,否则实例看不到自己的文件。
8) 观测面:dshs doctor / dshs cluster status。
验证(本次均已实跑)
- test/lease.test.mjs:SQLite 10/10 == PG 10/10
- 组件级端到端 5 个:verify-cluster-{agent,lease,fs,migrate,live}.mjs
- 真跨机(47 Manager / 106 Worker,跨云 + 反向隧道)verify-cluster-cross.mjs 九步全绿
- 域名形态访问 verify-cluster-domain.mjs(<user>.域名 → Manager → 远端实例;越权 403)
- 冒烟 scripts/smoke-*:6/8,失败项与改动前基线完全相同(无回归)
- 生产切换与回滚剧本见 dsh-server-docs/交接单/T08-集群化落地-兼容单例模式.md §16
46 lines
1.8 KiB
Bash
46 lines
1.8 KiB
Bash
#!/usr/bin/env bash
|
||
# 起一台 **cluster 模式的 Manager**(T08 跨机演练用;在 Manager 那台机器上跑)。
|
||
#
|
||
# 现场的对照(2026-09-15 演练实测):
|
||
# · Manager 在 **47**(本脚本所在机器),监听 `127.0.0.1:13080`(**不公网暴露**)
|
||
# · Worker agent 在 **106**,经 SSH 反向隧道出现在本机 `127.0.0.1:19000` / `19001`
|
||
# · 控制面 PG 也在 **106**,经同一条隧道出现在本机 `127.0.0.1:15432`
|
||
# 用法:bash scripts/start-cluster-manager.sh (env 见下方 manager.env)
|
||
在 47 上:建 manager.env、bootstrap 管理员、起 cluster Manager(127.0.0.1:13080)
|
||
set -uo pipefail
|
||
cd /opt/dshs-cluster || exit 1
|
||
|
||
cat > /opt/dshs-cluster/manager.env <<'ENVEOF'
|
||
DSHS_DEPLOY_MODE=cluster
|
||
DSHS_DB_URL=postgres://dshs:[email protected]:15432/dshs_cross
|
||
DSHS_DATA_ROOT=/opt/dshs-cluster/data
|
||
DSHS_CLUSTER_HOST_ID=m-47
|
||
DSHS_CLUSTER_AGENT_URL=http://127.0.0.1:19000
|
||
DSHS_CLUSTER_AGENT_TOKEN=cross-machine-token
|
||
DSHS_CLUSTER_INSTANCE_HOST=127.0.0.1
|
||
DSHS_CLUSTER_WORKER_DATA_ROOT=/opt/dshs-cluster/live-data
|
||
DSHS_CLUSTER_CAPACITY_MB=-1
|
||
DSHS_CLUSTER_REGISTER_SELF=0
|
||
DSHS_CLUSTER_LEASE_TTL_MS=30000
|
||
ENVEOF
|
||
|
||
set -a
|
||
# shellcheck disable=SC1091
|
||
. /opt/dshs-cluster/manager.env
|
||
set +a
|
||
|
||
echo "--- bootstrap 管理员 ---"
|
||
node lib/cli.js bootstrap-admin --username root --password crossmgr123 2>&1 | tail -1
|
||
|
||
echo "--- 起 Manager ---"
|
||
pkill -f "dshs-cluster/lib/cli.js --port 13080" 2>/dev/null
|
||
sleep 1
|
||
nohup node lib/cli.js --port 13080 --host 127.0.0.1 --log-level warn > /tmp/manager-47.log 2>&1 &
|
||
sleep 7
|
||
|
||
echo "--- 自检 ---"
|
||
echo " login.html : $(curl -s -o /dev/null -w '%{http_code}' -m 6 http://127.0.0.1:13080/login.html)"
|
||
echo " 进程 : $(pgrep -cf 'dshs-cluster/lib/cli.js --port 13080')"
|
||
echo " 日志尾部 :"
|
||
tail -4 /tmp/manager-47.log 2>/dev/null | sed 's/^/ /'
|