背景:把平台从「单机单进程」改造成「1 组 Manager + N 台 Worker + 共享归属状态」,
硬约束 = 全程兼容单例模式(deployMode 默认 local;生产切换前 47 一行未动)。
主要改动
1) 数据模型 v7(SQLite 与 PG 两方言同步):新增 dsh_hosts 注册表 +
dsh_instances.{host_id,epoch,heartbeat_at,lease_until};claimInstance 原子抢占
(UPDATE … WHERE host_id IS NULL OR lease_until < now)+ pinInstanceHost 钉住归属。
2) 租约与 fencing:src/supervisor/lease.ts(acquire/renew/release + stillHolder 判据 +
ttl > 2×renew 硬校验);心跳里续租,失权即向 worker 下发更高 epoch(self-fencing)。
⚠️ release 只清租约(lease_until),**保留 host_id** —— host_id 是「用户数据在哪台」的锚点。
3) Worker agent(src/worker/agent.ts,子命令 dshs worker):实例生命周期 + 文件面 /fs/*
+ 幂等键(operationId)+ 鉴权(timingSafeEqual);Worker 不写控制面数据
(apiKey/uid 由 Manager 随 launch 投递,R5 收窄)。
4) 远端 Spawner + LeasedSpawner:按 host 路由(**粘性优先**:有历史归属且那台 up 就留在原地,
否则按容量选最空的)+ 容量准入 + deployMode=cluster 装配(systemd drop-in,可回滚)。
5) bwrap 修正:**所有挂载点的中间目录统一前置 + 去重 + 由外到内**(「就近创建」会在嵌套前缀下
遮掉已绑挂载点 ⇒ bwrap: Can't chdir);且**只能用 --tmpfs**,用 --perms 会让 47 的
bwrap 0.4.0 直接拒启动(沙箱全挂)。
6) 跨机隧道 src/worker/tunnel.ts:SSH ControlMaster + 动态 -R 转发;**自愈由 agent 本地
20s 定时器驱动**(不能只放 /healthz —— 心跳本身经隧道进来,断了就没人触发它)。
7) 文件面按归属路由(RemoteUserFs):实例与文件必须落在同一台机器,否则实例看不到自己的文件。
8) 观测面:dshs doctor / dshs cluster status。
验证(本次均已实跑)
- test/lease.test.mjs:SQLite 10/10 == PG 10/10
- 组件级端到端 5 个:verify-cluster-{agent,lease,fs,migrate,live}.mjs
- 真跨机(47 Manager / 106 Worker,跨云 + 反向隧道)verify-cluster-cross.mjs 九步全绿
- 域名形态访问 verify-cluster-domain.mjs(<user>.域名 → Manager → 远端实例;越权 403)
- 冒烟 scripts/smoke-*:6/8,失败项与改动前基线完全相同(无回归)
- 生产切换与回滚剧本见 dsh-server-docs/交接单/T08-集群化落地-兼容单例模式.md §16
77 lines
4.0 KiB
Bash
77 lines
4.0 KiB
Bash
#!/usr/bin/env bash
|
||
# 切换 A 步(在 47 上跑):
|
||
# ① 备份 /opt/dshs/lib → /opt/dsh/backups/lib-<ts>/
|
||
# ② 覆盖 /opt/dshs/lib(T08 集群版代码)
|
||
# ③ 装 **本地 Worker**(w-47,19100,无隧道 —— Manager 同机直连)
|
||
# ④ 把既有用户(admin/guest)的归属**预置**为 w-47(否则粘性落点无处可粘、新老用户会被按容量随机调度)
|
||
# ⑤ 在 PG 里注册 w-47 / w-106 两台 worker
|
||
set -uo pipefail
|
||
TS=$(date +%Y%m%d-%H%M%S)
|
||
W47_TOKEN="dshs-worker-47-c4b7e19f"
|
||
W106_TOKEN="dshs-worker-7f3a91c05e"
|
||
PGURL="postgres://dshs:[email protected]:15432/dshs"
|
||
TARBALL=/tmp/dshs-lib-new.tgz
|
||
|
||
echo "=== ① 备份 /opt/dshs/lib ==="
|
||
mkdir -p "/opt/dsh/backups/lib-$TS"
|
||
cp -a /opt/dshs/lib "/opt/dsh/backups/lib-$TS/lib" && echo " ✓ 备份到 /opt/dsh/backups/lib-$TS/lib($(find /opt/dsh/backups/lib-$TS -type f | wc -l) 文件)"
|
||
|
||
echo "=== ② 覆盖 lib ==="
|
||
[ -f "$TARBALL" ] || { echo " ✗ 缺少 $TARBALL"; exit 1; }
|
||
rm -rf /opt/dshs/lib && tar -xzf "$TARBALL" -C /opt/dshs
|
||
echo " ✓ 已覆盖;cluster 特征检查: $(grep -l "DEPLOY_MODE" /opt/dshs/lib/config.js >/dev/null 2>&1 && echo '有 cluster 代码 ✓' || echo '✗ 未见 cluster 代码')"
|
||
echo " lease/agent/tunnel: $(ls /opt/dshs/lib/supervisor/lease.js /opt/dshs/lib/worker/agent.js /opt/dshs/lib/worker/tunnel.js 2>/dev/null | wc -l)/3"
|
||
|
||
echo "=== ③ 本地 Worker 单元(w-47,无隧道) ==="
|
||
cat > /etc/dshs-worker.env <<ENV
|
||
DSHS_DATA_ROOT=/var/lib/dshs
|
||
DSHS_ISOLATION_MODE=account
|
||
DSHS_DSH_BIN=/usr/local/bin/dsh
|
||
DSHS_BASE_UID=100000
|
||
DSH_INSTANCE_NODE_OPTIONS=--max-old-space-size=160
|
||
DSH_INSTANCE_UNIVER_SOCKET=auto
|
||
DSHS_CLUSTER_AGENT_TOKEN=$W47_TOKEN
|
||
ENV
|
||
chmod 600 /etc/dshs-worker.env
|
||
cat > /etc/systemd/system/dshs-worker.service <<UNIT
|
||
[Unit]
|
||
Description=DSHS cluster worker agent (this host = 47, local users' instances)
|
||
After=network-online.target
|
||
|
||
[Service]
|
||
Type=simple
|
||
EnvironmentFile=/etc/dshs-worker.env
|
||
ExecStart=/usr/local/bin/node /opt/dshs/lib/cli.js worker --port 19100 --host 127.0.0.1 --host-id w-47 --instance-host 127.0.0.1 --log-level info
|
||
Restart=on-failure
|
||
RestartSec=3
|
||
KillMode=mixed
|
||
|
||
[Install]
|
||
WantedBy=multi-user.target
|
||
UNIT
|
||
systemctl daemon-reload; systemctl enable dshs-worker >/dev/null 2>&1
|
||
systemctl restart dshs-worker; sleep 5
|
||
echo " dshs-worker=$(systemctl is-active dshs-worker) healthz=$(curl -s -m 6 http://127.0.0.1:19100/healthz | head -c 120)"
|
||
|
||
echo "=== ④ 既有用户归属预置为 w-47(粘性锚点) ==="
|
||
PGPASSWORD=dshs_cluster_2026 /usr/bin/psql -h 127.0.0.1 -p 15432 -U dshs -d dshs -tAc \
|
||
"insert into dsh_instances (id, user_id, role, status, host_id, epoch, heartbeat_at, lease_until)
|
||
select 'dsh-'||id, id, 'main', 'stopped', 'w-47', 0, 0, 0 from users
|
||
on conflict (id) do update set host_id='w-47', lease_until=0" 2>&1 | tail -1
|
||
PGPASSWORD=dshs_cluster_2026 /usr/bin/psql -h 127.0.0.1 -p 15432 -U dshs -d dshs -tAc \
|
||
"select u.username||' -> '||coalesce(i.host_id,'NULL') from users u left join dsh_instances i on i.user_id=u.id" 2>&1 | sed 's/^/ /'
|
||
|
||
echo "=== ⑤ 注册两台 worker ==="
|
||
PGPASSWORD=dshs_cluster_2026 /usr/bin/psql -h 127.0.0.1 -p 15432 -U dshs -d dshs -tAc \
|
||
"insert into dsh_hosts (id, endpoint, agent_token, capacity_mb, used_mb, status)
|
||
values ('w-47','http://127.0.0.1:19100','$W47_TOKEN',1024,0,'up'),
|
||
('w-106','http://127.0.0.1:19000','$W106_TOKEN',2560,0,'up')
|
||
on conflict (id) do update set endpoint=excluded.endpoint, agent_token=excluded.agent_token, capacity_mb=excluded.capacity_mb, status='up'" 2>&1 | tail -1
|
||
PGPASSWORD=dshs_cluster_2026 /usr/bin/psql -h 127.0.0.1 -p 15432 -U dshs -d dshs -tAc \
|
||
"select id||' cap='||capacity_mb||' status='||status||' ep='||endpoint from dsh_hosts order by id" 2>&1 | sed 's/^/ /'
|
||
|
||
echo "=== 回滚剧本(现在就记下) ==="
|
||
echo " rm -f /etc/systemd/system/dshs.service.d/cluster.conf && systemctl daemon-reload && \\"
|
||
echo " systemctl restart dshs # 回 SQLite 单机;lib 回滚 = cp -a /opt/dsh/backups/lib-$TS/lib /opt/dshs/lib"
|
||
echo " 备份时间戳: $TS"
|