feat(cluster): 集群化落地 —— Manager/Worker 拆分 + 归属租约 + 跨机验证(T08)
背景:把平台从「单机单进程」改造成「1 组 Manager + N 台 Worker + 共享归属状态」,
硬约束 = 全程兼容单例模式(deployMode 默认 local;生产切换前 47 一行未动)。
主要改动
1) 数据模型 v7(SQLite 与 PG 两方言同步):新增 dsh_hosts 注册表 +
dsh_instances.{host_id,epoch,heartbeat_at,lease_until};claimInstance 原子抢占
(UPDATE … WHERE host_id IS NULL OR lease_until < now)+ pinInstanceHost 钉住归属。
2) 租约与 fencing:src/supervisor/lease.ts(acquire/renew/release + stillHolder 判据 +
ttl > 2×renew 硬校验);心跳里续租,失权即向 worker 下发更高 epoch(self-fencing)。
⚠️ release 只清租约(lease_until),**保留 host_id** —— host_id 是「用户数据在哪台」的锚点。
3) Worker agent(src/worker/agent.ts,子命令 dshs worker):实例生命周期 + 文件面 /fs/*
+ 幂等键(operationId)+ 鉴权(timingSafeEqual);Worker 不写控制面数据
(apiKey/uid 由 Manager 随 launch 投递,R5 收窄)。
4) 远端 Spawner + LeasedSpawner:按 host 路由(**粘性优先**:有历史归属且那台 up 就留在原地,
否则按容量选最空的)+ 容量准入 + deployMode=cluster 装配(systemd drop-in,可回滚)。
5) bwrap 修正:**所有挂载点的中间目录统一前置 + 去重 + 由外到内**(「就近创建」会在嵌套前缀下
遮掉已绑挂载点 ⇒ bwrap: Can't chdir);且**只能用 --tmpfs**,用 --perms 会让 47 的
bwrap 0.4.0 直接拒启动(沙箱全挂)。
6) 跨机隧道 src/worker/tunnel.ts:SSH ControlMaster + 动态 -R 转发;**自愈由 agent 本地
20s 定时器驱动**(不能只放 /healthz —— 心跳本身经隧道进来,断了就没人触发它)。
7) 文件面按归属路由(RemoteUserFs):实例与文件必须落在同一台机器,否则实例看不到自己的文件。
8) 观测面:dshs doctor / dshs cluster status。
验证(本次均已实跑)
- test/lease.test.mjs:SQLite 10/10 == PG 10/10
- 组件级端到端 5 个:verify-cluster-{agent,lease,fs,migrate,live}.mjs
- 真跨机(47 Manager / 106 Worker,跨云 + 反向隧道)verify-cluster-cross.mjs 九步全绿
- 域名形态访问 verify-cluster-domain.mjs(<user>.域名 → Manager → 远端实例;越权 403)
- 冒烟 scripts/smoke-*:6/8,失败项与改动前基线完全相同(无回归)
- 生产切换与回滚剧本见 dsh-server-docs/交接单/T08-集群化落地-兼容单例模式.md §16
This commit is contained in:
1 parent
68c0a320ed
commit
c70d5d860e
47 files changed
+5367
-14
No files matched your search
+144
-2
@@ -13,10 +13,13 @@ import { chown, mkdir, readFile, stat, writeFile } from 'node:fs/promises'
|
||||
import type { ServerConfig } from '../config.js'
|
||||
import { createDbAdapter, type CredentialLandingRow, type DbAdapter, type PublicUser } from '../db/index.js'
|
||||
import { createUserFs } from '../fs/provider.js'
|
||||
import { RemoteUserFs } from '../fs/remote-user-fs.js'
|
||||
import type { UserFs } from '../fs/user-fs.js'
|
||||
import { decrypt, deriveKey } from '../crypto.js'
|
||||
import { hashUid } from '../isolation.js'
|
||||
import { LocalSpawner } from '../supervisor/orchestrator.js'
|
||||
import { LeasedSpawner } from '../supervisor/leased-spawner.js'
|
||||
import { RemoteSpawner, type ClusterHost } from '../supervisor/remote-spawner.js'
|
||||
import { registerDshProxy } from '../supervisor/proxy.js'
|
||||
import type { Spawner } from '../supervisor/spawner.js'
|
||||
import {
|
||||
@@ -259,8 +262,147 @@ export async function buildServer(config: ServerConfig): Promise<FastifyInstance
|
||||
if (config.deployMode === 'k8s') {
|
||||
throw new Error('deployMode "k8s" is not supported by this build: only the single-machine backend ships')
|
||||
}
|
||||
const supervisor: Spawner = new LocalSpawner(config, resolveApiKey, resolveUid)
|
||||
const userFs = createUserFs(config)
|
||||
// cluster 模式(T08 S3/S4):实例在 worker 上,Manager 只投递操作 + 代理。
|
||||
// fail-loud:没配 agent 地址就直接报错,别等第一个用户点进来才发现。
|
||||
if (config.deployMode === 'cluster' && config.clusterAgentUrl === '') {
|
||||
throw new Error('deployMode=cluster requires DSHS_CLUSTER_AGENT_URL (e.g. http://127.0.0.1:9000)')
|
||||
}
|
||||
// cluster:RemoteSpawner(传输)+ LeasedSpawner(**归属租约**)——
|
||||
// 后者保证"能不能拉起先问归属",这是多机下防双写同一个 home 的承重件(设计 §3.2)。
|
||||
let leased: LeasedSpawner | undefined
|
||||
// ── 多 worker 的 host 目录(T08 S6)────────────────────────────────────
|
||||
// 由 `dsh_hosts` 派生并**随用随刷新**(TTL 30 s)⇒ **新增 worker 不必重启 Manager**。
|
||||
// 同时供三处使用:RemoteSpawner 的按 host 路由、LeasedSpawner 的 fence 目标、
|
||||
// 以及 `selectHost` 的容量准入 —— 都读**同一份**内存目录,避免三套各自漂移。
|
||||
const hostDirectory = new Map<string, ClusterHost>()
|
||||
hostDirectory.set(config.clusterHostId, {
|
||||
hostId: config.clusterHostId,
|
||||
agentUrl: config.clusterAgentUrl,
|
||||
token: config.clusterAgentToken,
|
||||
instanceHost: config.clusterInstanceHost,
|
||||
})
|
||||
const hostsProvider = async (): Promise<ClusterHost[]> => {
|
||||
for (const row of await db.listDshHosts()) {
|
||||
hostDirectory.set(row.id, {
|
||||
hostId: row.id,
|
||||
agentUrl: row.endpoint,
|
||||
token: row.agentToken,
|
||||
instanceHost: config.clusterInstanceHost,
|
||||
})
|
||||
}
|
||||
return [...hostDirectory.values()]
|
||||
}
|
||||
/**
|
||||
* 按用户归属解析 host(实例面与**文件面**共用这一份,避免两套路由漂移)。
|
||||
*
|
||||
* 为什么两处都要用:用户工作区在**那台 worker 的本地盘**;若文件面固定打一台 agent,
|
||||
* 就会出现「实例跑在 A、mkdir/上传写到 B」⇒ 实例看不到自己的文件、甚至 cwd 不存在而崩
|
||||
* (2026-09-15 生产切换暴露)。
|
||||
*/
|
||||
const hostIdForUser = async (userId: string): Promise<string | undefined> =>
|
||||
(await db.findUserInstance(userId, 'main'))?.hostId ?? undefined
|
||||
|
||||
/**
|
||||
* **文件面专用**路由:没有归属就**先选机并钉住**。
|
||||
*
|
||||
* 为什么不能直接用 hostIdForUser:新用户还没有归属,"写文件"和"launch"会各自选一次机,
|
||||
* 两次可能选到不同机器 ⇒「文件写到 A、实例起在 B」⇒ 实例看不到自己的文件(2026-09-15 实测)。
|
||||
* 首次触达工作区就把归属钉住,后续(含 launch)全走粘性 ⇒ 两面必然一致。
|
||||
*/
|
||||
const hostIdForFile = async (userId: string): Promise<string | undefined> => {
|
||||
const owned = await hostIdForUser(userId)
|
||||
if (owned !== undefined && owned !== null) return owned
|
||||
const chosen = (await selectHost(userId)) ?? config.clusterHostId
|
||||
if (chosen === '') return undefined
|
||||
await db.pinInstanceHost(userId, chosen)
|
||||
return chosen
|
||||
}
|
||||
/**
|
||||
* 选机:**① 粘性优先 ② 再按容量准入**。
|
||||
*
|
||||
* ⚠️ 顺序不能颠倒(2026-09-15 生产切换时补的缺口):用户工作区在**本地盘**、跟着机器走,
|
||||
* 把"已有历史数据的用户"调度到另一台 ⇒ 他打开实例看到**空工作区**。
|
||||
* ⇒ 有历史归属且那台还 `up` 就留在原地;只有**从未有过归属**(新用户)才按容量挑最空的。
|
||||
* `capacityMb <= 0` = 未声明(不设限);`-1` = 显式禁用承载。
|
||||
*/
|
||||
const reserveMb = Number(process.env.DSHS_CLUSTER_RESERVE_MB ?? '512')
|
||||
const selectHost = async (userId?: string): Promise<string | undefined> => {
|
||||
const rows = await db.listDshHosts()
|
||||
const eligible = rows.filter((h) => h.status === 'up' && h.capacityMb !== -1)
|
||||
if (userId !== undefined) {
|
||||
const owned = (await db.findUserInstance(userId, 'main'))?.hostId ?? null
|
||||
if (owned !== null && eligible.some((h) => h.id === owned)) return owned
|
||||
}
|
||||
const candidates = eligible.filter(
|
||||
(h) => h.capacityMb <= 0 || h.usedMb + reserveMb <= h.capacityMb,
|
||||
)
|
||||
if (candidates.length === 0) return undefined // 无候选 ⇒ 回退到配置里那台
|
||||
candidates.sort((a, b) => a.usedMb - b.usedMb)
|
||||
return candidates[0].id
|
||||
}
|
||||
const supervisor: Spawner =
|
||||
config.deployMode === 'cluster'
|
||||
? (leased = new LeasedSpawner(
|
||||
new RemoteSpawner({
|
||||
agentUrl: config.clusterAgentUrl,
|
||||
token: config.clusterAgentToken,
|
||||
instanceHost: config.clusterInstanceHost,
|
||||
defaultHostId: config.clusterHostId,
|
||||
hostsProvider,
|
||||
resolveApiKey,
|
||||
resolveUid,
|
||||
// 按 host 路由:每次操作都落到"该用户实例所在那台"(与文件面同一份)
|
||||
hostIdFor: hostIdForUser,
|
||||
}),
|
||||
db,
|
||||
{
|
||||
hostId: config.clusterHostId,
|
||||
agentUrl: config.clusterAgentUrl,
|
||||
agentToken: config.clusterAgentToken,
|
||||
capacityMb: Number(process.env.DSHS_CLUSTER_CAPACITY_MB ?? '0'),
|
||||
// 专用 Manager 部署设 DSHS_CLUSTER_REGISTER_SELF=0(见 LeasedSpawner 的注释)
|
||||
registerSelf: (process.env.DSHS_CLUSTER_REGISTER_SELF ?? '1') !== '0',
|
||||
ttlMs: Number(process.env.DSHS_CLUSTER_LEASE_TTL_MS ?? '30000'),
|
||||
renewMs: Number(process.env.DSHS_CLUSTER_LEASE_RENEW_MS ?? '10000'),
|
||||
selectHost,
|
||||
agentFor: (hostId: string) => {
|
||||
const h = hostDirectory.get(hostId)
|
||||
return h === undefined ? undefined : { agentUrl: h.agentUrl, token: h.token }
|
||||
},
|
||||
},
|
||||
))
|
||||
: new LocalSpawner(config, resolveApiKey, resolveUid)
|
||||
// 注册本机 + 起心跳(异步,不阻塞启动;心跳失败只影响该 worker 的状态位)
|
||||
if (leased !== undefined) {
|
||||
void leased.start().catch((err: unknown) => {
|
||||
console.error('[cluster] heartbeat/register failed to start:', err)
|
||||
})
|
||||
}
|
||||
const userFs = createUserFs(config, {
|
||||
hostIdFor: hostIdForFile,
|
||||
agentFor: (hostId: string) => {
|
||||
const h = hostDirectory.get(hostId)
|
||||
return h === undefined ? undefined : { agentUrl: h.agentUrl, token: h.token }
|
||||
},
|
||||
})
|
||||
// T08 S5:cluster 模式下**所有 worker 的 dataRoot 必须是同一绝对路径**(基线约定,
|
||||
// 设计 §14.3)。不一致会让 `resolvePath` 算出的"实例眼里的路径"与实际不符 ⇒
|
||||
// 文件面与 launch 的 folder 都会错。这里在启动时**报出来**,别等用户点进去才发现。
|
||||
if (userFs instanceof RemoteUserFs) {
|
||||
void userFs
|
||||
.probeWorkerRoot()
|
||||
.then((root) => {
|
||||
if (root !== undefined && root !== userFs.workerDataRoot) {
|
||||
console.error(
|
||||
`[cluster] worker dataRoot 与配置不一致:agent 报 ${root},本进程按 ${userFs.workerDataRoot} 计算路径。` +
|
||||
'请把 DSHS_CLUSTER_WORKER_DATA_ROOT 设为 worker 上的实际值(所有 worker 必须同路径)。',
|
||||
)
|
||||
}
|
||||
})
|
||||
.catch(() => {
|
||||
/* 探测失败不阻塞启动:会有心跳/调用失败暴露 */
|
||||
})
|
||||
}
|
||||
|
||||
const app = Fastify({
|
||||
logger: { level: config.logLevel },
|
||||
|
||||
Reference in new issue
Block a user