feat(cluster): 集群化落地 —— Manager/Worker 拆分 + 归属租约 + 跨机验证(T08)

背景:把平台从「单机单进程」改造成「1 组 Manager + N 台 Worker + 共享归属状态」,
硬约束 = 全程兼容单例模式(deployMode 默认 local;生产切换前 47 一行未动)。

主要改动
1) 数据模型 v7(SQLite 与 PG 两方言同步):新增 dsh_hosts 注册表 +
   dsh_instances.{host_id,epoch,heartbeat_at,lease_until};claimInstance 原子抢占
   (UPDATE … WHERE host_id IS NULL OR lease_until < now)+ pinInstanceHost 钉住归属。
2) 租约与 fencing:src/supervisor/lease.ts(acquire/renew/release + stillHolder 判据 +
   ttl > 2×renew 硬校验);心跳里续租,失权即向 worker 下发更高 epoch(self-fencing)。
   ⚠️ release 只清租约(lease_until),**保留 host_id** —— host_id 是「用户数据在哪台」的锚点。
3) Worker agent(src/worker/agent.ts,子命令 dshs worker):实例生命周期 + 文件面 /fs/*
   + 幂等键(operationId)+ 鉴权(timingSafeEqual);Worker 不写控制面数据
   (apiKey/uid 由 Manager 随 launch 投递,R5 收窄)。
4) 远端 Spawner + LeasedSpawner:按 host 路由(**粘性优先**:有历史归属且那台 up 就留在原地,
   否则按容量选最空的)+ 容量准入 + deployMode=cluster 装配(systemd drop-in,可回滚)。
5) bwrap 修正:**所有挂载点的中间目录统一前置 + 去重 + 由外到内**(「就近创建」会在嵌套前缀下
   遮掉已绑挂载点 ⇒ bwrap: Can't chdir);且**只能用 --tmpfs**,用 --perms 会让 47 的
   bwrap 0.4.0 直接拒启动(沙箱全挂)。
6) 跨机隧道 src/worker/tunnel.ts:SSH ControlMaster + 动态 -R 转发;**自愈由 agent 本地
   20s 定时器驱动**(不能只放 /healthz —— 心跳本身经隧道进来,断了就没人触发它)。
7) 文件面按归属路由(RemoteUserFs):实例与文件必须落在同一台机器,否则实例看不到自己的文件。
8) 观测面:dshs doctor / dshs cluster status。

验证(本次均已实跑)
- test/lease.test.mjs:SQLite 10/10 == PG 10/10
- 组件级端到端 5 个:verify-cluster-{agent,lease,fs,migrate,live}.mjs
- 真跨机(47 Manager / 106 Worker,跨云 + 反向隧道)verify-cluster-cross.mjs 九步全绿
- 域名形态访问 verify-cluster-domain.mjs(<user>.域名 → Manager → 远端实例;越权 403)
- 冒烟 scripts/smoke-*:6/8,失败项与改动前基线完全相同(无回归)
- 生产切换与回滚剧本见 dsh-server-docs/交接单/T08-集群化落地-兼容单例模式.md §16
This commit is contained in:
admin committed 2026-09-15 18:47:02 +08:00
1 parent 68c0a320ed
commit c70d5d860e
47 files changed
+5367 -14

No files matched your search

+169
View File
@@ -0,0 +1,169 @@
/**
* 实例归属租约(T08 S2;设计 §3.1–§3.2)。
*
* **为什么必须有它**:local 模式靠"进程内 Map + 单机"天然保证「一个用户同时只有一个活实例」;
* 多机后这个保证只能落到 **DB 的原子 CAS** 上 —— 否则两个 worker 会同时写同一个 `$DSH_HOME`
* (会话日志 append 冲突 ⇒ **数据损坏**,本库最贵的一类事故)。
*
* 三条不变量(都不许省):
* 1. **单写者**:抢占必须原子(`claimInstance` 的 `UPDATE … WHERE 无人持有 OR 租约过期`),
* 调用方以"是否真正更新到行"判定成败,**不许"先读后写"**。
* 2. **TTL > 2 × 续租间隔**:留足抖动余量;否则自身网络一抖就会误判自己失权
* (或更糟 —— 误判别人已死)。构造时**硬校验**,fail-loud。
* 3. **fencing**:每次操作带 `epoch`;不匹配 = 已被他人抢占 ⇒ 调用方必须**自杀**
* (self-fencing,如停掉自己那个实例),而不是继续写。
*
* ⚠️ 与 **R9** 的关系:本模块只提供"判定与递增 epoch"的机械能力,
* **不提供**"判定对方已死 → 接管"的自动化。没有心跳判据时单方面接管是被明令禁止的;
* 因此 `expiredAll()` 只用于**巡检/报告/人工确认后的动作**,绝不自动接管。
*
* @module dshs/supervisor/lease
*/
import type { DbAdapter } from '../db/adapter.js'
import type { ClaimResult, DshInstance } from '../db/types.js'
/** 默认租约存活 30 s(设计 §3.2 的建议时序)。 */
export const DEFAULT_LEASE_TTL_MS = 30_000
/** 默认续租间隔 10 s(不变量:ttl > 2 × renew)。 */
export const DEFAULT_LEASE_RENEW_MS = 10_000
export interface LeaseOptions {
/** 租约存活时长(ms)。不变量:必须 **> 2 × renewMs**。 */
ttlMs?: number
/** 续租间隔(ms)。 */
renewMs?: number
/** 注入时钟 —— 仅用于本类自己的判定(**SQL 里的 now 仍是 `Date.now()`**)。 */
now?: () => number
}
/**
* 某实例当前是否由「我」合法持有(fencing 判据)。
*
* 用在两处:① 收到心跳/指令前自检"我还是不是持有者"② 老 worker 复活后判断
* 自己**是否已被接管** ⇒ 是则自杀(防双写)。
*/
export function stillHolder(
instance: DshInstance | undefined,
hostId: string,
epoch: number,
): boolean {
return instance !== undefined && instance.hostId === hostId && instance.epoch === epoch
}
/** 一个用户实例的租约句柄(holding = 我持有 + 我的 epoch)。 */
export class InstanceLease {
readonly hostId: string
private readonly db: DbAdapter
private readonly ttl: number
private readonly renewInterval: number
private readonly now: () => number
/**
* userId → **我认领到的那把租约**(epoch + 落在哪个 host)。
*
* 为什么必须记住 hostId:多 worker 后 `renewInstanceLease(userId, hostId, epoch)` 要能在
* **正确的那个 host** 上校验;只记 epoch 会在多机下续错对象(2026-09-15 T08 S6)。
*/
private readonly held = new Map<string, { epoch: number; hostId: string }>()
constructor(db: DbAdapter, hostId: string, options: LeaseOptions = {}) {
this.db = db
this.hostId = hostId
this.ttl = options.ttlMs ?? DEFAULT_LEASE_TTL_MS
this.renewInterval = options.renewMs ?? DEFAULT_LEASE_RENEW_MS
this.now = options.now ?? Date.now
// 不变量:TTL 必须显著大于续租间隔,否则单次网络抖动就会造成"自己失权"或"误判他人已死"。
if (this.ttl <= 2 * this.renewInterval) {
throw new Error(
`lease: ttlMs(${this.ttl}) must be > 2 × renewMs(${this.renewInterval}) —— ` +
'否则时钟/网络抖动会破坏单写者保证(设计 §3.2 不变量 2)',
)
}
}
get ttlMs(): number {
return this.ttl
}
get renewMs(): number {
return this.renewInterval
}
/** 我当前认领的实例(userId → {epoch, hostId})。 */
holdings(): ReadonlyMap<string, { epoch: number; hostId: string }> {
return this.held
}
/**
* 抢占某用户 main 实例的归属。
*
* `ok:false` = 有人在管 ⇒ **退让**:不要接管、不要重试到死,交给上层决定
* (拉长等待 / 报告管理员)。成功时记住 epoch 供续租与 fencing 使用。
*/
async acquire(
userId: string,
hostId?: string,
meta?: { folder?: string; patch?: string },
): Promise<ClaimResult> {
const target = hostId ?? this.hostId
// meta(folder/patch)随认领一起落库 ⇒ 迁移才能复现启动参数
const res = await this.db.claimInstance(userId, target, this.ttl, meta)
if (res.ok) this.held.set(userId, { epoch: res.epoch, hostId: target })
return res
}
/**
* 续租。返回 false = **我已失权**(被他人以更高 epoch 抢占,或行被删)⇒ 调用方
* 必须 self-fence(停掉自己那个实例),并清掉本地记录。
*/
async renew(userId: string): Promise<boolean> {
const held = this.held.get(userId)
if (held === undefined) return false
const ok = await this.db.renewInstanceLease(userId, held.hostId, held.epoch, this.ttl)
if (!ok) this.held.delete(userId)
return ok
}
/** 批量续租(心跳 tick 用)。返回失权的 userId 列表(调用方据此 self-fence)。 */
async renewAll(): Promise<string[]> {
const lost: string[] = []
for (const userId of [...this.held.keys()]) {
if (!(await this.renew(userId))) lost.push(userId)
}
return lost
}
/** 主动释放(停实例时)。成功后不再持有该用户。 */
async release(userId: string): Promise<boolean> {
const held = this.held.get(userId)
if (held === undefined) return false
const ok = await this.db.releaseInstanceLease(userId, held.hostId, held.epoch)
if (ok) this.held.delete(userId)
return ok
}
/** 重新读取 DB 里的真实归属,校正本地记录(对账用)。 */
async refresh(userId: string): Promise<DshInstance | undefined> {
const inst = await this.db.findUserInstance(userId, 'main')
const held = this.held.get(userId)
if (!stillHolder(inst, held?.hostId ?? this.hostId, held?.epoch ?? -1)) this.held.delete(userId)
return inst
}
/** **我名下**租约已过期的实例(供巡检;**不等于可以接管**,见模块头与 R9)。 */
async expiredHere(): Promise<DshInstance[]> {
const all = await this.db.listExpiredInstanceLeases(this.now())
const mine = new Set([...this.held.values()].map((h) => h.hostId))
mine.add(this.hostId)
return all.filter((inst) => inst.hostId !== null && mine.has(inst.hostId))
}
/** 全集群租约已过期的实例(管理面/巡检用;**不自动接管**)。 */
async expiredAll(): Promise<DshInstance[]> {
return this.db.listExpiredInstanceLeases(this.now())
}
/** 对账:**一次拿回本机全部实例**(替代逐用户查询,设计 §11.6)。 */
async mine(): Promise<DshInstance[]> {
return this.db.listInstancesByHost(this.hostId)
}
}
+284
View File
@@ -0,0 +1,284 @@
/**
* 给任意 `Spawner` 套上**归属租约**(T08 S4;设计 §1.2/§3.2/§11.5)。
*
* 它补上集群模式下 Manager 侧最关键的一环:**"能不能拉起"必须先问过归属**。
* 本地模式靠"进程内 Map + 单机"天然保证单写者;多机后这个保证只能落在 DB 的原子 CAS 上
* —— 两个 Manager 各持一把租约、抢同一个用户,就是**双写同一个 home** = 数据损坏。
*
* 三个动作:
* 1. `launch` 前 **claim**:拿不到就抛 `LeaseBusyError`(**退让**,不是接管 —— 见 R9);
* 2. 心跳里 **renewAll**:续租;同时把"我已失权"的实例用 **`/fence`** 通知 worker 停掉
* (self-fencing 的 Manager 侧对齐,设计 §11.5);
* 3. `stop` 时 **release**:归属交还,别人立刻可以接管(不用等 TTL)。
*
* ⚠️ 归属只由 Manager 写(设计 §1.3 数据分层的判据 1)。worker 侧只被通知。
*
* @module dshs/supervisor/leased-spawner
*/
import { AGENT_TOKEN_HEADER } from '../worker/agent.js'
import type { DbAdapter } from '../db/adapter.js'
import type { Endpoint, Instance, Spawner, UserStatus } from './spawner.js'
import { InstanceLease, type LeaseOptions } from './lease.js'
/** 归属被别人持有时抛出 —— 调用方应**退让**(等待/报告),**不得接管**(R9)。 */
export class LeaseBusyError extends Error {
readonly userId: string
readonly holder: string | null
readonly leaseUntil: number
constructor(userId: string, holder: string | null, leaseUntil: number) {
super(`instance ${userId} is held by ${holder ?? 'someone'} until ${new Date(leaseUntil).toISOString()}`)
this.name = 'LeaseBusyError'
this.userId = userId
this.holder = holder
this.leaseUntil = leaseUntil
}
}
export interface LeasedSpawnerOptions extends LeaseOptions {
/** 本机(= 它所属的 worker)在 `dsh_hosts.id` 里的标识。 */
hostId: string
/** worker agent 基址。 */
agentUrl: string
/** 与 agent 约定的共享密钥。 */
agentToken: string
/** 心跳间隔(ms)。默认 = 续租间隔。 */
heartbeatMs?: number
/** 该 worker 的内存预算(MB);**0 = 不承载实例**(只做门户/控制,设计 §15.3)。 */
capacityMb?: number
/**
* **选机**(T08 S6):返回这次要把实例放到的 `hostId`。
*
* ⚠️ 实现必须**先粘性、再容量**(2026-09-15 生产切换时补的设计缺口):
* 用户的工作区是**跟机器走的**(本地盘)⇒ 把"已有历史数据在某台"的用户调度到另一台,
* 他打开实例会看到**空工作区**。所以:有历史归属且那台还 `up` ⇒ **留在原地**;
* 只有"从没有过归属"(新用户)才按容量挑最空的。
* 传入 `userId` 就是为了让实现能做这件事。
*/
selectHost?: (userId?: string) => Promise<string | undefined>
/** **hostId → agent 地址/密钥**(多机时 fence 要发给"实例所在的那台")。 */
agentFor?: (hostId: string) => { agentUrl: string; token: string } | undefined
/**
* 启动时把自己注册进 `dsh_hosts`(幂等)。
*
* ⚠️ **一个 agent 只应有一条 host 记录**:`registerSelf` 只在"本 Manager 与 worker 同机"
* (1a 形态)时该开。**专用 Manager 部署必须关掉**(`DSHS_CLUSTER_REGISTER_SELF=0`),
* 否则会多出一条指向同一 agent 的 host 记录 ⇒ 同一个用户可能被两个 hostId 各自认领。
*/
registerSelf?: boolean
/** 关闭心跳(测试里手动 tick 用)。 */
manual?: boolean
}
/** 心跳里上报给 `dsh_hosts` 的本机观测值。 */
export interface HostObservation {
ok: boolean
instances: number
lastError?: string
}
export class LeasedSpawner implements Spawner {
private readonly lease: InstanceLease
private timer: NodeJS.Timeout | undefined
private lastObservation: HostObservation | undefined
private readonly heartbeatMs: number
constructor(
private readonly inner: Spawner,
private readonly db: DbAdapter,
private readonly options: LeasedSpawnerOptions,
) {
this.lease = new InstanceLease(db, options.hostId, options)
this.heartbeatMs = options.heartbeatMs ?? this.lease.renewMs
}
get hostId(): string {
return this.options.hostId
}
/** 最近一次心跳观测(管理面/诊断用)。 */
observation(): HostObservation | undefined {
return this.lastObservation
}
/** 本 Manager 当前持有的实例(userId → {epoch, hostId})。 */
holdings(): ReadonlyMap<string, { epoch: number; hostId: string }> {
return this.lease.holdings()
}
/** 注册本机 + 起心跳。窗口未开时先注册一次(否则管理面看不到这台 worker)。 */
async start(): Promise<void> {
if (this.options.registerSelf !== false) {
await this.db.upsertDshHost({
id: this.options.hostId,
endpoint: this.options.agentUrl,
agentToken: this.options.agentToken,
capacityMb: this.options.capacityMb ?? 0,
})
}
await this.tick() // 立即一次,管理面马上能看到心跳
if (this.options.manual === true) return
this.timer = setInterval(() => void this.tick(), this.heartbeatMs)
this.timer.unref?.()
}
/** 只停**心跳定时器**(不改实例)—— 注意别和 `Spawner.stop(userId)` 混淆,故另起名。 */
stopHeartbeat(): void {
if (this.timer !== undefined) {
clearInterval(this.timer)
this.timer = undefined
}
}
/** 一次心跳:续租 → 失权则 fence → 上报本机状态。 */
async tick(): Promise<void> {
// ① 续租;失权的 userId 会被清出本地记录
const lost = await this.lease.renewAll()
for (const userId of lost) {
// ② 我已失权 ⇒ 让 worker 停掉那个实例(下发的 epoch 取 DB 当前值 +1,确保高于它的记录)
await this.fenceOnAgent(userId)
}
await this.reportHost()
}
private async fenceOnAgent(userId: string): Promise<void> {
try {
const inst = await this.db.findUserInstance(userId, 'main')
// 多机(T08 S6):必须发给**实例所在的那台** —— 发错 host 等于没拦(旧持有者继续写)
const target = inst?.hostId === null || inst?.hostId === undefined
? { agentUrl: this.options.agentUrl, token: this.options.agentToken }
: (this.options.agentFor?.(inst.hostId) ?? { agentUrl: this.options.agentUrl, token: this.options.agentToken })
await this.post(`/fence`, { userId, epoch: (inst?.epoch ?? 0) + 1 }, target)
} catch {
// 通知失败不抛:下一轮心跳会重试;即便一直失败,租约已过期 ⇒ 新持有者会重建实例
}
}
private async reportHost(): Promise<void> {
const base = this.options.agentUrl.replace(/\/$/, '')
try {
const res = await fetch(`${base}/healthz`, { signal: AbortSignal.timeout(this.heartbeatMs) })
const body = (await res.json()) as { ok?: boolean; instances?: number }
this.lastObservation = { ok: body.ok === true, instances: body.instances ?? 0 }
await this.db.setDshHostStatus(
this.options.hostId,
this.lastObservation.ok ? 'up' : 'down',
undefined,
Date.now(),
)
} catch (err) {
this.lastObservation = { ok: false, instances: 0, lastError: err instanceof Error ? err.message : String(err) }
await this.db.setDshHostStatus(this.options.hostId, 'down', undefined, Date.now())
}
}
private async post(
path: string,
body: Record<string, unknown>,
target?: { agentUrl: string; token: string },
): Promise<unknown> {
const use = target ?? { agentUrl: this.options.agentUrl, token: this.options.agentToken }
const base = use.agentUrl.replace(/\/$/, '')
const res = await fetch(`${base}${path}`, {
method: 'POST',
headers: { [AGENT_TOKEN_HEADER]: use.token, 'content-type': 'application/json' },
body: JSON.stringify(body),
signal: AbortSignal.timeout(10_000),
})
if (!res.ok) throw new Error(`agent POST ${path} → ${res.status}`)
return res.json()
}
// ── Spawner 实现 ────────────────────────────────────────────────────────
async launch(
userId: string,
folder: string,
patch?: string,
opts?: { force?: boolean; epoch?: number; hostId?: string },
): Promise<Instance> {
// **先选机、再认领**(T08 S6):租约的 host_id 必须与"实例真正落在哪台"一致,
// 否则续租/释放会指向错误的对象(多机下就是静默的脑裂入口)。
// 显式 `opts.hostId`(迁移的目标机)优先于自动选机。
const hostId = opts?.hostId ?? (await this.options.selectHost?.(userId)) ?? this.options.hostId
const claim = await this.lease.acquire(userId, hostId, { folder, patch })
if (!claim.ok) throw new LeaseBusyError(userId, claim.holder, claim.leaseUntil)
try {
return await this.inner.launch(userId, folder, patch, { ...opts, epoch: claim.epoch, hostId })
} catch (err) {
// 拉起失败就**立刻交还归属** —— 否则要白等一个 TTL 才能重试(用户侧表现为"卡住")
await this.lease.release(userId)
throw err
}
}
async restartMain(userId: string): Promise<Instance | undefined> {
const current = await this.inner.status(userId)
if (current.main === undefined) return undefined
const { folder, patch } = current.main
await this.stop(userId)
return this.launch(userId, folder, patch)
}
/** 只重启**本 Manager 持有**的实例 —— 别人的归属不该被我重启(会与其持有者抢同一个 home)。 */
async restartAllMains(): Promise<void> {
for (const userId of [...this.lease.holdings().keys()]) {
try {
await this.restartMain(userId)
} catch {
// 单个失败不打断其余
}
}
}
async spawnWatchdog(userId: string): Promise<Instance | undefined> {
return this.inner.spawnWatchdog(userId)
}
async status(userId: string): Promise<UserStatus> {
return this.inner.status(userId)
}
async endpointFor(userId: string): Promise<Endpoint | undefined> {
return this.inner.endpointFor(userId)
}
async stop(userId: string, hostId?: string): Promise<void> {
// 显式 host 优先;否则用**我认领时那台**(认领记录里有)—— 别让 stop 落到别的 worker 上
const target = hostId ?? this.lease.holdings().get(userId)?.hostId
await this.inner.stop(userId, target)
await this.lease.release(userId)
}
async teardown(): Promise<void> {
this.stopHeartbeat()
await this.inner.teardown()
}
async waitForLaunchTokenForUser(userId: string, timeoutMs?: number): Promise<void> {
return this.inner.waitForLaunchTokenForUser(userId, timeoutMs)
}
async restartAndProbe(userId: string, settleMs?: number): Promise<{ ok: boolean; reason: string }> {
return this.inner.restartAndProbe(userId, settleMs)
}
touch(userId: string): void {
this.inner.touch(userId)
}
async ensureFileService(userId: string): Promise<void> {
return this.inner.ensureFileService(userId)
}
/**
* 透传可选观测面。`Spawner` 里这两个是**可选**方法 ⇒ 这里做条件委托:
* 内层有就转发(熔断/配额是 worker 本地自管的概念,设计 §1.2),没有就回 null。
*/
breakerInfo(userId: string): { opens: number; openedAt: number; cooldownUntil: number } | null {
return this.inner.breakerInfo?.(userId) ?? null
}
quotaInfo(userId: string): { baseMb: number; memMb: number; heapMb: number } | null {
return this.inner.quotaInfo?.(userId) ?? null
}
}
+103 -1
View File
@@ -22,7 +22,7 @@ import {
realpathSync,
writeFileSync,
} from 'node:fs'
import { join } from 'node:path'
import { dirname, join } from 'node:path'
import type { ServerConfig } from '../config.js'
import { handoffPath, homeRoot, userRoot, workspaceRoot } from '../fs/workspace.js'
import {
@@ -129,6 +129,47 @@ function withHeap(base: string | undefined, memMb: number): string {
}
/**
* 列出 `dest` 与 `stopAt` 之间的**祖先目录**(由外到内),用于 bwrap 的 `--tmpfs`(见
* {@link mountParentDirArgs}:`--tmpfs` 自带 0755,且**兼容 47 上的 bwrap 0.4.0**)。
*
* 背景(T08 S1.6,2026-09-15 实测):bwrap **只创建挂载点本身**,沿途缺失的父目录由它自建,
* 而权限是 **`0700 root:root`** —— 实测 `--bind /opt/a/b/c /opt/a/b/c` 会得到 `/opt`、
* `/opt/a`、`/opt/a/b` **全是 0700**。后果:**实例以非 root 的 uid 穿越这些路径时 EACCES**。
* 已实测到的两处症状:
* ① 宿主上 `/etc/ssl/openssl.cnf` 是指向 `/etc/pki/tls/openssl.cnf` 的**符号链接** ⇒ 解析要穿过
* `/etc/pki`(0700)⇒ node 报 `OpenSSL configuration error … Permission denied`、**exitCode 13**
* (106 / OpenCloudOS 9.6 实测;47 上**没有**该文件故静默跳过 ⇒ 同一份代码一台能跑一台崩);
* ② **用户工作区在沙箱内不可穿越** ⇒ 实例按**绝对路径**读写自己的文件被拒。
* 修法:把这些祖先目录**显式建成 0755**。**权限不扩大** —— 这些目录里只有随后绑定的白名单内容
* (整绑 `/etc/pki` 的替代方案已否决:会带入 `/etc/pki/tls/private/postfix.key`,违反 R5)。
*/
function mountParentDirList(dest: string, stopAt: string): string[] {
const dirs: string[] = []
let cur = dirname(dest)
while (cur !== stopAt && cur !== '/' && cur !== '' && cur !== '.') {
dirs.push(cur)
cur = dirname(cur)
}
return dirs.reverse() // 由外到内(`--perms` 只作用于紧接着的那一个 `--dir`)
}
/**
* 把 {@link mountParentDirList} 的结果摊平成 bwrap 参数。
* ⚠️ **不要再"就近调用"它**(例如插在 `--bind` 之前)—— 见 `bwrapArgs` 里"统一前置"
* 那段注释:就近创建会在嵌套前缀下遮掉已绑好的挂载点。当前实现只在**一处**统一使用。
*/
function mountParentDirArgs(dest: string, stopAt: string): string[] {
// ⚠️ **必须用 `--tmpfs`,不能用 `--perms 0755 --dir`**(2026-09-15 实测,差点打断生产):
// · `--perms` 是 bubblewrap **0.5+** 才有的选项;**47 上是 0.4.0** ⇒ 传了直接
// `bwrap: Unknown option --perms` ⇒ **沙箱起不来 = 所有实例全挂**(106 是 0.11.0,能过)。
// · 而 `--tmpfs` **自带 0755**(本文件另一处注释也这么写:`bwrap 的 --tmpfs 权限是 755`),
// 且在 0.4.0 上就可用。
// 47 上实测:改用 `--tmpfs` 后 `/etc` 可见条目 **78 → 78(零变化)**,且 `/etc/pki` 权限
// 由 `drwx------` 变为 `drwxr-xr-x`(可穿越)。代价 = 每个中间目录多一个空 tmpfs 挂载(极小)。
return mountParentDirList(dest, stopAt).flatMap((d) => ['--tmpfs', d])
}
/**
* Local backend: owns the lifecycle of per-user DSH process pairs via
* child_process. State is in-memory. Implements {@link Spawner}.
@@ -361,6 +402,26 @@ export class LocalSpawner implements Spawner {
return { main: this.mains.get(userId), watchdog: this.watchdogs.get(userId) }
}
/**
* 整机视角的实例清单(T08 S3:worker agent 的 `/instances` 用)。
*
* 口径 = **每个用户的 main 实例**(watchdog 是一次性 headless,不进对账口径)。
* 设计上这是 `/instances` "一次拿回整机"的实现,替代逐用户查询(设计 §11.6)。
*/
listUserInstances(): Instance[] {
return [...this.mains.values()]
}
/**
* 当前 main 实例的 launch token(T08 S3 P0-6)。
*
* 本地模式下 token 从子进程 stdout 解析出来;跨机后 **agent 必须把它回传 Manager**,
* 否则「登录直达会话」(档案 06/13/15)与实例侧 401 自愈(档案 24/50/51)都会失效。
*/
launchTokenOf(userId: string): string | undefined {
return this.mains.get(userId)?.launchToken
}
/** Endpoint the proxy forwards to (local → the running main's loopback port). */
async endpointFor(userId: string): Promise<Endpoint | undefined> {
const port = this.mains.get(userId)?.port
@@ -680,6 +741,30 @@ export class LocalSpawner implements Spawner {
'/etc/ssl',
]
const out: string[] = []
// 2026-09-15(T08 S1.6)**中间挂载点必须可穿越(0755)**。
//
// 现象(106 / OpenCloudOS 9.6 实测):实例起不来,子进程报
// `/usr/bin/node: OpenSSL configuration error: … Permission denied:
// … fopen(/etc/ssl/openssl.cnf, rb)` ⇒ **exitCode 13**。
// 根因:bwrap 会为 `--ro-bind-try /etc/pki/tls/certs …` 这类路径**自动补齐父目录**,
// 而这些自动创建的目录权限是 **0700(drwx------ root:root)** ⇒ 非 root 的实例
// **无法穿越**;宿主上 `/etc/ssl/openssl.cnf` 恰好是**指向 `/etc/pki/tls/openssl.cnf`
// 的符号链接** ⇒ 解析要穿过 `/etc/pki` → 被拒 → 报 **EACCES(不是 ENOENT)**
// → node 读 OpenSSL 配置**硬失败**。
// 为什么 47 没事:Alibaba Cloud Linux 3 上**没有** `/etc/ssl/openssl.cnf`
// ⇒ node 静默跳过 ⇒ **同一份代码一台能跑、一台崩**(机器基线差异,设计 §14.3)。
// 修法:在绑定**之前**把白名单路径在 `/etc` 下的所有中间目录显式建成 0755。
// 权限**不扩大**:`/etc` 在本沙箱里是 tmpfs,这些目录里只有下面白名单绑定的内容,
// 不新增任何宿主可见面。("整绑 `/etc/pki`"的替代方案已否决 —— 会顺带带入
// `/etc/pki/tls/private/postfix.key`,违反 **R5 权限只准收窄**。)
// 注意顺序:由外到内(内层挂载点要求外层已存在)。
const intermediates = new Set<string>()
for (const p of allow) {
for (const d of mountParentDirList(p, '/etc')) intermediates.add(d)
}
for (const d of [...intermediates].sort((a, b) => a.split('/').length - b.split('/').length)) {
out.push('--tmpfs', d) // 见 mountParentDirArgs 的注释:`--tmpfs` 自带 0755 且兼容 bwrap 0.4.0
}
for (const p of allow) {
let src = p
try {
@@ -691,6 +776,23 @@ export class LocalSpawner implements Spawner {
}
return out
})(),
// ── 所有挂载点的**中间目录**统一在这里建好(T08 S1.6 修正版)────────────────
// 为什么必须"统一前置 + 去重 + 由外到内"(2026-09-15 实测踩到的真 bug):
// 用户根(`--bind root root`)与共享技能层(`--ro-bind-try skill skill`)**可能嵌套在
// 同一前缀下**。若按"就近创建"把技能层的中间目录插在 `--bind root root` **之后**,
// 那么后挂的 `--tmpfs <共同祖先>` 会把**已经绑好的用户根整个遮掉** ⇒ bwrap 报
// `Can't chdir to <userRoot>/ws/xxx: No such file or directory` ⇒ 实例崩溃循环。
// 前置 + 去重后,中间目录只建一次,后续所有 bind 都落在它里面,谁也不遮谁。
// 权限不扩大:这些目录里只有随后绑定的白名单内容。
...(() => {
const dirs = new Set<string>()
for (const dest of [root, this.config.bundledSkillDir].filter((d) => d !== '')) {
for (const d of mountParentDirList(dest, '/')) dirs.add(d)
}
return [...dirs]
.sort((a, b) => a.split('/').length - b.split('/').length)
.flatMap((d) => ['--tmpfs', d])
})(),
'--dev', '/dev', '--proc', '/proc',
'--bind', tmpDir, '/tmp',
'--bind', root, root,
+330
View File
@@ -0,0 +1,330 @@
/**
* `Spawner` 的**远端实现**(T08 S3 单机 / S6 多机;设计 §1.1/§11)。
*
* 路由层只依赖 `Spawner` 接口(见 `spawner.ts` 头注释),所以本类**不触碰路由与代理**
* —— 代理层把 `endpointFor` 返回的 `{host, port}` 直连即可(`proxy.ts` 的 TCP 目标
* 与 Host 头本就是分开处理的,跨机不需要改信任逻辑)。
*
* 三条协议纪律:
* 1. **幂等键复用**:同一次调用的重试**复用同一个 `operationId`** —— 否则 agent 会把
* "Manager 超时后重发"当成新请求,起出两个实例(设计 §11.3 手段 3)。
* 2. **重试有界**:3 次(200ms / 1s / 3s),仍失败就**抛错**,由上层决定退让或告警;
* 错误信息里带上状态码与响应体片段,避免"静默失败"。
* 3. **按 host 路由**(S6):每个用户的操作都落到**它实例所在的那台** —— 依据是
* `dsh_instances.host_id`,由上层以 `hostIdFor` 注入(本类不直接连 DB)。
*
* ⚠️ 本类**不持有归属租约**:租约由 Manager 侧的 `InstanceLease`/`LeasedSpawner` 管理。
*
* @module dshs/supervisor/remote-spawner
*/
import { randomUUID } from 'node:crypto'
import { AGENT_TOKEN_HEADER } from '../worker/agent.js'
import type { Endpoint, Instance, Spawner, UserStatus } from './spawner.js'
/** 一台 worker 的接入信息。 */
export interface ClusterHost {
hostId: string
/** agent 基址。 */
agentUrl: string
/** 与 agent 约定的共享密钥。 */
token: string
/** 代理时使用的主机(同机 1a = `127.0.0.1`;跨机填 Worker 内网 IP)。 */
instanceHost?: string
}
export interface RemoteSpawnerOptions {
/** agent 基址,如 `http://127.0.0.1:9000`。 */
agentUrl: string
/** 与 agent 约定的共享密钥。 */
token: string
/** 代理时使用的主机(同机 1a = `127.0.0.1`;跨机填 Worker 内网 IP)。 */
instanceHost?: string
/** 单次请求超时(ms)。控制通道是短请求,默认 10 s(设计 §11.2)。 */
timeoutMs?: number
/** 注入 fetch(测试用)。 */
fetchImpl?: typeof fetch
/**
* 解析该用户的模型 key 与 uid —— Manager 侧解析后**随 launch 投递**给 agent。
* 为什么不投递"让 Worker 自己查凭据库":那是**权限扩大**(Worker 就能读全量用户的 key),
* 而投递是收窄到"本次实例那一把"。**注意**:这条只管**控制面凭据**,与"Worker 能不能有
* 自己的库"无关(插件数据在实例 home 里,见设计 §1.3 数据分层)。
*/
resolveApiKey?: (userId: string) => Promise<string | null>
resolveUid?: (userId: string) => Promise<number>
/** 默认 host 的 id(不提供 `hosts` 时的单机形态用它)。 */
defaultHostId?: string
/** **多机(S6)**:除默认 host 外的其它 worker。给了就按 `hostId` 路由。 */
hosts?: ClusterHost[]
/** **多机(S6)**:`userId` → 它实例所在的 `hostId`(上层查 `dsh_instances.host_id` 注入)。 */
hostIdFor?: (userId: string) => Promise<string | undefined>
/**
* **host 目录的来源**(S6):从 `dsh_hosts` 读。给它就**不必预知 worker 列表**,
* 且**新增 worker 无需重启 Manager**(TTL 内自动生效)。
*/
hostsProvider?: () => Promise<ClusterHost[]>
/** 目录缓存时长(ms)。默认 30 s —— 与心跳同量级。 */
directoryTtlMs?: number
}
const RETRY_DELAYS_MS = [200, 1000, 3000]
export class RemoteSpawner implements Spawner {
private readonly defaultHost: ClusterHost
private readonly hosts = new Map<string, ClusterHost>()
private readonly timeoutMs: number
private readonly doFetch: typeof fetch
private readonly resolveApiKey?: (userId: string) => Promise<string | null>
private readonly resolveUid?: (userId: string) => Promise<number>
private readonly hostIdFor?: (userId: string) => Promise<string | undefined>
private readonly hostsProvider?: () => Promise<ClusterHost[]>
private readonly directoryTtlMs: number
private directoryLoadedAt = 0
constructor(options: RemoteSpawnerOptions) {
this.defaultHost = {
hostId: options.defaultHostId ?? 'local',
agentUrl: options.agentUrl.replace(/\/$/, ''),
token: options.token,
instanceHost: options.instanceHost ?? '127.0.0.1',
}
this.hosts.set(this.defaultHost.hostId, this.defaultHost)
for (const host of options.hosts ?? []) {
this.hosts.set(host.hostId, { ...host, agentUrl: host.agentUrl.replace(/\/$/, '') })
}
this.timeoutMs = options.timeoutMs ?? 10_000
this.doFetch = options.fetchImpl ?? fetch
this.resolveApiKey = options.resolveApiKey
this.resolveUid = options.resolveUid
this.hostIdFor = options.hostIdFor
this.hostsProvider = options.hostsProvider
this.directoryTtlMs = options.directoryTtlMs ?? 30_000
}
/**
* 按需刷新 host 目录(TTL 内不重复查询)。
* **默认 host 始终在表里**(配置里那台),即使它还没注册进 `dsh_hosts`。
*/
private async ensureHosts(): Promise<void> {
if (this.hostsProvider === undefined) return
if (Date.now() - this.directoryLoadedAt < this.directoryTtlMs) return
this.directoryLoadedAt = Date.now()
try {
for (const host of await this.hostsProvider()) {
this.hosts.set(host.hostId, { ...host, agentUrl: host.agentUrl.replace(/\/$/, '') })
}
this.hosts.set(this.defaultHost.hostId, this.defaultHost)
} catch {
// 查库失败就沿用旧目录(可能是全库不可用的前兆,由心跳/管理面暴露)
}
}
/** 已知的 host 目录(管理面/诊断用)。 */
knownHosts(): ClusterHost[] {
return [...this.hosts.values()]
}
/** 强制刷新目录(管理面/测试用)。 */
async reloadHosts(): Promise<void> {
this.directoryLoadedAt = 0
await this.ensureHosts()
}
/** hostId → 接入信息;未知 host 回退到默认(单机形态下这就是唯一那台)。 */
hostById(hostId: string | null | undefined): ClusterHost {
if (hostId === null || hostId === undefined) return this.defaultHost
return this.hosts.get(hostId) ?? this.defaultHost
}
/**
* 决定这次操作落到哪台。优先级:**显式指定**(迁移目标机、launch 时选好的机)
* → **该用户实例的归属**(`hostIdFor`)→ 默认 host。
*/
private async hostFor(userId: string, explicit?: string): Promise<ClusterHost> {
await this.ensureHosts()
if (explicit !== undefined) {
let found = this.hosts.get(explicit)
if (found === undefined) {
// 目录有 30s TTL:显式指定的 host 可能"刚 join 还没进目录" ⇒ 强制刷一次
this.directoryLoadedAt = 0
await this.ensureHosts()
found = this.hosts.get(explicit)
}
// ⛔ 仍然找不到就**报错**,绝不回退到默认 host ——
// "租约认领在 A、实例却起在 B"是多机下最危险的静默失败(归属与实例分离)。
if (found === undefined) throw new Error(`unknown host "${explicit}":不在 host 目录里,拒绝改投到别的 worker`)
return found
}
if (this.hostIdFor !== undefined) {
const owned = await this.hostIdFor(userId)
if (owned !== undefined) {
const found = this.hosts.get(owned)
if (found !== undefined) return found
}
}
return this.defaultHost
}
/** 带重试的请求。`operationId` 由调用方生成并在重试间**保持不变**(幂等)。 */
private async call<T>(
host: ClusterHost,
method: 'GET' | 'POST',
path: string,
body?: Record<string, unknown>,
operationId?: string,
): Promise<T> {
const payload = body === undefined ? undefined : { ...body, ...(operationId === undefined ? {} : { operationId }) }
let lastErr: unknown
for (let attempt = 0; attempt <= RETRY_DELAYS_MS.length; attempt += 1) {
if (attempt > 0) await new Promise((r) => setTimeout(r, RETRY_DELAYS_MS[attempt - 1]))
try {
const res = await this.doFetch(`${host.agentUrl}${path}`, {
method,
headers: {
[AGENT_TOKEN_HEADER]: host.token,
...(payload === undefined ? {} : { 'content-type': 'application/json' }),
},
body: payload === undefined ? undefined : JSON.stringify(payload),
signal: AbortSignal.timeout(this.timeoutMs),
})
if (res.ok) return (await res.json()) as T
const text = await res.text()
// 4xx 是"协议/参数错",重试没意义;5xx 与网络错才重试。
if (res.status < 500) throw new Error(`agent ${method} ${path} → ${res.status}: ${text.slice(0, 200)}`)
lastErr = new Error(`agent ${method} ${path} → ${res.status}: ${text.slice(0, 200)}`)
} catch (err) {
lastErr = err
}
}
throw lastErr instanceof Error ? lastErr : new Error(String(lastErr))
}
async launch(
userId: string,
folder: string,
patch?: string,
opts?: { force?: boolean; epoch?: number; hostId?: string },
): Promise<Instance> {
const host = await this.hostFor(userId, opts?.hostId)
// 同一个 operationId 贯穿这次调用的所有重试 ⇒ agent 侧幂等回放(不会起两个实例)。
const operationId = randomUUID()
const apiKey = this.resolveApiKey === undefined ? null : await this.resolveApiKey(userId)
const uid = this.resolveUid === undefined ? undefined : await this.resolveUid(userId)
const res = await this.call<{ instance: Instance; note?: string }>(
host,
'POST',
'/launch',
{ userId, folder, patch, apiKey, uid, epoch: opts?.epoch },
operationId,
)
return res.instance
}
async restartMain(userId: string, hostId?: string): Promise<Instance | undefined> {
const current = await this.status(userId)
if (current.main === undefined) return undefined
const { folder, patch } = current.main
const host = await this.hostFor(userId, hostId)
await this.stop(userId, host.hostId)
return this.launch(userId, folder, patch, { hostId: host.hostId })
}
/** 重启**所有 host 上**能看到的实例(跨机聚合;上层 `LeasedSpawner` 会限定在自己持有的范围内)。 */
async restartAllMains(): Promise<void> {
for (const host of this.hosts.values()) {
let instances: Array<{ userId: string }> = []
try {
const res = await this.call<{ instances: Array<{ userId: string }> }>(host, 'GET', '/instances')
instances = res.instances
} catch {
continue // 该 host 不可达:跳过(心跳/告警负责暴露)
}
for (const inst of instances) {
try {
await this.restartMain(inst.userId, host.hostId)
} catch {
// 单台失败不打断其余(与 LocalSpawner 的语义一致)
}
}
}
}
async spawnWatchdog(userId: string): Promise<Instance | undefined> {
const host = await this.hostFor(userId)
const res = await this.call<{ instance: Instance | null }>(host, 'POST', `/watchdog/${encodeURIComponent(userId)}`)
return res.instance ?? undefined
}
async status(userId: string): Promise<UserStatus> {
const host = await this.hostFor(userId)
const res = await this.call<{ main: Instance | null }>(host, 'GET', `/status/${encodeURIComponent(userId)}`)
return res.main === null ? {} : { main: res.main }
}
/** 代理目标:由**实例所在那台** agent 给端口(未运行 → undefined,代理会走冷启动分支)。 */
async endpointFor(userId: string): Promise<Endpoint | undefined> {
const host = await this.hostFor(userId)
const res = await this.call<{ running: boolean; host?: string; port?: number }>(
host,
'GET',
`/endpoint/${encodeURIComponent(userId)}`,
)
return res.running && res.host !== undefined && res.port !== undefined
? { host: res.host, port: res.port }
: undefined
}
async stop(userId: string, hostId?: string): Promise<void> {
const host = await this.hostFor(userId, hostId)
await this.call(host, 'POST', '/stop', { userId }, randomUUID())
}
async teardown(): Promise<void> {
// 远端实例的寿命长于任何单个 Manager 副本 ⇒ 由 Manager 的归属/租约管理,不在关闭时清。
}
/**
* 等 launch token 出现(本地模式 = 启动完成的信号)。
* 跨机后 token 由 agent 回传,语义不变;超时返回(不抛)—— 与本地实现一致。
*/
async waitForLaunchTokenForUser(userId: string, timeoutMs = 20_000): Promise<void> {
const deadline = Date.now() + timeoutMs
while (Date.now() < deadline) {
try {
const st = await this.status(userId)
if (st.main !== undefined && st.main.launchToken !== undefined) return
if (st.main === undefined) return // 没实例/已停 ⇒ 立即返回(与本地实现一致)
} catch {
// 网络抖动:继续等
}
await new Promise((r) => setTimeout(r, 200))
}
}
async restartAndProbe(userId: string, settleMs?: number): Promise<{ ok: boolean; reason: string }> {
const host = await this.hostFor(userId)
return this.call<{ ok: boolean; reason: string }>(
host,
'POST',
`/restart-probe/${encodeURIComponent(userId)}`,
settleMs === undefined ? {} : { settleMs },
)
}
/** 活动信号:转发给**实例所在那台** agent,让它自己的 idle-reap 不误杀(fire-and-forget)。 */
touch(userId: string): void {
void this.hostFor(userId)
.then((host) =>
this.doFetch(`${host.agentUrl}/touch/${encodeURIComponent(userId)}`, {
method: 'POST',
headers: { [AGENT_TOKEN_HEADER]: host.token },
}),
)
.catch(() => {
/* 活动信号丢了不影响正确性 */
})
}
async ensureFileService(_userId: string): Promise<void> {
// worker 本机就有用户卷(local 语义)⇒ 无需 sidecar。跨机文件面由 RemoteUserFs(/fs/*)承担。
}
}
+17 -2
View File
@@ -92,7 +92,22 @@ export interface Endpoint {
* `LocalSpawner` materializes it to a file inside the user's own volume.
*/
export interface Spawner {
launch(userId: string, folder: string, patch?: string, opts?: { force?: boolean }): Promise<Instance>
/**
* 拉起实例。
*
* `opts.epoch`(T08 S4):**集群模式下 epoch 是 launch 契约的一部分** —— Manager 先抢占
* 归属拿到 epoch,再把它随 launch 下发,worker 记下来用于 **self-fencing**
* (收到更高 epoch 就停掉自己那个实例)。本地模式忽略该字段。
*
* `opts.hostId`(T08 S6):**多 worker 时指定落到哪台** —— 由上层选好机、并已用它认领租约,
* 因此这里必须与租约的 `host_id` 一致(否则归属与实例分离)。单机/1a 忽略。
*/
launch(
userId: string,
folder: string,
patch?: string,
opts?: { force?: boolean; epoch?: number; hostId?: string },
): Promise<Instance>
restartMain(userId: string): Promise<Instance | undefined>
/**
* 档案 78:熔断观测面(可选 —— 熔断是**本地模式**概念,k8s 模式没有)。
@@ -112,7 +127,7 @@ export interface Spawner {
spawnWatchdog(userId: string): Promise<Instance | undefined>
status(userId: string): Promise<UserStatus>
endpointFor(userId: string): Promise<Endpoint | undefined>
stop(userId: string): Promise<void>
stop(userId: string, hostId?: string): Promise<void>
teardown(): Promise<void>
/** 等待该用户 main 实例打印 launch token(本地模式 = 启动完成的信号)。无实例 /
* 已崩溃 / 已停 → 立即返回;k8s 模式无 token 概念 → no-op。供 enter 复用分支在返回