feat(cost): 省积分机制 —— 会话预算告警 + Bash 大输出门禁

本会话实测:553 轮 × 平均 42 万 token = 2.33 亿 input,output 仅占 0.35%(286:1)。
根因 = **对话历史是 append-only**:工具输出(`function_call_result`)一旦生成就**永久留在 messages 里、
每轮全量重发**,模型无权删自己的历史 ⇒ **唯一能"去掉"的时机 = 输出被生成之前**
(平台 `contextWindow=100 万` / 压缩阈值 90% 只是放大器,不是根因)。

- `scripts/stop-dialog-guard.py`:
  ① 修 `-S -E` ⇒ stdin 回退 cp936 ⇒ 含中文 payload 解析失败导致的**静默空转**(改走 buffer 显式 UTF-8,
     读写都加固;入口即留痕;根治"日志缺失时无法区分『没被调用』与『被静默 return』")
  ② 新增 `session_budget()`:读转录 `usage.input_tokens` ⇒ 当前体量 + 累计工具调用 + **上一轮增量**
  ③ 阈值 `BUDGET_TOKENS=120000` / `BUDGET_TOOLS=80` ⇒ 超预算经 `UserPromptSubmit` 注入告警
- `scripts/bash-output-guard.py`(新):`PreToolUse(Bash)` 拦"几乎必然巨大"的 6 条读命令
  (`cat *.log/jsonl` 无管道 / `grep -r` 无管道 / `ls -laR` / `find /` / `journalctl` 无 `-n` / `dmesg` 无 `head`),
  deny 文案**必给等价限流写法**。
  **设计(用户质疑"全拦也有问题")**:⛔ **不全拦** —— 误拦挡住正事比漏拦更贵 ⇒
  默认只拦高置信度、**拿不准一律放行**、`rg`/`git log`/`du·tree` 仅 `hard` 模式拦、
  可 `off` 急停(`<工作区>/.workbuddy/bash-guard-mode`)、任何异常 **fail-open**。
  实测 8/8:soft 拦 cat/ls-lR、放行 rg/git-log/grep+head;hard 全拦;off 全放;deny JSON 正确输出。
- (`settings.json` 的 `PreToolUse` 已加 `"matcher": "Bash"` 条目 —— **该文件不在仓库内**,
  且 hooks 是**启动时快照** ⇒ **Bash 门禁需完全重启才生效**;脚本内容本身每次现读、改完即生效。)

⚠️ 记录两个已踩的坑:Python `%` 格式串里的**裸 `%` 必须写 `%%`**(否则 `TypeError` ⇒ 在 `except: pass` 里
**静默失效**,表现为"日志有 DENY 但 stdout 空"= 看起来拦了其实没生效)⇒ **hook 必须"写日志 + emit"双动作**。
This commit is contained in:
admin committed 2026-09-15 21:17:08 +08:00
1 parent b617cdbaa7
commit 03c8363960
2 files changed
+249 -4

No files matched your search

+76 -4
View File
@@ -244,6 +244,70 @@ CONTEXT = (
)
# ─────────────────────────────────────────────────────────────
# 会话预算(2026-09-15 用户选 A 案):到 ~15 万 token / ~120 次工具调用 ⇒ 提醒开新会话
# 为什么加进本脚本、而不新装一个 hook:`settings.json` 的 hooks 是**应用启动时快照**(新增条目要重启),
# 而**脚本内容每次调用现读** ⇒ 改这里即刻生效。数据源 = 转录 `type=function_call` → `message.usage.input_tokens`
# (最近一条 usage 即"当前上下文体量",精确,不靠估算)。
# 依据(实测某会话):上下文 5.2 万 → 59.2 万;累计 input 1.93 亿 / output 51.9 万(**371:1**);
# 其中 33 次缓存失效,每次都把 ~50 万 token **按全价**重算 ⇒ 会话越长,单次失效越贵。
BUDGET_TOKENS = 120000
BUDGET_TOOLS = 80
def session_budget(path):
"""返回 (当前上下文 token, 工具调用累计次数);读不到返回 (None, None)。"""
try:
if os.path.getsize(path) > 64 * 1024 * 1024:
return None, None
except OSError:
return None, None
last_in, n_calls, prev_in = None, 0, None
try:
with io.open(path, encoding='utf-8', errors='replace') as f:
for ln in f:
if '"function_call"' not in ln:
continue
n_calls += 1
if '"usage"' not in ln:
continue
try:
o = json.loads(ln)
except ValueError:
continue
if o.get('type') != 'function_call':
continue
u = ((o.get('message') or {}).get('usage')) or {}
v = u.get('input_tokens')
if v:
prev_in, last_in = last_in, int(v)
except OSError:
return None, None, None
return last_in, n_calls, prev_in
def budget_note(tokens, ncalls, prev=None):
"""超预算 ⇒ 返回注入用的提醒串;未超 ⇒ 但增量异常也提醒(<1 万 token 的增量不打扰)。"""
over_t = tokens is not None and tokens >= BUDGET_TOKENS
over_n = ncalls is not None and ncalls >= BUDGET_TOOLS
grew = (tokens is not None and prev is not None and tokens - prev >= 40000) if not (over_t or over_n) else False
if not (over_t or over_n or grew):
return ''
which = ('上下文 %s token' % tokens) if over_t else (
'工具调用 %s 次' % ncalls) if over_n else ('**上一轮新增 %s token**' % (tokens - prev))
return (
'💰 【会话预算告警】本会话已到 **%s**(阈值 %s token / %s 次工具调用)。'
'代价机制(2026-09-15 实测):对话历史是**追加式**的 —— 一次工具调用的输出(`function_call_result`)'
'**永久留在历史里、每轮全量重发**(某会话 553 轮 × 平均 42 万 = **2.33 亿 input**,output 仅 0.35%%)。'
'**处置:本轮收口后开新会话**(新会话起点约 5 万 ⇒ 每轮降到 1/12);'
'并把该记的状态写进记忆 + 按 `dsh-change-workflow` 交付门禁落交接。'
'⚠️ 压增长的三条硬纪律:**❶ 大输出先落盘、只读关键行**(`> /tmp/x.txt` 后 `sed -n`);'
'**❷ 命令层限流**(`| head -30` / `| cut -c1-120` / `grep -c` 代替 `grep`);'
'**❸ 让脚本内部聚合、只 print 摘要** —— ⛔ 禁 `cat` 大文件、无 `head` 的 `grep -r`、`ls -laR`。'
% (which, BUDGET_TOKENS, BUDGET_TOOLS)
)
def mode_of(root):
try:
m = io.open(os.path.join(root or '.', '.workbuddy', 'stop-guard-mode'), encoding='utf-8').read()
@@ -266,11 +330,19 @@ def user_prompt_mode(payload):
text = transcribe_last_assistant(tp)
tail = tail_lines(text, 1) if text else ''
hit = bool(tail) and not RE_QUOTE.search(tail) and bool(RE_BAN.search(tail))
log(root0 or '.', 'invoked(user-prompt)|mode=%s|上轮收尾=征询句:%s|%s'
% (mode, hit, (tail.replace('\n', ' ')[:60] if tail else '(取不到上一轮文本)')))
if hit and mode == 'inject':
toks, ncalls, prev = session_budget(tp)
note = budget_note(toks, ncalls, prev)
log(root0 or '.', 'invoked(user-prompt)|mode=%s|上轮收尾=征询句:%s|上下文=%s tok(+%s)|工具=%s 次|预算告警=%s|%s'
% (mode, hit, toks, (toks - prev) if (toks and prev) else '-', ncalls, bool(note),
(tail.replace('\n', ' ')[:60] if tail else '(取不到上一轮文本)')))
if mode == 'inject' and (hit or note):
ctx = ''
if hit:
ctx += CONTEXT
if note:
ctx += ('\n\n' + note) if ctx else note
_emit({'hookSpecificOutput': {'hookEventName': 'UserPromptSubmit',
'additionalContext': CONTEXT}})
'additionalContext': ctx}})
def main():