Files
mcn-short-video/project/短视频脚本创作/V1.0/subskills/browser-harness/interaction-skills/screenshots.md
T
maogeigei 11fe8d42e2 refactor(skill): 主技能 subskill 目录更名 subskills(09-02 用户要求)
- V1.0/subskill/ → V1.0/subskills/(browser-harness/mcn-dou-analysis/mcn-script-review/mcn-video-prompt,git 识别 R100 纯重命名保留历史)
- 路径说明同步:V1.0/SKILL.md(96/105行)、Lite1.0/SKILL.md(66行)、mcn-dou-analysis/SKILL.md 红线行、操作规范.md 159/175行、mcn-work-shop app.js SKILL_HINT_ACCOUNT
- 用户环境 ~/.workbuddy/skills/短视频脚本创作/ 与源仓库 V1.0 同 inode(junction),自动同步无需单独改
- mcn-dou-analysis 内部 subskills/nuwa-skill-main 为正常结构不受影响
2026-09-02 10:36:23 +08:00

1.0 KiB
Raw Blame History

Screenshots

capture_screenshot() writes a PNG of the current viewport. The file is in device pixels — on a 2× display a 2296×1143 CSS viewport produces a 4592×2286 PNG.

That matters for two reasons:

  1. Click coordinates are CSS pixels. Don't read a target off the image and pass it to click_at_xy() directly without dividing by devicePixelRatio. The simplest workflow is to take the screenshot, look at it in a viewer that shows CSS coordinates, or measure relative positions and use js("window.devicePixelRatio") to convert.

  2. Some LLMs reject images > 2000 px per side. Long sessions on 2× displays will eventually hit this. Pass max_dim=1800 to downscale the file before it gets into the conversation:

capture_screenshot("/tmp/shot.png", max_dim=1800)

The downscale only happens when the image actually exceeds max_dim, so it's safe to leave on for every shot.

Use full-page screenshots (full=True) only when you need to see content below the fold — they are much larger and slower than viewport-only.