- V1.0/subskill/ → V1.0/subskills/(browser-harness/mcn-dou-analysis/mcn-script-review/mcn-video-prompt,git 识别 R100 纯重命名保留历史) - 路径说明同步:V1.0/SKILL.md(96/105行)、Lite1.0/SKILL.md(66行)、mcn-dou-analysis/SKILL.md 红线行、操作规范.md 159/175行、mcn-work-shop app.js SKILL_HINT_ACCOUNT - 用户环境 ~/.workbuddy/skills/短视频脚本创作/ 与源仓库 V1.0 同 inode(junction),自动同步无需单独改 - mcn-dou-analysis 内部 subskills/nuwa-skill-main 为正常结构不受影响
1.0 KiB
Screenshots
capture_screenshot() writes a PNG of the current viewport. The file is in device pixels — on a 2× display a 2296×1143 CSS viewport produces a 4592×2286 PNG.
That matters for two reasons:
-
Click coordinates are CSS pixels. Don't read a target off the image and pass it to
click_at_xy()directly without dividing bydevicePixelRatio. The simplest workflow is to take the screenshot, look at it in a viewer that shows CSS coordinates, or measure relative positions and usejs("window.devicePixelRatio")to convert. -
Some LLMs reject images > 2000 px per side. Long sessions on 2× displays will eventually hit this. Pass
max_dim=1800to downscale the file before it gets into the conversation:
capture_screenshot("/tmp/shot.png", max_dim=1800)
The downscale only happens when the image actually exceeds max_dim, so it's safe to leave on for every shot.
Use full-page screenshots (full=True) only when you need to see content below the fold — they are much larger and slower than viewport-only.