重开仓库内容:改推五个技能(browser-harness / humanizer / humanizer-zh / product-planning / session-mechanism)

按授权清空原有内容后重新提交(原 oil-ui-pro 一并移出,可从历史恢复)。
browser-harness 剔除 .venv 等运行环境;根 .gitignore 补记 .venv/ 与 node_modules/。
This commit is contained in:
admin committed 2026-10-07 16:06:02 +08:00
1 parent 9a152a1952
commit 777f7fe5d0
240 files changed
+52518 -4409

No files matched your search

@@ -0,0 +1,48 @@
# Connection & Tab Visibility
## The omnibox popup problem
When Chrome opens fresh, the only CDP `type: "page"` targets are `chrome://inspect` and `chrome://omnibox-popup.top-chrome/` (a 1px invisible viewport). If the daemon attaches to the omnibox popup, all subsequent work — including `new_tab()` and `goto_url()` — happens on tabs that exist in CDP but may not be visible in the Chrome UI.
The daemon's `attach_first_page()` handles this by creating an `about:blank` tab when no real pages exist. If you still end up on an invisible tab, use `switch_tab()` which calls `Target.activateTarget` to bring the tab to front.
## Startup sequence
1. Check if a daemon is already running with `daemon_alive()`
2. If stale sockets exist but daemon is dead, clean them up
3. List open tabs with `list_tabs()` to see what's available
4. `ensure_real_tab()` attaches to a real page
5. `switch_tab(target_id)` both attaches AND activates (brings to front)
```python
if not daemon_alive():
import os, ipc
ipc.cleanup_endpoint("default")
pid = ipc.pid_path("default")
if pid.exists(): pid.unlink()
ensure_daemon()
tabs = list_tabs()
for t in tabs:
print(t["url"][:60])
tab = ensure_real_tab()
```
## Bringing Chrome to front
If Chrome is behind other windows or on another desktop:
```python
import subprocess
subprocess.run(["osascript", "-e", 'tell application "Google Chrome" to activate'])
```
## Navigating
Prefer navigating an existing tab over `new_tab()`. Tabs created via CDP's `Target.createTarget` are visible but may open behind the active tab.
```python
tab = ensure_real_tab()
goto_url("https://example.com")
```
@@ -0,0 +1,3 @@
# Cookies
Document how to get cookies, save cookies, and set cookies without confusing browser state with page state.
@@ -0,0 +1,3 @@
# Cross-Origin Iframes
Focus on `iframe_target(...)`, target attachment, and when compositor-level coordinate clicks are lower-friction than cross-target DOM work.
@@ -0,0 +1,64 @@
# Dialogs
Browser dialogs (`alert`, `confirm`, `prompt`, `beforeunload`) freeze the JS thread. Two approaches depending on timing.
## Detection
`page_info()` auto-surfaces any open dialog: if one is pending it returns `{"dialog": {"type", "message", ...}}` instead of the usual viewport dict (because the page's JS is frozen anyway). So if you call `page_info()` after an action and see a `dialog` key, handle it before doing anything else.
## Reactive: dismiss via CDP (preferred)
Works even when JS is frozen. Handles all dialog types including `beforeunload`.
```python
# Dismiss and read the message
cdp("Page.handleJavaScriptDialog", accept=True) # accept / click OK
cdp("Page.handleJavaScriptDialog", accept=False) # cancel / click Cancel
# Read what the dialog said (from buffered CDP events)
events = drain_events()
for e in events:
if e["method"] == "Page.javascriptDialogOpening":
print(e["params"]["type"]) # "alert", "confirm", "prompt", "beforeunload"
print(e["params"]["message"]) # the dialog text
```
Undetectable by antibot — no JS injected into the page.
## Proactive: stub via JS
Prevents dialogs from ever appearing. Good when you expect multiple `alert()`/`confirm()` calls in sequence.
```python
js("""
window.__dialogs__=[];
window.alert=m=>window.__dialogs__.push(String(m));
window.confirm=m=>{window.__dialogs__.push(String(m));return true;};
window.prompt=(m,d)=>{window.__dialogs__.push(String(m));return d||'';};
""")
# ... do actions that trigger dialogs ...
msgs = js("window.__dialogs__||[]")
```
Tradeoffs:
- Stubs are lost on page navigation -- must re-run the snippet
- `confirm()` always returns `true` (auto-approves)
- Detectable by antibot (`window.alert.toString()` reveals non-native code)
- Does NOT handle `beforeunload`
## beforeunload specifically
Fires when navigating away from a page with unsaved changes (forms, editors, upload pages). The page freezes until the user clicks Leave/Stay.
```python
# Option A: dismiss after navigating (CDP-level, safe)
goto_url("https://new-url.com")
try:
cdp("Page.handleJavaScriptDialog", accept=True) # click "Leave"
except:
pass # no dialog — normal
# Option B: prevent before navigating (JS injection, detectable)
js("window.onbeforeunload=null")
goto_url("https://new-url.com")
```
@@ -0,0 +1,3 @@
# Downloads
Separate browser-triggered downloads from direct `http_get(...)` fetches, and document the minimal signals that prove a download actually started.
@@ -0,0 +1,3 @@
# Drag And Drop
Focus on when drag-and-drop can be driven with low-level input events versus when the site really expects a file upload or DOM-specific drag sequence.
@@ -0,0 +1,3 @@
# Dropdowns
Split dropdowns into native selects, custom overlays, searchable comboboxes, and virtualized menus, and always re-measure after opening because option geometry often appears late.
@@ -0,0 +1,3 @@
# Iframes
Cover same-origin iframe traversal through `contentDocument` / `contentWindow`, and keep the frame-local versus page-coordinate warning explicit for clicks.
@@ -0,0 +1,105 @@
# Make a video
Use captured frames as evidence. Never reenact a finished task or fabricate a
cleaner result.
## Workflow
Use the exact recording selected under `SKILL.md`; never replay browser work to
manufacture missing footage. For a post-task recording, verify `meta.json` and
`events.jsonl` match the task.
```bash
browser-harness video init <recording> --require-explicit
# write <recording>/edit-brief.json
browser-harness video review <recording>
# inspect video-review-contact-sheet.jpg and every image in .privacy-review/
browser-harness video export <recording> --reviewed
```
Omit `--require-explicit` only for a verified post-task recording. In a source
checkout use `./browser-harness`. Never edit generated `composition.js` or
`video.html`; change the brief or shared implementation. Export never
overwrites an existing video, so use `--output video-v2.mp4` for another cut.
## Cut
- Optimize for first-time comprehension; the raw trace is the debugging
artifact. Start with the task and a 2–5 step plan, then end on verified
outcomes.
- Build one causal chain: intent → action → visible result. Remove waits,
retries, and repetition that add no understanding, but show every item or
state explicitly claimed by the outcome.
- Narration is optional and sticky. Set a short present-tense thought only when
it changes, then omit `narration` while 2–3 screenshots advance underneath
it. Text and screenshots should not share a mechanical cadence.
- Preserve representative captured clicks, cursor motion, typing, and result
frames. Pair clicks with `afterEvent`; use `context: true` only to orient.
- Keep a useful wrong turn when it changed the approach. Explain it once as
Observed → Mistake → Correction; remove failures that teach nothing.
- Keep raw frames unlabelled. Subtitles and progress stay outside the app.
Use semantic routes and let the compiler own timing, camera, motion, and
visual style. The 22-second budget and 380 WPM cards are pause-friendly.
## Edit brief
Events are one-based entries in `recording-summary.json`; chapters are
zero-based plan entries. `frameEvent` may select a cleaner pre-action frame and
`afterEvent` should show the click result.
```json
{
"task": "Extract the top five stories and comments",
"summary": "Collect each discussion and save structured JSON.",
"plan": ["Collect stories", "Capture discussions", "Verify JSON"],
"actions": [
{
"event": 3,
"frameEvent": 2,
"afterEvent": 4,
"chapter": 0,
"route": "Hacker News / Front page",
"afterRoute": "Hacker News / Discussion",
"narration": "Open the first discussion.",
"label": "Open discussion"
},
{
"event": 8,
"afterEvent": 9,
"chapter": 1,
"route": "Hacker News / Discussion",
"afterRoute": "Hacker News / Next discussion",
"label": "Continue in rank order"
}
],
"explanations": [{
"afterAction": 2,
"title": "Why the first approach failed",
"observed": "Navigation links appeared in the result",
"mistake": "I selected every page link",
"correction": "Restrict extraction to story rows"
}],
"outcomeTitle": "Five discussions captured",
"outcomeSummary": "The requested JSON is verified.",
"outcomes": ["Five current stories saved", "Comment trees preserved"],
"privacy": {
"reviewedFrames": ["0002.jpg", "0004.jpg", "0008.jpg", "0009.jpg"],
"redact": {"0004.jpg": [{"x": 10, "y": 10, "w": 120, "h": 32}]}
}
}
```
Keep only actions that change the viewer's understanding. Each action requires
`event`, `chapter`, and a short semantic `route`; optional fields are
`frameEvent`, `afterEvent`, `afterRoute`, `narration`, `label`, `detour`,
`error`, `context`, and `showTyping`. Narration is at most seven words; when
omitted, the previous narration persists across the new screenshot.
Explanations reveal Observed → Mistake → Correction. Outcomes must be verified.
Typed text is hidden unless inspected and explicitly enabled with
`showTyping: true`; passwords cannot be revealed. Private app URLs, identities,
credentials, tokens, tenant data, and unrelated people stay private. Use opaque
redaction rectangles in page coordinates and list every used frame in
`privacy.reviewedFrames` only after inspecting its final full-resolution image.
Public task evidence such as authors, post text, and link domains may remain.
The detector is a backstop, not a privacy guarantee.
@@ -0,0 +1,3 @@
# Network Requests
Document how to watch or infer network activity when page state is ambiguous, especially for submit flows, downloads, and SPA actions that succeed without obvious DOM changes.
@@ -0,0 +1,3 @@
# Print As PDF
Cover both direct PDF generation via CDP and sites that only expose a visible "Print" button which must be clicked before handling the browser print flow.
@@ -0,0 +1,90 @@
# Profile sync
Make a remote Browser Use browser start already logged in, by uploading cookies from a local Chrome profile.
## One-time install
```bash
curl -fsSL https://browser-use.com/profile.sh | sh
```
Downloads `profile-use` (macOS / Linux, x64 / arm64). The Python helpers shell out to it; you don't run `profile-use` directly.
## Python API (pre-imported in `browser-harness`)
```python
list_cloud_profiles()
# [{id, name, userId, cookieDomains, lastUsedAt}, ...] — every profile under this API key
list_local_profiles()
# [{BrowserName, ProfileName, DisplayName, ProfilePath, ...}, ...] — detected on this machine
sync_local_profile(profile_name, browser=None,
cloud_profile_id=None, # update an existing cloud profile instead of creating new
include_domains=None, # only these domains (and subdomains); leading dot optional
exclude_domains=None) # drop these domains; applied before include
# Shells out to `profile-use sync`. Returns the cloud profile UUID
# (the existing one if cloud_profile_id was passed, else the newly-created one).
start_remote_daemon("work", profileName="my-work") # name→id resolved client-side
start_remote_daemon("work", profileId="<uuid>") # or pass UUID directly
stop_remote_daemon("work") # shut the daemon and PATCH the cloud browser to stop — billing ends
```
`sync_local_profile` prints `♻️ Using existing cloud profile` when `cloud_profile_id` is accepted, or `📝 Creating remote profile...` → `✓ Profile created: <uuid>` when it creates a new one. Check that line if you want to confirm which path ran.
## Chat-driven flow (don't guess — ask the user)
Cookies are real auth. Don't sync or pick a profile unilaterally.
```python
# 1. Show what's already in the cloud.
for p in list_cloud_profiles():
print(f"{p['name']:25} {len(p['cookieDomains']):3} domains {p['id']}")
```
→ Agent: *"You have these cloud profiles (<N> domains each). Want to reuse one, sync a local profile, or start clean?"*
```python
# 2a. Reuse cloud → one call.
start_remote_daemon("work", profileName="browser-use.com")
# 2b. Sync local first. Show the options:
for lp in list_local_profiles():
print(lp["DisplayName"])
```
→ Agent: *"Which local profile?"* → user picks → before syncing, inspect domain-level cookie counts with `profile-use inspect --profile <name>` (or `--verbose` for individual cookies) and report the summary; never dump 500 cookies into chat.
```python
# 3. Sync + use. Returns the cloud UUID.
uuid = sync_local_profile("browser-use.com")
start_remote_daemon("work", profileId=uuid)
# 3b. Refresh that same cloud profile later (idempotent — no duplicate profiles).
sync_local_profile("browser-use.com", cloud_profile_id=uuid)
# 3c. Scoped: push *only* Stripe cookies into a dedicated cloud profile.
sync_local_profile("browser-use.com",
cloud_profile_id=uuid,
include_domains=["stripe.com"])
```
## What actually gets synced
**Cookies only.** No localStorage, no IndexedDB, no extensions. Enough for session-cookie sites (Google, GitHub, Stripe, most SaaS); not for sites that store auth in localStorage.
Cookies mutated during a remote session only persist on a clean `PATCH /browsers/{id} {"action":"stop"}` — the daemon does this on shutdown when `BU_BROWSER_ID` + `BROWSER_USE_API_KEY` are set (default for remote daemons). Sessions that hit the timeout lose in-session state.
## Cloud profile CRUD
- UI: https://cloud.browser-use.com/settings?tab=profiles
- API: `GET /profiles`, `GET/PATCH/DELETE /profiles/{id}` (paths are relative to `BU_API = "https://api.browser-use.com/api/v3"` in `admin.py`). Fields: `id`, `name`, `userId`, `lastUsedAt`, `cookieDomains[]`. `list_cloud_profiles()` wraps this.
- Name → UUID: `profileName=` on `start_remote_daemon` resolves client-side; no API change needed.
- Need the UUID for an existing profile? `matches = [p["id"] for p in list_cloud_profiles() if p["name"] == "<name>"]` — then verify `len(matches) == 1` before using it. Profile names are not unique; syncs create duplicates unless you pass `cloud_profile_id=`.
- Lower-level raw calls: `from browser_harness.admin import _browser_use; _browser_use("/profiles/<id>", "DELETE")`. Pass the path *without* the `/api/v3` prefix — it's already on `BU_API`.
## Traps
- **Default proxy (`proxyCountryCode="us"`) blocks some destinations** with `ERR_TUNNEL_CONNECTION_FAILED` (e.g. `cloud.browser-use.com` itself). `proxyCountryCode=None` disables the BU proxy; a different country code picks a different exit.
- **Prefer a dedicated work profile over your personal one.** Especially while testing.
- **Older than `profile-use` v1.0.5?** Pre-1.0.5 the sync needed the Chrome profile to be closed (exclusive SQLite lock on the `Cookies` DB). v1.0.5+ copies the profile dir to a temp and syncs from the copy — Chrome can stay open.
@@ -0,0 +1,17 @@
# Screenshots
`capture_screenshot()` writes a PNG of the current viewport. The file is in **device pixels** — on a 2× display a 2296×1143 CSS viewport produces a 4592×2286 PNG.
That matters for two reasons:
1. **Click coordinates are CSS pixels.** Don't read a target off the image and pass it to `click_at_xy()` directly without dividing by `devicePixelRatio`. The simplest workflow is to take the screenshot, look at it in a viewer that shows CSS coordinates, or measure relative positions and use `js("window.devicePixelRatio")` to convert.
2. **Some LLMs reject images > 2000 px per side.** Long sessions on 2× displays will eventually hit this. Pass `max_dim=1800` to downscale the file before it gets into the conversation:
```python
capture_screenshot("/tmp/shot.png", max_dim=1800)
```
The downscale only happens when the image actually exceeds `max_dim`, so it's safe to leave on for every shot.
Use full-page screenshots (`full=True`) only when you need to see content below the fold — they are much larger and slower than viewport-only.
@@ -0,0 +1,3 @@
# Scrolling
Separate page scroll, nested containers, virtualized lists, and dropdown menus, and identify which element is actually consuming wheel events before scrolling.
@@ -0,0 +1,3 @@
# Shadow DOM
Focus on recursive `shadowRoot` traversal, and note when coordinate clicking is simpler than piercing deeply nested component trees.
@@ -0,0 +1,69 @@
# Tabs
Use **CDP for control**, **UI automation for user-visible order**.
## Pure CDP (portable: macOS / Linux / Windows)
```python
tabs = list_tabs() # includes chrome:// pages too
real_tabs = list_tabs(include_chrome=False)
tid = new_tab("https://example.com") # create + attach
switch_tab(tid) # attach harness to tab
cdp("Target.activateTarget", targetId=tid) # show it in Chrome
print(current_tab())
print(page_info())
```
What CDP is good at:
- attach to a tab
- open a tab
- activate a known target
- inspect URL/title/viewport
- capture the attached tab's screenshot even if another tab is visibly frontmost
What CDP is bad at:
- matching the **left-to-right tab strip order** the user sees
- telling whether the attached target is an omnibox popup / internal page without URL filtering
## Visible order (platform UI)
### macOS
```applescript
tell application "Google Chrome"
set out to {}
set i to 1
repeat with t in every tab of front window
set end of out to {tab_index:i, tab_title:(title of t), tab_url:(URL of t)}
set i to i + 1
end repeat
return out
end tell
```
```applescript
tell application "Google Chrome"
set active tab index of front window to 2
activate
end tell
```
### Linux
No AppleScript. Same split still applies:
- use CDP for `new_tab`, attach, inspect, activate known targets
- use window-manager / browser UI automation when the user means visible order
Typical tools:
- `xdotool`
- `wmctrl`
- desktop-environment scripting (`gdbus`, KWin, GNOME Shell extensions, etc.)
## Rules that held up in practice
- `switch_tab()` is **not enough** if the user expects Chrome to visibly change.
- `Target.activateTarget` is the CDP-side "show this tab".
- `list_tabs()` includes `chrome://newtab/` by default; ask for `include_chrome=False` when you want only real pages.
- `chrome://omnibox-popup.top-chrome/` can appear as a fake page target; ignore it for user-facing tab lists.
- If a page has `w=0 h=0`, you may be attached to the wrong target or a non-window surface.
- For dynamic UIs, re-read element rects after opening dropdowns / modals before coordinate-clicking.
@@ -0,0 +1 @@
# Uploads
@@ -0,0 +1,3 @@
# Viewport
Cover how viewport size changes affect layout, coordinate clicks, and any workflow that depends on stable geometry.