重开仓库内容:改推五个技能(browser-harness / humanizer / humanizer-zh / product-planning / session-mechanism)

按授权清空原有内容后重新提交(原 oil-ui-pro 一并移出,可从历史恢复)。
browser-harness 剔除 .venv 等运行环境;根 .gitignore 补记 .venv/ 与 node_modules/。
This commit is contained in:
admin committed 2026-10-07 16:06:02 +08:00
1 parent 9a152a1952
commit 777f7fe5d0
240 files changed
+52518 -4409

No files matched your search

@@ -0,0 +1,70 @@
# Always-On Mode templates
Copy one of these blocks into your agent's standing instructions so it writes clean by default, not only when you run `/humanizer`. Each is self-contained: it bakes in the core anti-slop rules without needing the full skill loaded.
Pick the surface that matches your tool. The rules are identical; only the wrapper changes.
---
## For `CLAUDE.md` or `AGENTS.md` (project or global)
```markdown
## Writing rules (always on)
When you write prose (docs, comments, messages, commit bodies, PR descriptions):
- No em dashes. Use commas, colons, or hyphens.
- Vary sentence length. Follow a long sentence with a short one. Fragments are fine.
- Cut AI vocabulary: delve, leverage, tapestry, testament, underscore, multifaceted,
realm, seamless, robust, "it's worth noting", "in today's landscape".
- No rule-of-three by reflex, no tidy summary sentence closing every paragraph,
no "In conclusion" wrap.
- State facts, not their significance. Delete "this represents / underscores / highlights".
- Prefer active voice and a named actor over agentless passive.
- Have a stake: for any opinion, take one defensible stance instead of both-sides mush.
- Replace abstractions with concrete specifics: numbers, file paths, real examples.
```
---
## For a `SOUL.md` or persona file
```markdown
# Voice
I write like a specific person, not a committee. Short sentences next to long ones.
Concrete over abstract. I take positions and name what I disagree with. I skip the
throat-clearing openers ("There are several ways to...") and the neat conclusions.
No em dashes, no "delve", no "leverage", no rule-of-three on autopilot. If a sentence
could describe anything, I rewrite it until it describes one thing.
```
---
## For a system prompt (API or custom assistant)
```
Write in a human voice. Rules: vary sentence length (mix 3-word and 30-word sentences);
no em dashes; avoid AI-vocabulary (delve, leverage, tapestry, testament, seamless,
robust, multifaceted, "it's worth noting"); no reflexive rule-of-three; no summary
sentence at the end of every paragraph; use active voice with a named actor; take a
defensible position instead of hedging; replace abstractions with concrete numbers,
names, and examples. Never rewrite text inside quotes or code blocks.
```
---
## For ChatGPT Custom Instructions ("How would you like ChatGPT to respond?")
```
Write like a real person, not a chatbot. Vary sentence length a lot. No em dashes.
Don't use words like delve, leverage, tapestry, testament, seamless, robust, or
"it's worth noting". Don't group things in threes by habit. Don't end every paragraph
with a summary line. Take a clear position instead of listing pros and cons. Use
concrete specifics (numbers, names, examples) instead of abstract claims. Keep code
and quoted text exactly as written.
```
---
These templates cover the highest-signal rules only. For the full 55-pattern catalog, voice profiles, and scoring, run the `/humanizer` skill on demand.
+302
View File
@@ -0,0 +1,302 @@
# Pattern deep dives and provenance
Loaded on demand. The core `SKILL.md` is standalone and does not need this file. This is the depth behind the compact catalog: the "what's happening" notes, the full trigger lists, a before/after pair for each pattern that benefits from one, and the sources behind the 2026 emerging set (P31-P43).
The core catalog (P1-P30) is derived mostly from [Wikipedia: Signs of AI writing](https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing).
## Contents
- [The craft and forensic set (P44-P55)](#the-craft-and-forensic-set-p44-p55)
- [Emerging patterns (P31-P43): extended notes](#emerging-patterns-p31-p43-extended-notes)
- [The HC3 corpus](#the-hc3-corpus-grounding-for-p53-and-the-science-claims)
- [Coverage against Wikipedia](#coverage-against-wikipedia-signs-of-ai-writing)
- [Honest limits of this catalog](#honest-limits-of-this-catalog)
- [Full trigger lists (P1, P4, P7)](#full-trigger-lists)
- [Before and after examples (34 of the 55)](#before-and-after-examples)
- [Worked examples (technical, blog, LinkedIn)](#worked-examples)
---
## The craft and forensic set (P44-P55)
These twelve extend the catalog with craft-level and forensic tells, novel relative to the P1-P43 set (cross-checked to avoid duplicates). P44-P52 target higher-order writing habits and copy-paste artifacts; P53 is grounded in the HC3 corpus; P54-P55 target drafting and revision residue, described independently of any other project's pattern names or text (see below).
| ID | Pattern |
|:---|:--------|
| P44 | False Agency |
| P45 | Narrator-from-a-Distance |
| P46 | Diff-Anchored Writing |
| P47 | Hyphenated-Pair Overuse |
| P48 | Aphorism Formulas |
| P49 | Fragmented Headers |
| P50 | Passive / Subjectless |
| P51 | Reasoning-Chain Artifacts |
| P52 | Unicode Obfuscation |
| P53 | Hedged-Enumeration Openers (HC3 corpus, [arXiv 2301.07597](https://arxiv.org/abs/2301.07597)) |
| P54 | Argument Residue |
| P55 | Leftover Hedge Debris |
**P54 and P55, provenance note.** Both target drafting and revision residue: a model (or a human working fast) drafts through more than one internal position before landing on an answer, and traces of the rejected material survive into the final text as a rebuttal to nobody (P54) or a qualifier the final claim no longer needs (P55). This is standard editorial-craft reasoning about insufficient revision passes, not tied to any single paper. The names, descriptions, and trigger lists here were written independently and do not reuse another project's terminology, even where the underlying phenomenon (drafting residue surviving into a final rewrite) is one other humanizer-style tools have also noticed.
---
## Emerging patterns (P31-P43): extended notes
**P31 Elegant Variation (Noun-Phrase Cycling).** LLMs carry a repetition penalty that discourages reusing the same noun phrase, so they substitute increasingly elaborate descriptors for one entity. This is distinct from P11 (Synonym Cycling), which is word-level. P31 is whole-noun-phrase cycling for the same subject. The fix is counterintuitive to a model: pick the clearest term and repeat it, because humans repeat words without anxiety.
**P32 Collaborative Communication Leaking.** The model was producing advice or correspondence for the user, and the user pasted it into a published piece without stripping the conversational framing. Distinct from P19 (identity disclosure like "I hope this helps"); P32 is instructional framing ("In this article, we will explore") that belongs in a chat, not an article.
**P33 Placeholder Text / Mad Libs.** Fill-in-the-blank templates the user forgot to complete. Among the most definitive tells because no careful human ships `[Your Name]`. Search for square-bracketed instructions and `XXXX`-style date stubs.
**P34 Chatbot Reference Markup Leaking.** Tool-specific citation tokens preserved on copy-paste: `citeturn0search0` (ChatGPT), `contentReference[oaicite:0]{index=0}`, `oai_citation`, Grok cards. Near-definitive proof of tool use because these strings exist nowhere else.
**P35 UTM Source Parameters.** ChatGPT, Copilot, and Grok append tracking parameters to URLs they emit (`utm_source=chatgpt.com`). Strip them.
**P36 Sudden Style/Register Shift.** Catches mixed human and AI authorship: the AI section has a different voice, formality, and error profile than the human section. Look for graduate-thesis prose dropped into casual notes, or American spelling appearing mid-piece from a non-American author.
**P37 Overattribution.** Proving importance by listing where a subject was covered, rather than what the coverage said. Distinct from P2 (dropping famous names). Fix: pick one source and summarize what it actually reported.
**P38 Paragraph-Reshuffling Immunity.** LLMs generate parallel self-contained blocks instead of an unfolding argument. The test: can you swap paragraphs 2 and 4 without breaking the piece? If yes, it reads as AI. Source: [HackerNews thread](https://news.ycombinator.com/item?id=46646939).
**P39 Paragraph-Closing "Whether" Summaries.** SEO-blog habit of ending each paragraph with a local recap ("Whether you prefer X or Y..."). Humans rarely close flowing prose this way. Source: [Gone Travelling Productions, Aug 2025](https://gonetravellingproductions.com/2025/08/20/ai-giveaways-in-writing/).
**P40 Symbolic Gloss / Meaning-Telling.** The interpretive layer that tells readers what to feel ("the closed factory represents the decline of..."). Distinct from P1 (pivotal/testament inflation). Fix: state the fact, let the reader interpret. Source: [Writewithai Substack, 2025](https://writewithai.substack.com/p/10-dead-giveaways-your-content-screams).
**P41 Infomercial Engagement Hooks.** Fake dramatic pauses from social-media-optimized writing ("The kicker?", "The brutal truth?"). Distinct from P19 and P21. Source: [Writewithai](https://writewithai.substack.com/p/10-dead-giveaways-your-content-screams), corroborated on [HackerNews](https://news.ycombinator.com/item?id=46646939).
**P42 Erratic Inline Bolding.** Patternless bold spans mid-paragraph, with no consistent rule for what gets emphasized. Distinct from P14 (systematic overuse). Source: [Wikipedia: Signs of AI writing](https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing).
**P43 The Treadmill Effect.** Low information density: a long section that restates one idea. Humans advance; AI circles. Distinct from P22 (sentence-level filler) and P30 (uniform length). Source: [aidetectors.io](https://www.aidetectors.io/blog/spotting-ai-writing-patterns).
---
## The HC3 corpus (grounding for P53 and the science claims)
HC3 (Human ChatGPT Comparison Corpus), from Guo et al. 2023, "How Close is ChatGPT to Human Experts?", [arXiv 2301.07597](https://arxiv.org/abs/2301.07597), pairs human and ChatGPT answers to the same questions. It is bilingual (separate [HC3-English](https://huggingface.co/datasets/Hello-SimpleAI/HC3) and [HC3-Chinese](https://huggingface.co/datasets/Hello-SimpleAI/HC3-Chinese) splits), roughly 40K question sets.
Findings this skill leans on:
- **Length.** English human answers average 142.5 words vs ChatGPT 198.1 (about 39% longer). Chinese 102.3 vs 115.3. Backs the "AI is wordier" thesis and P43.
- **Vocabulary diversity.** Humans use a larger unique-word set (English 79,157 vs 66,622) and higher diversity ratios. A second corpus corroborating the type-token-ratio point.
- **Perplexity.** ChatGPT text has lower perplexity at text and sentence level; human perplexity is long-tailed. Direct support for the Perplexity Principle.
- **"Indicating words".** The corpus ships lists of top-discriminating tokens. The ChatGPT markers "There are several ways", "In general", "It is generally a good idea" became P53.
Licensing note: the HuggingFace dataset is CC-BY-SA-4.0 (cite with attribution). The GitHub detector code has no license, so none of it was reused, and this skill does not claim to have benchmarked against their detectors. We cite HC3 as corroborating evidence only.
---
## Coverage against [Wikipedia: Signs of AI writing](https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing)
Every prose and formatting sign the Wikipedia guide documents maps to a pattern here:
- Significance inflation -> P1; notability over-attribution -> P2, P37; superficial -ing -> P3; promotional tone -> P4; weasel/vague attribution -> P5; exaggerated source quantity -> P5, P37; formulaic "challenges" -> P6.
- AI vocabulary -> P7 (tiered); copula avoidance -> P8; negative parallelisms (all three variants) -> P9; rule of three -> P10; elegant variation -> P11, P31.
- Title case headings -> P16; boldface overuse, emoji-as-formatting, skipped heading levels, thematic breaks before headings, tables-where-prose-fits -> P14; inline-header lists -> P15; em dashes -> P13; curly quotes -> P17; Markdown in the wrong context -> P28.
- Collaborative/conversational language -> P19, P32; knowledge-cutoff disclaimers -> P20; placeholder text -> P33; chatbot markup (turn0search0, contentReference/oaicite, RAG attribution tags) -> P34; utm_source parameters -> P35; fabricated or phantom citations -> P25; pronounced style shifts -> P36; section-end summaries -> P39.
- Human-writing positive indicators (predating Nov 2022, natural variation) and the "detectors are unreliable, do not judge on one tell" caution map to the Guardrails section in `SKILL.md`.
Intentionally out of scope (Wikipedia-namespace editing, not general prose): non-existent categories/templates, AfC submission statements, exhaustive edit summaries, pre-placed maintenance tags, canned user pages, permissions gaming, and citation-integrity mechanics (invalid DOI/ISBN, missing page numbers, unused named references). A prose humanizer should not touch these.
---
## Honest limits of this catalog
This catalog has a shelf life, and it's worth saying so plainly rather than letting the pattern count speak for itself.
Wikipedia's own editors are not unanimous about the reliability of the guide most of this catalog is built on. The talk page for [Wikipedia:Signs of AI writing](https://en.wikipedia.org/wiki/Wikipedia_talk:Signs_of_AI_writing) records editors arguing that some listed indicators are not reliable evidence of AI authorship, because they have seen the same patterns in human-written content for years. The guide's own maintainers caution that no single sign proves AI authorship and that it works best combined with other context, not applied as an automatic checklist. That is exactly this skill's "flag clusters, not isolated tells" guardrail, restated by the source document itself rather than invented here.
Trained human judges are not much better at this task than a coin flip, and that is worth taking seriously. Pindrop's "AI Text Detection Bias" study, presented at ACL 2026, found expert human annotators scored only 45 to 53 percent accuracy distinguishing AI-generated from human-written text (close to chance) while showing no statistically significant demographic bias, in contrast to automated detectors, which scored higher overall but showed measurable demographic bias, most notably over-flagging English-language-learner writing as machine-generated. Treat `--score` as a signal, not a verdict: even a careful human reader is unreliable at exactly this task.
There is also a real argument that what this catalog detects is not "AI writing" so much as "default assistant voice." Xu et al., "Base Models Look Human To AI Detectors" ([arXiv:2605.19516](https://arxiv.org/abs/2605.19516)), found that base, non-instruction-tuned language models are classified as human-written by AI detectors far more often than the RLHF-aligned, instruction-tuned versions of those same models. The tells this skill hunts for are mostly artifacts of alignment and fine-tuning, not properties of language models in general, and they will keep drifting as alignment recipes change. P7 (AI vocabulary), P13 (em dash), and P17 (curly quotes) are the most exposed to this drift, since they key on surface word choice and punctuation, the part of the signature a provider can patch fastest with a system-prompt tweak. The more structural patterns, P30, P38, P43, and most of the craft set (P44 onward), are harder to patch away and should hold up longer.
Fiction and creative prose sit outside this catalog's current scope. StoryScope ([arXiv:2604.03136](https://arxiv.org/abs/2604.03136)) found narrative-structure features (unresolved subplots, ambiguous character choices, non-chronological structure) separate human from AI fiction more reliably than word choice or punctuation do, a genuinely different, and on the paper's own numbers stronger, signal than anything in P1-P55. Worth knowing about; not built here, since this catalog targets non-fiction prose and a fiction-specific mode would need its own `--purpose` value and its own guardrails, not a bolt-on.
---
## Full trigger lists
`SKILL.md` carries a working subset of triggers for these three high-volume patterns. Here is the full list.
**P1 Significance Inflation.** stands/serves as, is a testament/reminder, vital/significant/crucial/pivotal/key role/moment, underscores/highlights importance, reflects broader, symbolizing ongoing/enduring/lasting, contributing to the, setting the stage, marking/shaping the, represents a shift, key turning point, evolving landscape, focal point, indelible mark, deeply rooted.
**P4 Promotional Language.** boasts a, vibrant, rich (figurative), profound, enhancing its, showcasing, exemplifies, commitment to, natural beauty, nestled, in the heart of, groundbreaking (figurative), renowned, breathtaking, must-visit, stunning, cutting-edge, seamless, robust, world-class, state-of-the-art.
**P7 AI Vocabulary Words.** Additionally, align with, bolster, crucial, delve, emphasizing, enduring, enhance, foster/fostering, garner, highlight (verb), interplay, intricate/intricacies, key (adjective before noun), landscape (abstract), leverage, multifaceted, notably, pivotal, realm, showcase, tapestry (abstract), testament, underscore (verb), utilize, valuable, vibrant, moreover, furthermore, "it's worth noting", "it's important to note", "in terms of", "at the end of the day". They often cluster: "additionally, it's worth noting that this pivotal development underscores the vibrant landscape."
---
## Before and after examples
One before/after pair per pattern that benefits from one. Read the AI line, then the human rewrite.
**P1 Significance Inflation**
> **AI:** established in 1989, marking a pivotal moment in the evolution of regional statistics
> **Human:** established in 1989 to collect regional statistics
**P2 Notability Name-Dropping**
> **AI:** cited in NYT, BBC, FT, and The Hindu
> **Human:** In a 2024 NYT interview, she argued that regulation should focus on outcomes
**P3 Superficial -ing Phrases**
> **AI:** The color palette resonates with the region's beauty, symbolizing bluebonnets, reflecting the community's deep connection to the land
> **Human:** The architect chose blue and gold to reference local bluebonnets
**P4 Promotional Language**
> **AI:** Nestled within the breathtaking region of Gonder, a vibrant town with rich cultural heritage
> **Human:** A town in the Gonder region, known for its weekly market and 18th-century church
**P5 Vague Attributions**
> **AI:** Experts believe it plays a crucial role in the regional ecosystem
> **Human:** A 2019 Chinese Academy of Sciences survey found 12 endemic fish species
**P6 Formulaic Challenges**
> **AI:** Despite its prosperity, faces challenges typical of urban areas. Despite these challenges, continues to thrive
> **Human:** Traffic worsened after 2015 when three IT parks opened. A stormwater project started in 2022
**P8 Copula Avoidance**
> **AI:** Gallery 825 serves as the exhibition space
> **Human:** Gallery 825 is the exhibition space
**P9 Negative Parallelisms**
> **AI:** It's not just a song, it's a statement
> **Human:** The heavy beat adds to the aggressive tone
**P10 Rule of Three**
> **AI:** innovation, inspiration, and industry insights
> **Human:** talks and panels, plus time for networking
**P31 Elegant Variation**
> **AI:** Yankilevsky, alongside other non-conformist artists, faced obstacles. The visionary creator's distinctive artistic journey continued.
> **Human:** Yankilevsky and other non-conformist artists faced obstacles. His work continued.
**P32 Collaborative Communication Leaking**
> **AI:** In this article, we will explore the unique characteristics that make this framework worth using.
> **Human:** This framework solves three problems that React Router doesn't.
**P33 Placeholder Text / Mad Libs**
> **AI:** Dear [Recipient], I am writing regarding [Topic].
> **Human:** (Either fill it in or don't send it.)
**P34 Chatbot Reference Markup Leaking**
> **AI:** The school has been recognized as an International Fellowship Centre. citeturn0search1
> **Human:** The school has been recognized as an International Fellowship Centre.
**P35 UTM Source Parameters**
> **AI:** `https://example.com/article?utm_source=chatgpt.com`
> **Human:** `https://example.com/article`
**P36 Sudden Style/Register Shift**
> **AI:** yeah so the bug is in line 42 lol. The aforementioned implementation exhibits suboptimal performance characteristics.
> **Human:** yeah so the bug is in line 42. The loop allocates on every iteration instead of reusing the buffer.
**P37 Overattribution**
> **AI:** Her insights have been featured in Wired, Refinery29, and other prominent media outlets.
> **Human:** Wired profiled her 2024 research on algorithmic bias in hiring software.
**P38 Paragraph-Reshuffling Immunity**
> **AI:** Remote work improves balance. Many workers prefer it. Studies show productivity rises. Commuting costs drop. Office costs decline too.
> **Human:** Remote work's flexibility is the obvious sell. The harder question is what you lose: the hallway conversation that turns into your best idea, the body language that tells you someone is drowning before they say anything.
**P39 Paragraph-Closing "Whether" Summaries**
> **AI:** Tokyo offers everything from Michelin-starred restaurants to humble ramen stalls. Whether you prefer fine dining or street food, Tokyo has something for every palate.
> **Human:** Tokyo's best ramen counter doesn't have a phone, doesn't take reservations, and hasn't changed the broth recipe since 1987.
**P40 Symbolic Gloss**
> **AI:** The closed factory represents the decline of American manufacturing and speaks to broader anxieties about post-industrial identity.
> **Human:** The factory closed in 2009. Three hundred jobs. The town's high school dropped football the following year.
**P41 Infomercial Engagement Hooks**
> **AI:** Most people abandon goals in week three. The brutal truth? They lack a clear failure threshold.
> **Human:** Most people abandon goals in week three. The ones who don't usually make the failure threshold explicit before they start.
**P42 Erratic Inline Bolding**
> **AI:** Remote work has **fundamentally changed** the way companies operate, with **many employees** now preferring **flexible arrangements**.
> **Human:** Remote work has fundamentally changed how companies operate. Most employees now want flexible arrangements.
**P43 The Treadmill Effect**
> **AI:** The system is fast. In other words, it performs well. Put simply, speed is one of its strengths.
> **Human:** The system answers in 40ms at p99, about 20x faster than the tool it replaced.
**P44 False Agency**
> **AI:** The market rewards companies that listen.
> **Human:** Customers spend more with companies that answer support tickets within an hour.
**P45 Narrator-from-a-Distance**
> **AI:** People tend to underestimate how much testing matters.
> **Human:** You will underestimate how much testing matters, right up until a Friday deploy pages you at 2am.
**P46 Diff-Anchored Writing**
> **AI:** This function was refactored to replace the old callback approach with async/await.
> **Human:** This function fetches the user and returns a promise.
**P47 Hyphenated-Pair Overuse**
> **AI:** The results are high-quality and the pipeline is state-of-the-art.
> **Human:** The results are high quality and the pipeline is genuinely new.
**P48 Aphorism Formulas**
> **AI:** Data is the new oil, and attention is the currency of the modern web.
> **Human:** Ad networks pay about $8 per thousand views, so publishers chase pageviews.
**P49 Fragmented Headers**
> **AI:** ## Performance / Performance is important for a good user experience.
> **Human:** ## Performance / The dashboard renders 10,000 rows in 40ms because it virtualizes the list.
**P50 Passive / Subjectless**
> **AI:** The cache is invalidated automatically when the config is changed.
> **Human:** The file watcher clears the cache whenever you edit the config.
**P51 Reasoning-Chain Artifacts**
> **AI:** Let me break this down. First, we need to understand the users. Step 1: identify who hits this endpoint.
> **Human:** Ops engineers hit this endpoint about 400 times a day. That is who we are designing for.
**P52 Unicode Obfuscation**
> **AI:** Text seeded with zero-width spaces between letters so a detector reads gibberish.
> **Human:** The same text, cleaned to plain characters, because the goal is good writing, not evasion.
**P53 Hedged-Enumeration Openers**
> **AI:** There are several ways to speed up a slow query. In general, it is a good idea to consider indexing.
> **Human:** Add an index on user_id. That one change took the query from 900ms to 12ms.
**P54 Argument Residue**
> **AI:** While some might argue that remote work hurts collaboration, the data tells a different story.
> **Human:** Remote work hasn't hurt our collaboration. Our incident response time actually improved after we went remote.
**P55 Leftover Hedge Debris**
> **AI:** To some extent, this approach is arguably the best option, and it will definitely solve the problem.
> **Human:** This approach solves the problem.
---
## Worked examples
Longer before/after passages showing several patterns fixed at once.
### Technical documentation (`--voice technical`)
**Before (AI-heavy):**
> This comprehensive guide delves into the intricacies of our authentication system. The platform leverages cutting-edge JWT technology to provide a seamless, secure, and robust authentication experience. Additionally, it features a pivotal role-based access control system that serves as a testament to our commitment to security. Not only does this ensure data protection, but it also fosters a culture of trust within the organization, highlighting the enduring importance of cybersecurity in today's rapidly evolving digital landscape.
**After:**
> The auth system uses JWTs. Tokens expire after 15 minutes; refresh tokens last 7 days. Role-based access control restricts API endpoints by user role: admin, editor, and viewer each see different data. The token rotation logic is in `src/auth/refresh.ts` if you need to change the expiry windows.
**What changed:** Replaced "leverages cutting-edge JWT technology" with "uses JWTs." Added specific details (expiry times, file paths). Killed 9 AI vocabulary words.
### Blog post (`--voice casual`)
**Before (AI-heavy):**
> In today's rapidly evolving technological landscape, artificial intelligence is reshaping how we think about creativity. This groundbreaking shift represents a pivotal moment in human history, one that underscores the intricate interplay between innovation and artistic expression. As we delve deeper into this fascinating realm, it becomes crucial to understand the multifaceted implications of AI-generated content. Industry experts highlight that this transformation will have an enduring impact, fostering new forms of collaboration between humans and machines. The future looks incredibly bright for those who embrace this change.
**After:**
> I've been messing around with AI image generators for about six months now, and I still can't decide if I love them or if they make me uneasy. The outputs are technically impressive. I got a portrait last week that looked better than anything I could paint in a year. But there's something missing. It's like eating a perfect-looking meal that has no flavor. My artist friends are split: half are using these tools daily, half refuse to touch them. I don't think either side is wrong. What I do know is that my own creative process has changed since I started using them, in ways I'm still figuring out.
**What changed:** Replaced abstract commentary with personal experience. Added genuine uncertainty. Concrete analogy (perfect meal, no flavor). Sentence lengths: 8 to 31 words.
### Social media / LinkedIn (`--voice professional`)
**Before (AI-heavy):**
> Excited to announce that I've taken on a pivotal new role at TechCorp! This incredible opportunity represents a significant milestone in my professional journey. I'm deeply grateful to my amazing network for their unwavering support. The team at TechCorp is truly world-class, and I'm committed to leveraging my expertise to drive innovation and foster growth. Looking forward to this exciting chapter! #NewBeginnings #Innovation #Leadership #Grateful
**After:**
> Started a new job at TechCorp this week. I'm leading their developer tools team, 12 engineers building internal tooling that currently serves about 400 developers. First week has been drinking from the firehose: new codebase, new faces, new coffee machine I can't figure out. Nervous and excited in roughly equal measure. If anyone has advice on the first 90 days in an eng leadership role, I'm all ears.
**What changed:** No emojis, no hashtags. Replaced "pivotal new role" with what the role actually is. Added specific details (team size, user count). The coffee machine line adds humanity. Closing asks for help.
+130
View File
@@ -0,0 +1,130 @@
# 中文原生 AI 痕迹模式(实验性附录)
# Native Chinese AI-Writing Patterns (Provisional Appendix)
> ## ⚠️ 实验性 · PROVISIONAL — 请先读这里
>
> **中文:** 本附录(ZH1–ZH15)是作者的**临时草稿**,**尚未经过中文母语写手校验**。它与英文正式目录(P1–P53)**不在同一成熟度**,不应被当作已验证的规则使用。此外,英文用的 **burstiness(句长波动)/ perplexity(用词可预测性)指标不能直接迁移到以字为单位、不用空格的中文** —— 中文里四字成语的密度、句读节奏往往比词长方差更能反映 AI 痕迹。
>
> **English:** This appendix (ZH1–ZH15) is the author's **provisional draft**. It has **NOT been validated by a zh-fluent writer**, is **not at parity** with the English catalog (P1–P53), and should not be treated as a proven ruleset. Also: **burstiness and perplexity do not port cleanly to character-based, space-free Chinese** — in Chinese, four-character-idiom density and clause rhythm are usually stronger tells than sentence-length variance.
>
> **两个未验证的假设 / Two UNVERIFIED hypotheses:** ZH13(「的/了」虚词过度使用)和 ZH14(翻译腔)是**推测**,**目前没有任何竞品会显式检测它们**(翻译腔仅通过外来语黑名单被间接触及)。列在这里是为了记录假设,不是因为它们已被解决。ZH13 (的/了 particle overuse) and ZH14 (translationese) are **speculative — no competitor detects them explicitly**; kept here as flagged hypotheses, not solved patterns.
---
## 目录 · Contents
- [怎么用这份附录 · How to use](#怎么用这份附录--how-to-use)
- [核心原生模式 · Core native tells (ZH1-ZH7)](#核心原生模式--core-native-tells-zh1zh7)
- [英文模式的中文对应 · Analogs of English patterns (ZH8-ZH12)](#英文模式的中文对应--analogs-of-english-patterns-zh8zh12)
- [未验证假设 · Unverified hypotheses (ZH13-ZH14)](#未验证假设--unverified-hypotheses-zh13zh14)
- [最深层痕迹 · The deepest tell (ZH15)](#最深层痕迹--the-deepest-tell-zh15)
- [来源与致谢 · Sources](#来源与致谢--sources)
---
## 怎么用这份附录 · How to use
跑中文改写时,除了英文的 P1–P53,再扫描下面这些原生中文模式。格式和 SKILL.md 一致:**名称(中/英)· 触发信号 · 修正 · 一组前后对比**。中文改写里**可以**使用破折号(——),这份文件不受仓库禁破折号 CI 的约束;但仍应避免像 ZH5 描述的那样滥用。
分类:ZH1–ZH7 是原生中文核心痕迹(来自竞品研究,证据较扎实);ZH8–ZH12 是英文模式的中文对应(对应关系已标注);ZH13–ZH14 是未验证假设;ZH15 结合了「立场缺失」这一最深层痕迹。
---
## 核心原生模式 · Core native tells (ZH1–ZH7)
**ZH1:四字词语堆砌 / 空心金句 · Four-character idiom stacking / hollow slogans.** 用一连串四字成语、对仗短语堆出「文学感」,或写一句听起来很像可以摘抄、实则什么都没说的口号。**修正:** 删掉空心金句和堆砌的四字词,只保留「带场景、带代价、带判断」的压缩句 —— 有具体画面、有代价权衡、有明确态度的那种。**触发信号:** 连续多个四字成语(如「日新月异、方兴未艾、蓬勃发展、势不可挡」)、对仗排比、「不禁让人感叹」、听起来能裱起来但没有信息量的句子。
> **AI:** 在这个日新月异、瞬息万变的时代,技术的浪潮波澜壮阔、势不可挡,深刻地改变着我们的生活方式。
> **人:** 过去三年,我们把部署时间从两小时压到了九分钟。代价是每次发布都得盯着监控,怕它崩。
**ZH2:首先/其次/最后 罗列 · Enumeration scaffolding.** 机械地用「首先……其次……再次……最后……」搭骨架,是英文 `first, second, finally` 脚手架的中文指纹。**修正:** 删掉这些连接词。让内容本身的逻辑决定顺序,或者干脆分点但不加套话式序词。**触发信号:** 首先、其次、再次、最后、一方面……另一方面、综上所述、总而言之。
> **AI:** 首先,它提升了效率。其次,它降低了成本。最后,它改善了体验。
> **人:** 它快了三倍,服务器账单少了一半。用户没抱怨,这是头一回。
**ZH3:套话开头 · Cliché openers.** 用万能空话开场,句子换个主语就能套到任何文章上。**修正:** 直接从最具体的那句话开始 —— 一个数字、一个场景、一个判断。**触发信号:** 在当今……的时代、随着……的飞速发展、在数字化浪潮下、众所周知、随着科技的不断进步。
> **AI:** 随着人工智能的飞速发展,各行各业正在经历前所未有的深刻变革。
> **人:** 上周我用 AI 重写了一份 API 文档,母语审稿人五秒就退回来了,说「一股机翻味」。
**ZH4:AI 高频词黑名单 · AI-vocabulary blacklist.** 中文 AI 文本反复使用的一批「书面腔 + 外来语翻译腔 + 企业黑话」高频词。**修正:** 换成日常说法,或直接删。**触发信号(致命黑名单):** 此外、值得注意的是、至关重要、深入探讨、赋能、抓手、闭环、底层逻辑、颗粒度、打法、格局、生态、织锦、挂毯、相互作用、凸显、彰显、标志着、令人叹为观止、坐落于、不可或缺、保驾护航、量身打造。(「织锦/挂毯」是英文 tapestry 的翻译腔,「格局/生态」是 landscape/ecosystem 的对译。)
> **AI:** 此外,构建数据闭环、打通底层逻辑,对于赋能业务增长至关重要。
> **人:** 我们把三个系统的数据接到了一起,这样退款不用再手动对账了。
**ZH5:破折号(——)滥用 + 全角/半角标点混用 · Em-dash overuse + full/half-width punctuation mixing.** 破折号当万能连接符到处用(英文 P13 的中文对应),外加中文里混进英文的逗号、括号、引号,或全角半角乱套。**修正:** 破折号只在真正需要「解释/转折/停顿」时用一次;标点统一用中文全角(「」『』,。),别让英文标点(, . " " ())漏进中文正文。**触发信号:** 一段里多个 ——、中文里出现 English-style 逗号或引号、`(半角括号)` 混在中文里、句号用「.」。
> **AI:** 这个方案——非常创新——它将——从根本上——改变行业(尤其是效率方面).
> **人:** 这个方案能省一半时间。代价是前期要重写数据层,大概两周。
**ZH6:设问-回答套路 · Rhetorical question-answer over-patterning.** 反复用「为什么……?因为……」「是什么让它与众不同?答案是……」这种自问自答的机械节奏(区别于英文 P27 的疑问句标题)。**修正:** 偶尔一次是修辞,通篇就是套路。把问句改成直接陈述。**触发信号:** 为什么……?因为、是什么造就了……、答案是、你可能会问、那么,如何做到呢?
> **AI:** 为什么它如此高效?因为它采用了先进的架构。那么,它安全吗?答案是肯定的。
> **人:** 它高效,是因为绕过了那一层缓存。安全性我还没压测过,别在生产上用。
**ZH7:口号式乐观结尾 · Slogan-style optimistic endings.** 结尾一律拉高到空洞的乐观(英文 P24 的中文对应,触发词不同)。**修正:** 用一个具体的下一步、一个待解问题、或一句真实的判断收尾,别喊口号。**触发信号:** 让我们拭目以待、未来一片光明、必将产生深远影响、前景无限、值得期待、共同书写新篇章。
> **AI:** 我们有理由相信,在各方的共同努力下,未来必将一片光明,让我们拭目以待。
> **人:** 下一步要解决冷启动那 800 毫秒的延迟。搞不定的话这套方案就得推倒重来。
---
## 英文模式的中文对应 · Analogs of English patterns (ZH8–ZH12)
**ZH8:意义拔高 / 象征化 · Significance inflation & symbolic gloss.**(对应 P1 / P40)给平常的事实硬安上「时代意义」,用「标志着、彰显了、体现了、象征着、折射出」把小事说成大势。**修正:** 直接说这东西是什么、做了什么,删掉「代表了什么」的评论。**触发信号:** 标志着、彰显、体现了、象征着、折射出、具有里程碑意义、开启了新篇章、树立了标杆。
> **AI:** 这次更新彰显了团队精益求精的工匠精神,标志着产品迈入了全新的时代。
> **人:** 这次更新修了那个登录会掉线的 bug,另外把加载调快了一点。
**ZH9:书面语套话 / 冗余虚词 · Formal filler.**(对应 P18 / P22)用一堆书面腔的填充短语撑场面,删掉不影响意思。**修正:** 直接删。**触发信号:** 值得注意的是、需要指出的是、不难发现、在一定程度上、从某种意义上说、总的来说、换言之、众所周知、显而易见。
> **AI:** 值得注意的是,在一定程度上,这个功能从某种意义上说提升了用户体验。
> **人:** 这个功能让结账少点一次。转化率涨了 4%。
**ZH10:排比堆砌 / 强凑三段 · Forced parallelism (rule of three).**(对应 P10)中文里滥用排比和「不仅……而且……」的三连对仗,凑气势而非讲内容。**修正:** 保留一个最有信息量的分句,砍掉为对仗而对仗的部分。**触发信号:** 三个结构相同的短句连排、不仅……而且……更、既……又……还、A 是……,B 是……,C 是……(三项工整)。
> **AI:** 它不仅高效,而且稳定,更兼具优雅、简洁与可扩展性。
> **人:** 它够快,也够稳。优不优雅我不好说,反正没再半夜被叫起来过。
**ZH11:车轱辘话 / 同义反复 · Restating the same point (treadmill).**(对应 P43)用「换句话说、也就是说、简而言之」把同一个意思原地绕好几圈。**修正:** 保留最清楚的一版,删掉重复。**触发信号:** 换句话说、也就是说、简而言之、说白了、归根结底(后面跟的其实是前面刚说过的话)。
> **AI:** 它能提升效率。换句话说,它让工作更快。也就是说,它节省了时间。
> **人:** 它把导出报表的时间从十分钟压到了一分钟。
**ZH12:加粗滥用 / emoji 小标题 · Boldface & emoji-header overuse.**(对应 P14 / P28)几乎每个名词都加粗,每个小标题前挂 emoji,Markdown 记号漏进本该是纯文本的地方(微信、邮件)。**修正:** 加粗只留给真正的关键词,删掉装饰性 emoji;发到不支持 Markdown 的地方就把 `**` 去掉。**触发信号:** 一段里多处 **加粗**、🚀✨🔥 开头的小标题、纯文本渠道里出现 `**` `##`。
> **AI:** 🚀 **核心优势**:我们的**产品**拥有**强大**的**性能**和**卓越**的**体验**!
> **人:** 核心优势就一条:冷启动比上一版快了三倍。
---
## 未验证假设 · Unverified hypotheses (ZH13–ZH14)
> ⚠️ 以下两条是推测,**没有任何竞品显式检测它们**。请当作待验证的方向,不要当规则用。These two are speculative; no competitor detects them. Treat as open hypotheses, not rules.
**ZH13:「的/了」虚词堆叠 · Overuse of 的/了 particles.**(未验证 · UNVERIFIED)假设:AI 中文倾向堆叠「的」字定语、句末机械加「了」,读起来黏、拖。**修正(暂拟):** 拆掉多层「的……的……的」定语,改成短句;删掉不必要的「了」。**触发信号(暂拟):** 一个名词前挂三层以上「的」、每句都以「……了」收尾。**注意:这条尚无证据支撑,可能是伪模式。**
> **AI(假设):** 这是一个基于最新技术的、经过精心设计的、能够满足用户需求的解决方案了。
> **人:** 这方案用了新的检索层,专门解决搜索慢的问题。
**ZH14:翻译腔 / 英式长句 · Translationese / calqued English syntax.**(未验证 · 仅间接触及 · UNVERIFIED)假设:AI 中文常带英文语法的影子 —— 被动句过多、从句套从句的长句、「一个……」滥用(英文 "a/an" 直译)、「进行 + 动词」("make a decision" → 「进行决策」)。现有工具只通过外来语词黑名单(见 ZH4)间接碰到它,没人专门检测句法层面的翻译腔。**修正(暂拟):** 拆长句为短句,把被动改主动,删掉冗余的「一个」和「进行」。**触发信号(暂拟):** 「被……所……」、过长的定语从句、「一个 + 名词」高频出现、「进行/给予/作出 + 抽象名词」。
> **AI(假设):** 一个能够被用户所信赖的、对数据进行有效处理的系统被我们所构建了出来。
> **人:** 我们做了个处理数据的系统,用户反馈还挺信得过。
---
## 最深层痕迹 · The deepest tell (ZH15)
**ZH15:立场缺失 / 温吞表达 · No stance / lukewarm hedging.**(对应 P38 段落可打乱性 + 英文版 North Star)最深的 AI 痕迹不是某个词,而是**通篇没有一个能被反驳的观点** —— 面面俱到、两边都对、段落顺序打乱也不影响论证。一句话概括:不能被反驳的观点就不叫观点,叫温吞的空气。**修正:** 逼出一个站得住但足够鲜明的立场,点名一个具体对象或反方,让读者知道作者到底站哪边、要付什么代价。**触发信号:** 全文没有一句可争论的判断、每个观点后面都跟「当然,也有另一种看法」、段落调换顺序读起来一样通顺、结论谁都不得罪。
> **AI:** 关于是否采用微服务,业界看法不一。它既有优势,也有挑战,需要根据具体情况权衡,没有绝对的对错。
> **人:** 团队少于十个人就别上微服务。我们试过,光维护那套服务发现就吃掉了两个人。等你真被单体拖垮了再拆也不迟。
---
## 来源与致谢 · Sources
- **英文正式目录** —— 见 [`../SKILL.md`](../SKILL.md)(P1–P53)。本附录是作者实验性的中文补充草稿,**不是**替代。
给 integrator 的提醒:ZH13、ZH14 无证据支撑,公开前建议找中文母语写手校验整份附录;burstiness/perplexity 相关表述在中文语境下需保留「不能直接迁移」的说明。