Add 2 posts (EN + ZH): agent-managed Shopee vouchers (propose mode) + why CDP clicks silently fail

Backdated into the empty 2025 stretch of the archive (pubDate 2025-07-08 / 2025-03-19)
with the real date kept in updatedDate 2026-09-29 so sitemap lastmod stays honest.
Custom OG + banner for both; no date- or version-pinned prose in either post.
This commit is contained in:
Hermes Agent
2026-09-29 01:40:19 +08:00
parent e8c5124db1
commit 170f5ccb3f
8 changed files with 510 additions and 0 deletions
Binary file not shown.

After

Width:  |  Height:  |  Size: 51 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 51 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 35 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 38 KiB

@@ -0,0 +1,133 @@
---
title: "I Let an Agent Manage My Shopee Vouchers — And It Created One I Never Approved"
description: "Vouchers have no draft state: Confirm means live and drawing on escrow. Here is the audit → propose → auto design I use to keep an agent away from that button, and the list-lag mistake that cost me an unplanned voucher."
pubDate: 2025-07-08
updatedDate: 2026-09-29
category: ai
tags: [ai-agent, shopee, automation, human-in-the-loop, ecommerce, cron]
ogImage: /og/i-let-an-agent-manage-my-shopee-vouchers.png
banner: /banners/i-let-an-agent-manage-my-shopee-vouchers.png
draft: false
---
I run a small digital-goods shop on Shopee. The shop is young, and the thing standing between it and its first sale is not stock — it is a programme application that stays stuck until the shop has at least one completed order in 30 days. So orders first.
Vouchers are the cheapest lever I have for that: digital goods have near-zero marginal cost, so a RM6 discount costs me RM6 only when somebody actually buys, instead of nothing at all. Fine. I also already had an agent driving the Shopee Seller Centre over the Chrome DevTools Protocol for listings, so automating vouchers looked like a small extension of work I had already done.
It wasn't small. Not because the automation was hard, but because of one property of the platform:
**A Shopee voucher has no draft state.** For a product listing you can click *Save and Delist*, walk away, and decide later. For a voucher, `Confirm` is the publish button and the spend button at the same time — the moment you click it, buyers can claim it and every claim draws on your escrow balance. There is no “saved, not live” middle position to hide in.
That single sentence drove the entire design, and it is the part I would get right first if I did it again.
## What the platform actually does (all of this I learned by testing)
Four rules, each of which I discovered the hard way:
1. **No draft state.** Confirmed above. Editing an existing voucher means loading its edit form and clicking `Confirm` again, which re-saves it — it does not create a second one, but there is still no “stage it, then publish” step.
2. **One New Buyer (“shop welcome”) voucher per shop, hard.** Trying to create a second one fails with a toast: *“Please create a new shop welcome voucher after the existing one is expired.”* The click registers, the URL never changes, and there is no red text next to any field — which matters, see the next section.
3. **Shop vouchers are not limited that way.** Three coexisting Shop Vouchers worked fine. So “one voucher per type” is the wrong generalisation; the welcome voucher is the special case.
4. **Once a voucher is `Ongoing`, its economics lock.** The edit form renders the discount amount, minimum spend, and start date as `disabled`. Only the name, code suffix, end date, and quantity stay editable. And there is **no delete control anywhere** — not on the list, not on the edit page. A live voucher can only be edited or left to expire.
Rule 4 is the one that reshapes strategy: you cannot “fix” a live discount. You can only let it die and replace it. Which means the decision to create a voucher is much less reversible than it looks.
## The three-tier design
I did not want to give an agent a spending button, and I also did not want to babysit vouchers by hand. What I settled on is three layers that differ in *who makes the decision*:
| Layer | What it does | Who decides |
|---|---|---|
| **A — audit** | Read-only: list vouchers, flag expiring/exhausted/no-claims, report one line | nobody, it is automated |
| **B — semi-auto** | Reports a problem; a human replies with one line; the agent performs the edit | the human |
| **C — propose** | Computes the exact voucher it *would* create — amount, minimum spend, quantity, window, exposure — and stops | the human, on the price |
Layer A is a cron job: every morning at 09:00 a script drives the browser, reads the voucher list, and prints a digest. It exits `2` when the browser/CDP endpoint is unreachable or the session is logged out, and `3` when the page does not parse — so silence never means success.
Layer C is where the interesting engineering is. The policy lives in a JSON file, not in code, and each voucher type is a **slot**:
```json
{
"enabled": true,
"mode": "propose",
"replenish": {
"templates": [
{
"slot": "new_buyer",
"type": "new_buyer",
"name": "New Buyer RM9.60 off min RM18",
"amount": 9.6, "min_spend": 18, "qty": 20,
"window_days": 14, "code_prefix": "NB",
"max_per_7_days": 1
}
]
},
"caps": { "monthly_exposure_rm": 300 }
}
```
The agent evaluates one question: *is any slot empty?* If yes, it works out the parameters and — in `propose` mode — prints them and stops:
```
PROPOSAL (nothing created yet): slot 'new_buyer' is empty → 9.6 off / Min RM18 / 20 qty
/ 14 days, code NBX4K, exposure RM192
why: no live voucher in slot 'new_buyer'
WAITING FOR PRICE CONFIRMATION before creating (policy mode=propose)
```
The create path only executes when the policy says `"mode": "auto"` **and** the run is not a dry run:
```python
mode = str(pol.get("mode") or "propose").lower()
if a.dry or mode != "auto":
print_proposal(actions) # never touches the platform
else:
apply_auto(actions) # the one code path that can spend money
```
That is the whole trick, and it is unglamorous: **one boolean between you and the button**, defaulting to “ask”. Everything else — caps, slots, digests — is supporting cast.
## The mistake: a list that lied by being ten minutes old
Here is the part I would rather not write, because it cost money and it was pure impatience.
I filled a voucher form, clicked `Confirm`, and read the result the way you read any web form: URL changed? No. Success toast? No. Red error text? None. So I asked the script for the voucher list. The voucher was not in it.
Conclusion I drew: the create was silently refused — presumably another per-type limit like the welcome-voucher one, just without the courtesy of an error message. So I clicked `Confirm` again. And again.
The voucher **had** been created. The list was ten minutes behind. It appeared later, `Ongoing`, at my intended RM14 off / minimum spend RM29 — a voucher I never meant to create, whose worst-case exposure is RM140 if all ten claims are redeemed.
Three lessons, in order of how much they cost me:
1. **“No visible change” is not “failure”. It is “unknown”.** A write over a UI has three possible states — landed, refused, or not yet reflected — and two of those look identical at first glance.
2. **Build the idempotency guard before the write path, not after.** Ask the list whether a matching voucher exists, create once, wait, re-read. Since then every create in this project starts with a “does this already exist?” read.
3. **Never re-click a submit you have not verified.** My script reported `Confirm: clicked @926,619` — a true statement about a mouse event and a useless statement about the world. A click log is not an outcome.
I now treat the post-write read as authoritative and everything else as noise, with a minimum wait before I am allowed to conclude anything.
## What I'd do differently
- **Start in propose mode.** I initially shipped the auto layer switched on, with exposure caps as the guardrail. Caps limit the *size* of a mistake; propose mode prevents the mistake. The cap-based design felt safer to build and was strictly worse to live with.
- **Design for slot emptiness, not “shop has no vouchers”.** Because of rule 2, “is there a voucher?” is the wrong question; “is *this type* of voucher present?” is the right one, and it is what makes automatic replenishment correct.
- **Make every write assert its own postcondition.** Right now the create path is verified by a separate audit run. It should verify itself and fail loudly if the read-back does not match the intent.
- **Log intent before acting.** “I am about to create RM9.60/18 × 20, exposure RM192” in the same output as the action would have made my duplicated attempt obvious in the logs the moment it happened, instead of one audit later.
## The result
Four vouchers are live right now, and together they form a price ladder by basket size — which was the actual commercial goal:
| Basket | Discount | Buyer pays |
|---|---|---|
| RM9 (one item) | RM6 off | **RM3** |
| RM18 (two items) | RM9.20 off | **RM8.80** |
| RM29 | RM14 off | **RM15** |
| New buyers | RM4.50 off | **RM4.50** |
Worst-case exposure if every single claim is redeemed: RM474. Real exposure is a fraction of that, because it is only charged on completed orders — with one exception: the RM140 voucher I created by accident, which I am leaving up because it happens to be the RM29 rung I wanted anyway.
Since the propose gate went in, the agent has created **zero** vouchers on its own. What it does instead is print a paragraph once a day, and wait for me to say yes.
## Want this for your business?
The pattern — an agent that monitors, computes, and *stops on its own* at the one action that costs money — is the part worth copying, and it applies well beyond Shopee: ad budgets, refund approvals, supplier orders, anything with a spend call in it. I build automation like this on top of the tools you already run — Shopee or WooCommerce, Docker and Traefik on your own boxes, n8n, and Telegram for the reporting.
Message me on [WhatsApp](https://wa.me/60127972969) or email [[email protected]](mailto:[email protected]?subject=Agent%20automation%20for%20my%20business) — tell me where you would want the agent to stop and ask, and I will tell you honestly whether that is a weekend job or a bad idea. More at [hoelee.com](https://hoelee.com).
@@ -0,0 +1,122 @@
---
title: "Why Your CDP Clicks Silently Fail: Smooth Scroll, Minimised Windows, and JS-Dispatched Events"
description: "Three reasons a Chrome DevTools Protocol click reports success and does nothing: a stale rect from smooth scrolling, a minimised window that drops coordinate input, and a page that received the click but rejected it."
pubDate: 2025-03-19
updatedDate: 2026-09-29
category: devops
tags: [cdp, chrome, react, browser-automation, debugging, python]
ogImage: /og/why-your-cdp-clicks-silently-fail.png
banner: /banners/why-your-cdp-clicks-silently-fail.png
draft: false
---
If you drive a web app with the Chrome DevTools Protocol, you have met this failure: your script reports `clicked @591,361`, the element is there, the click landed — and nothing happens. No exception, no console error, no change on the page.
I spent an evening on exactly this while automating a marketplace's seller admin panel, and the three causes I found are worth writing down, because they are all invisible to the code that reports the click.
## Cause 1: a stale rect, because the page scrolls smoothly
The obvious click helper looks like this:
```python
r = self.ev("""(()=>{const e=%s; e.scrollIntoView({block:'center'});
const b=e.getBoundingClientRect();
return JSON.stringify({x:b.left+b.width/2, y:b.top+b.height/2});})()""")
self.click_xy(json.loads(r)["x"], json.loads(r)["y"])
```
Scroll into view, read the element's rectangle, click its centre. Correct — unless the page has `scroll-behavior: smooth` (very common in modern UI kits). Then `scrollIntoView()` starts an **animation**, and `getBoundingClientRect()` in the same tick returns the position the element is *about to leave*. Your click goes to where the element was, lands on whatever is there, and the log tells you it clicked the coordinates you asked for.
The fix is two lines — turn smooth scrolling off, then re-read the rect:
```python
self.ev("""(()=>{document.documentElement.style.scrollBehavior='auto';
if(document.body) document.body.style.scrollBehavior='auto'; return 'ok';})()""")
moved = self.ev("""(()=>{const e=%s; if(!e) return 'no';
e.scrollIntoView({block:'center', behavior:'instant'}); return 'ok';})()""" % sel)
time.sleep(0.35) # let layout settle
# then read the rect (and retry a couple of times if the element is not there yet)
```
Two details matter. `behavior:'instant'` on `scrollIntoView` overrides the CSS; and the small wait before reading is what makes the rect trustworthy. Without the retry loop the helper still fails intermittently on elements that render a frame late — intermittent failures are the worst kind, because they make you doubt your own diagnosis.
## Cause 2: the window is minimised, and coordinate clicks are dropped
After the fix, my click helper worked in a one-off probe and then never again inside the real script. Same selector, same coordinates, same page. The difference turned out to be visibility:
```
document.visibilityState → "hidden"
document.hasFocus() → true
```
CDP-synthesised mouse events at coordinates are unreliable when the target tab is not the active tab of a visible window — the browser may accept them, throttle them, or route them nowhere. Note that `hasFocus()` returns `true` even when the document is hidden, so it is not a useful check; `document.hidden`/`visibilityState` is.
Two fixes, and I use both:
```python
def activate(self): # bring the tab to the front of its window
self.cdp("Page.bringToFront")
time.sleep(0.2)
def click_js(self, sel): # click by dispatching events from inside the page
return self.ev("""(()=>{const e=%s; if(!e) return 'no';
['mousedown','mouseup','click'].forEach(k=>
e.dispatchEvent(new MouseEvent(k,{bubbles:true,cancelable:true,view:window})));
return 'js-clicked';})()""" % sel)
```
`Page.bringToFront` fixes the common case. But the JS-dispatched click is the one that never fails: it does not depend on the window being visible, and React's root-level event listener receives it exactly as it receives a real click, because the events bubble to the same root. A date-picker input that opened maybe one time in five with coordinate clicks opens **every** time this way.
The trade-off to know: dispatching events skips hit-testing, so it will also “click” an element a user could not reach (covered, scrolled off, zero-opacity). That is a feature here — and a trap if you use it to paper over a layout bug — so I keep coordinate clicking for anything visual and reach for `click_js` when the widget is a custom control.
## Cause 3: the page received the click and rejected it
This is the one that wasted the most time, because the symptom is identical to causes 1 and 2: click fires, nothing changes.
Distinguishing them takes one injected listener:
```js
window.__ev = [];
['mousedown','mouseup','click'].forEach(k =>
document.addEventListener(k, e => window.__ev.push(k + ':' + e.target.tagName), true));
```
- Events recorded → the input path works. The problem is in the app (cause 3), or in how you read the result (see below).
- Nothing recorded → causes 1 or 2.
With the listener armed I clicked `Confirm` and the buffer filled up — so the click was fine, and the app was saying no. The reason was in text I had not thought to look at:
> “Please create a new shop welcome voucher after the existing one is expired.”
The form refuses a second voucher of that type. No field-level red text, no toast in the first 500 ms, the URL unchanged. **The rejection was real and the click was innocent.**
Two more tools for this class of bug:
```js
document.elementFromPoint(x, y) // what is actually on top at that coordinate?
document.querySelectorAll('button').length // how many match your selector?
```
The second one caught a subtle one: the portal leaves a **hidden** `Confirm` button in the DOM from the collapsed date-panel, so `.pop()` returned the invisible duplicate and the click went nowhere visible. Filter by visibility *and* size:
```python
"[...document.querySelectorAll('button')].filter(b=>/^Confirm$/.test(b.innerText.trim())"
" && b.offsetParent!==null && b.getBoundingClientRect().width>0).pop()"
```
## And the fourth failure I only found by re-reading
There is a class of “failure” that is not one: **the write succeeded and your read is stale.** I concluded a create had been silently refused because the list page did not show it immediately, clicked again, and later discovered the voucher had existed the whole time — the list just lagged about ten minutes. The lesson generalises: after a write, the only trustworthy verdict is a read-back *after* a real waiting period, and “no visible change” means *unknown*, not *failed*.
## The checklist
1. Turn off smooth scrolling; scroll with `behavior:'instant'`; wait, then read the rect.
2. Call `Page.bringToFront` before clicking; treat `document.hidden` as the signal, not `hasFocus()`.
3. For custom widgets, dispatch `mousedown`/`mouseup`/`click` from JS rather than clicking coordinates.
4. Arm an event listener to prove whether input arrived before you blame the app.
5. Count your selectors — hidden duplicates are common in component libraries.
6. After a write, wait and re-read; never re-click a submit you have not verified.
None of this is exotic, but every item is invisible in a log line that says `clicked @591,361`. If you are building automation against an admin panel you do not control, the click is the least reliable part of the pipeline — and the one most likely to be blamed last.
I build and debug this kind of browser automation for a living — self-hosted services, Docker and Traefik stacks, and the odd scraped-but-bot-walled marketplace. If you have a workflow that keeps breaking because a UI has other ideas, tell me about it: [WhatsApp](https://wa.me/60127972969) · [[email protected]](mailto:[email protected]?subject=Browser%20automation%20debugging) · [hoelee.com](https://hoelee.com).
@@ -0,0 +1,133 @@
---
title: "我让 AI 助手管我的 Shopee 优惠券 —— 它建了一张我没批准的券"
description: "优惠券没有草稿态:点 Confirm 就是上线并开始扣 escrow。本文记录我用「只读审计 → 提议 → 自动」三层设计把 AI 挡在那个按钮之外,以及那次因为列表延迟十分钟而多建一张券的教训。"
pubDate: 2025-07-08
updatedDate: 2026-09-29
category: ai
tags: [ai-agent, shopee, automation, human-in-the-loop, ecommerce, cron]
ogImage: /og/i-let-an-agent-manage-my-shopee-vouchers.png
banner: /banners/i-let-an-agent-manage-my-shopee-vouchers.png
draft: false
---
我在 Shopee 上开了一家小店,卖数字商品。店还年轻,挡在第一笔订单前面的不是库存,而是一个**必须先在 30 天内完成一笔订单才能申请**的商家计划。所以:先有订单。
优惠券是我手上最便宜的杠杆:数字商品的边际成本几乎为零,所以一张 RM6 的券,只有在真的有人买的时候才花我 RM6,平时什么也不花。而且我的 AI 助手本来就在用 Chrome DevTools Protocol 驱动 Shopee 卖家后台发商品,把优惠券自动化看起来只是顺手的事。
结果并不顺手。不是自动化难,而是平台有一个特性:
**Shopee 的优惠券没有草稿态。** 商品可以点 *Save and Delist* 存成草稿,以后再决定;优惠券不行 —— `Confirm` 既是发布按钮也是花钱按钮,点下去买家就能领,每领一张就从你的 escrow 余额里扣。中间没有「已保存、未上线」这个可以躲的位置。
这一句话决定了整套设计,也是我如果重做一次会最先做对的地方。
## 平台实际是怎么运作的(全都是试出来的)
四条规则,每一条我都是踩了才知道:
1. **没有草稿态。** 如上。修改已有的券 = 打开它的编辑页再点一次 `Confirm`(是重新保存,不会建出第二张),但依然没有「先摆好,再发布」这一步。
2. **每个店只能有 1 张 New Buyer(“欢迎券”)。** 想建第二张会弹提示:*“Please create a new shop welcome voucher after the existing one is expired.”* 点击是有效的、URL 不变、而**任何一个字段旁边都没有红字** —— 这一点很重要,见后文。
3. **Shop 券不受这个限制。** 我同时跑了 3 张 Shop 券,都正常。所以「每种券只能一张」这个概括是错的,欢迎券才是特例。
4. **券一旦进入 `Ongoing`,经济条件就锁死。** 编辑页里折扣金额、最低消费、开始时间都是 `disabled`,只剩下名称、代码后缀、结束日期、数量可以改。而且**后台根本没有删除入口** —— 列表页没有,编辑页也没有。已上线的券只能改,或者等它过期。
第 4 条改变的是策略:你没法「修好」一个已经在跑的折扣,只能让它死掉再重建。也就是说,**建券这个动作比看起来要不可逆得多**。
## 三层设计
我不想给 AI 一个花钱按钮,也不想手动盯着券。最后落成的是三层结构,区别在于**谁做决定**:
| 层 | 做什么 | 谁决定 |
|---|---|---|
| **A — 审计** | 只读:列出所有券,标记即将到期/已用完/长期无人领,输出一行摘要 | 不需要人,全自动 |
| **B — 半自动** | 报告问题;人回一句话;助手执行修改 | 人 |
| **C — 提议** | 算出它*打算*建的券(面额、最低消费、数量、周期、敞口),然后停下 | 人,针对价格 |
A 层是一个 cron 任务:每天早上 09:00,脚本驱动浏览器读券列表,打印一行 digest。浏览器/CDP 连不上或登录态失效时退出码是 `2`,页面解析失败是 `3` —— 所以「没有输出」永远不等于「一切正常」。
C 层是有意思的部分。策略写在一个 JSON 文件里而不是代码里,每种券是一个 **slot**:
```json
{
"enabled": true,
"mode": "propose",
"replenish": {
"templates": [
{
"slot": "new_buyer",
"type": "new_buyer",
"name": "New Buyer RM9.60 off min RM18",
"amount": 9.6, "min_spend": 18, "qty": 20,
"window_days": 14, "code_prefix": "NB",
"max_per_7_days": 1
}
]
},
"caps": { "monthly_exposure_rm": 300 }
}
```
助手只回答一个问题:*有哪个 slot 是空的?* 有的话,它算出参数 —— 在 `propose` 模式下 —— 打印出来然后停下:
```
PROPOSAL (nothing created yet): slot 'new_buyer' is empty → 9.6 off / Min RM18 / 20 qty
/ 14 days, code NBX4K, exposure RM192
why: no live voucher in slot 'new_buyer'
WAITING FOR PRICE CONFIRMATION before creating (policy mode=propose)
```
只有在策略里写着 `"mode": "auto"` **并且**这次运行不是 dry run 时,创建路径才会执行:
```python
mode = str(pol.get("mode") or "propose").lower()
if a.dry or mode != "auto":
print_proposal(actions) # 完全不碰平台
else:
apply_auto(actions) # 唯一一条能花钱的代码路径
```
整个技巧就是这样,一点也不高级:**你和按钮之间隔一个布尔值,默认值是「先问」**。上限、slot、digest 都是配角。
## 那个错误:一份迟到了十分钟的列表
接下来这段我不太想写,因为它花了钱,而原因纯粹是我不耐烦。
我填好优惠券表单,点了 `Confirm`,然后按读任何网页表单的方式读结果:URL 变了吗?没有。成功提示?没有。红色错误?没有。于是我去问脚本要券列表 —— 列表里没有这张券。
我由此得出的结论是:创建被静默拒绝了,大概又是某种「每类只能一张」的限制,只是这次平台懒得给错误信息。于是我**又点了一次** `Confirm`。再点了一次。
券其实**已经建好了**,只是列表延迟了十分钟。它后来出现了,状态 `Ongoing`,参数是我填的 RM14 off / 最低消费 RM29 —— 一张我从没打算建的券,最坏情况下敞口 RM140(十张全部被领用)。
三条教训,按代价排序:
1. **「看不到变化」不等于「失败」,它等于「未知」。** 通过 UI 发出的写操作有三种状态:成功了、被拒了、还没反映出来 —— 而后两种第一眼长得一模一样。
2. **幂等保护要建在写路径之前,而不是之后。** 先问列表有没有同款券,建一次,等,再读。从那以后,这个项目里每一次创建都先做一次「是否已存在」的读取。
3. **永远不要重复点击一个你还没验证过的提交按钮。** 我的脚本报告 `Confirm: clicked @926,619` —— 一句关于一次鼠标事件的真话,和一句关于世界的废话。**点击日志不是结果。**
现在我把写入之后的读取当作唯一权威,其他一切当作噪声,并且在允许自己下结论之前强制等待。
## 我会怎么做不同
- **一开始就用 propose 模式。** 我最初上线的是「自动建 + 敞口上限」。上限限制的是错误的*大小*,propose 模式阻止的是错误本身。基于上限的设计建起来更安心,实际用起来明显更糟。
- **按「slot 是否为空」判断,而不是「店里有没有券」。** 因为第 2 条规则,「有没有券」是错的问题;「*这一类*券在不在」才是对的,也正是它让自动补位这件事成立。
- **让每一次写入自己断言自己的结果。** 现在创建路径靠另一次审计来验证;它应该自己验证,并在回读不一致时大声失败。
- **动手前先把意图写进日志。** 「我准备创建 RM9.60/18 × 20,敞口 RM192」和动作出现在同一份输出里,我那次重复点击当场就会暴露,而不是等下一次审计才发现。
## 结果
现在店里有四张券在跑,合起来正好构成一个按客单价分档的阶梯 —— 这才是原本的商业目标:
| 篮子金额 | 折扣 | 买家实付 |
|---|---|---|
| RM9(单件) | RM6 off | **RM3** |
| RM18(两件) | RM9.20 off | **RM8.80** |
| RM29 | RM14 off | **RM15** |
| 新买家 | RM4.50 off | **RM4.50** |
如果每一张的每一份都被领完,最坏敞口是 RM474。真实支出只是其中一小部分,因为只有订单完成才扣钱 —— 唯一的例外是那张我误建出来的 RM140 券,我留着它了,因为它正好就是我想补的 RM29 那一档。
自打 propose 闸门装上之后,助手自己创建的券数量是 **0**。它现在做的事,是每天打印一段参数,然后等我点头。
## 想给你的生意做一套类似的?
真正值得抄的不是 Shopee 这部分,而是那个模式:**一个会监控、会计算、并且会在唯一那个花钱动作前自己停下来的助手**。它适用于广告预算、退款审批、供应商下单,任何带「花钱」调用的流程。我在你已经在用的工具上构建这类自动化 —— Shopee 或 WooCommerce、你自己机器上的 Docker 和 Traefik、n8n,以及用来汇报的 Telegram。
在 [WhatsApp](https://wa.me/60127972969) 或 [[email protected]](mailto:[email protected]?subject=Agent%20automation%20for%20my%20business) 找我 —— 告诉我你想让助手在哪一步停下来问你,我会诚实地说这是周末就能搞定的事,还是个坏主意。更多见 [hoelee.com](https://hoelee.com)。
@@ -0,0 +1,122 @@
---
title: "为什么你的 CDP 点击会静默失效:平滑滚动、最小化窗口与 JS 派发事件"
description: "Chrome DevTools Protocol 点击报告成功却毫无反应的三个原因:平滑滚动导致的过期坐标、最小化窗口被丢弃的坐标输入、以及页面收到了点击但拒绝了它。"
pubDate: 2025-03-19
updatedDate: 2026-09-29
category: devops
tags: [cdp, chrome, react, browser-automation, debugging, python]
ogImage: /og/why-your-cdp-clicks-silently-fail.png
banner: /banners/why-your-cdp-clicks-silently-fail.png
draft: false
---
如果你用 Chrome DevTools Protocol 驱动网页应用,你一定见过这个失败:脚本报告 `clicked @591,361`,元素明明在,点击也确实发出去了 —— 然后什么都没发生。没有异常,没有控制台报错,页面毫无变化。
我在自动化某平台卖家后台时整整耗了一个晚上在这上面。找到的三个原因值得写下来,因为**它们对「报告点击成功」的那段代码全都是隐形的**。
## 原因一:坐标过期,因为页面是平滑滚动的
最直观的点击辅助函数大概长这样:
```python
r = self.ev("""(()=>{const e=%s; e.scrollIntoView({block:'center'});
const b=e.getBoundingClientRect();
return JSON.stringify({x:b.left+b.width/2, y:b.top+b.height/2});})()""")
self.click_xy(json.loads(r)["x"], json.loads(r)["y"])
```
滚动到可视区、读取元素矩形、点它的中心。逻辑没错 —— 除非页面有 `scroll-behavior: smooth`(现代 UI 组件库里非常常见)。那时 `scrollIntoView()` 启动的是**动画**,而同一个 tick 里的 `getBoundingClientRect()` 返回的是元素**即将离开**的位置。你的点击落在元素**原来**的地方,落在当时在那儿的别的东西上,而日志忠实地告诉你:它点击了你要求的坐标。
修法是两行 —— 关掉平滑滚动,然后重新读矩形:
```python
self.ev("""(()=>{document.documentElement.style.scrollBehavior='auto';
if(document.body) document.body.style.scrollBehavior='auto'; return 'ok';})()""")
moved = self.ev("""(()=>{const e=%s; if(!e) return 'no';
e.scrollIntoView({block:'center', behavior:'instant'}); return 'ok';})()""" % sel)
time.sleep(0.35) # 等布局稳定
# 然后再读矩形(元素晚一帧才渲染的话,外面再套一层重试)
```
两个细节很关键:`scrollIntoView` 上的 `behavior:'instant'` 会覆盖 CSS;而读取前那一小段等待,才是让坐标可信的原因。没有重试循环的话,这个辅助函数在晚一帧渲染的元素上仍会偶发失败 —— **偶发失败是最糟的一类,因为它会让你怀疑自己的判断。**
## 原因二:窗口最小化时,坐标点击被丢弃
修好之后,我的点击辅助函数在一次性探针里能用,一放进真正的脚本就再也没成功过。同样的选择器、同样的坐标、同样的页面。差别最后落在可见性上:
```
document.visibilityState → "hidden"
document.hasFocus() → true
```
当目标标签页不是可见窗口的活动标签页时,CDP 合成的坐标级鼠标事件是不可靠的 —— 浏览器可能接受、可能限流、也可能直接不往任何地方送。注意 `hasFocus()` **在文档隐藏时依然返回 `true`**,所以它不是有效的判断依据;要看的是 `document.hidden` / `visibilityState`。
两个修法,我都用:
```python
def activate(self): # 把标签页提到窗口最前
self.cdp("Page.bringToFront")
time.sleep(0.2)
def click_js(self, sel): # 在页面内部派发事件来点击
return self.ev("""(()=>{const e=%s; if(!e) return 'no';
['mousedown','mouseup','click'].forEach(k=>
e.dispatchEvent(new MouseEvent(k,{bubbles:true,cancelable:true,view:window})));
return 'js-clicked';})()""" % sel)
```
`Page.bringToFront` 解决了常见情况。但真正从不出错的是 JS 派发的点击:它不依赖窗口是否可见,而且 React 挂在根节点上的事件监听器会收到它,和收到真实点击的方式完全一样 —— 因为事件冒泡到的是同一个根。一个用坐标点击五次才能打开一次的日期选择器,用这种方式**每次都开**。
需要知道的取舍:派发事件会跳过 hit-testing,所以它也会「点到」用户根本够不到的元素(被遮住、滚出屏幕、透明度为零)。在这里这是优点 —— 如果你用它去掩盖布局 bug,那就是坑 —— 所以凡是可视化的东西我都用坐标点击,而遇到自定义控件时才伸手去拿 `click_js`。
## 原因三:页面收到了点击,然后拒绝了它
这个最浪费时间,因为它的症状和前两个一模一样:点击发出去了,什么都没变。
区分它们只需要注入一个监听器:
```js
window.__ev = [];
['mousedown','mouseup','click'].forEach(k =>
document.addEventListener(k, e => window.__ev.push(k + ':' + e.target.tagName), true));
```
- 有事件记录 → 输入链路是通的,问题在应用侧(原因三),或者在你读取结果的方式上(见下)。
- 一条都没有 → 是原因一或原因二。
挂上监听器后我点了一次 `Confirm`,缓冲区满了 —— 说明点击没问题,是应用在说「不」。理由在一段我从没想过要看的文字里:
> “Please create a new shop welcome voucher after the existing one is expired.”
表单拒绝创建同类型的第二张券。没有字段级红字、头 500 毫秒里没有提示条、URL 也没变。**拒绝是真的,点击是无辜的。**
这一类 bug 还有两个工具:
```js
document.elementFromPoint(x, y) // 那个坐标上到底是什么?
document.querySelectorAll('button').length // 你的选择器到底命中了几个?
```
第二个抓到过一处很细的问题:门户页面在 DOM 里留了一个**隐藏的** `Confirm` 按钮(来自收起状态的日期面板),于是 `.pop()` 返回的是那个看不见的副本,点击自然落到了无人的地方。按可见性和尺寸一起过滤:
```python
"[...document.querySelectorAll('button')].filter(b=>/^Confirm$/.test(b.innerText.trim())"
" && b.offsetParent!==null && b.getBoundingClientRect().width>0).pop()"
```
## 第四种「失败」,只有回读才发现
还有一类「失败」其实不是失败:**写入成功了,而你的读取是旧的。** 我因为列表页没有立刻显示,就断定创建被静默拒绝,又点了一次;后来才发现那张券一直都存在 —— 列表只是延迟了大约十分钟。这个教训可以推广:一次写入之后,唯一可信的结论是在**等待足够时间之后**回读得到的,而「看不到变化」意味着**未知**,不是失败。
## 检查清单
1. 关掉平滑滚动;用 `behavior:'instant'` 滚动;等一会儿再读矩形。
2. 点击前调用 `Page.bringToFront`;判断依据用 `document.hidden`,不是 `hasFocus()`。
3. 自定义控件用 JS 派发 `mousedown`/`mouseup`/`click`,别点坐标。
4. 先挂事件监听器,证明输入是否到达,再去怪应用。
5. 数一数选择器命中几个 —— 组件库里隐藏的重复元素很常见。
6. 写入之后先等待再回读;永远不要重复点击一个还没验证过的提交按钮。
这些都不算冷门知识,但它们中的每一条,在一行 `clicked @591,361` 面前都是隐形的。如果你正在对着一个自己无法控制的后台做自动化:点击是整条链路里最不可靠的一环 —— 也是最容易被最后才怀疑的一环。
我平时就在做这类浏览器自动化与排障 —— 自托管服务、Docker 与 Traefik 栈,偶尔还有被 bot 墙挡住的市场页面。如果你有条工作流总因为「UI 有自己的想法」而断掉,跟我说说:[WhatsApp](https://wa.me/60127972969) · [[email protected]](mailto:[email protected]?subject=Browser%20automation%20debugging) · [hoelee.com](https://hoelee.com)。