Docs overhaul: beginner quick start + all-in-one NocoDB stack

- README rewritten around a copy-paste quick start (CLI and Portainer),
  verified alert test, day-to-day NocoDB operations, troubleshooting table, FAQ
- new docker-compose.allinone.yml: NocoDB (pinned 2026.09.0, SQLite) + monitor
  on a private network, with healthchecks and the Telegram IPv4 pin documented
- docs/: QUICKSTART-PORTAINER, TELEGRAM-SETUP, NOCODB-SETUP, ARCHITECTURE,
  OPERATIONS, TROUBLESHOOTING (replace DOCUMENTATION.md + COMPOSE-SETUP.md)
- secrets: SECRETS.md is gitignored and untracked; tracked template is
  SECRETS.example.md; real base id / chat id removed from .env.example
- LICENSE (MIT), .gitignore/.dockerignore tidied
- AGENTS.md: layout, iron rules, verification gates; host-specific deploy
  details moved to the gitignored OPS-INTERNAL.md
This commit is contained in:
2026-10-06 16:37:59 +08:00
parent 44851a1c29
commit 5d9db48a26
16 changed files with 1354 additions and 647 deletions
+141
View File
@@ -0,0 +1,141 @@
# Architecture
How the pieces fit together, and what every line of the compose file and Dockerfile is doing. Read this before changing the code.
---
## 1. Runtime
```
carousell-monitor container (python:3.11-alpine, no ports, no web UI)
│
├─ once at start ─ bootstrap(): create/fix the NocoDB tables, load every known
│ product_url into an in-memory seen-set
│
└─ every TICK_SECONDS (default 60s):
├─ read the Settings table (watch list)
├─ per watch whose check_interval_minutes has elapsed:
│ GET the Carousell search page
│ parse the embedded JSON state → SearchListing.listingCards[]
│ keep only URLs not in seen
│ INSERT them into Listings
│ first pass for a watch (last_checked_at empty) → notified = true (silent seed)
│ afterwards → notified = false (queued)
│ advance last_checked_at, then sleep FETCH_GAP_SECONDS
├─ send every queued listing (notified = false), one Telegram message each,
│ applying the ignore filters, then mark it notified
└─ write /data/health.json → the Docker HEALTHCHECK reads it
```
External calls: **Carousell** (search pages, and listing photos when alerting), **NocoDB** (REST on the private network), **Telegram** (Bot API). Nothing calls in.
## 2. The files
| File | Role |
|---|---|
| `monitor.py` | everything: config, HTTP, NocoDB schema bootstrap, extraction, IO, Telegram, health, alerting |
| `healthcheck.py` | Docker HEALTHCHECK probe — exits 0 only if `health.json` is fresh **and** `ok: true` |
| `Dockerfile` | `python:3.11-alpine`, copies the two scripts, declares the healthcheck |
| `docker-compose.allinone.yml` | NocoDB + monitor, private `carousell` network |
| `docker-compose.yml` | monitor only, joins an existing NocoDB network (production variant) |
| `test_pagination.py` | regression suite (stdlib, no network): NocoDB paging, the Telegram HTTP verb, failure thresholds |
| `.env.example` | the settings, documented |
## 3. Why it is built this way
**Two containers minimum.** NocoDB holds all state; the monitor is disposable. You can delete the monitor container, rebuild it, or change `monitor.py` and nothing is lost — the seen-set is rebuilt from `Listings` on start.
**Archive and notify are decoupled.** A listing is inserted with `notified = false`; a later pass sends it and flips the flag. So a Telegram outage or a wrong token never loses a listing — the queue just drains later. It also means "no alert" and "no data" are distinguishable failures.
**Filters are data, not config.** Ignored sellers and keywords live in NocoDB tables that are re-read every cycle, so muting something is a UI action, not a redeploy. Operational knobs (which searches, how often, whether to notify) are in `Settings`; only credentials and tuning live in the environment.
**First pass is silent.** `last_checked_at` empty means "never seeded": the current listings are archived with `notified = true`. Otherwise the first start would fire hundreds of messages.
**Dedupe on the canonical product URL**, `https://www.carousell.com.my/p/<id>/` — never the search URL and never the raw listing id alone. The search page's own links carry tracking parameters that change between loads.
**The failure threshold matters.** `FAILURE_RATIO_THRESHOLD` (default `1.0`) means a tick only counts as failed if *every* watch failed. Carousell soft-blocks individual searches from time to time; without this, one flaky search would mark the container unhealthy and fire a failure alert every tick.
**Telegram failures are alerted on their own edge.** `ERROR_ALERT_AFTER` consecutive failed ticks send one alert, and clearing the failure state sends one recovery notice — a state machine debounce, not a per-tick message.
## 4. The compose file, block by block
### `docker-compose.allinone.yml`
```yaml
services:
nocodb:
image: nocodb/nocodb:2026.09.0 # pinned on purpose; the monitor targets this meta API
volumes:
- nocodb-data:/usr/app/data # SQLite db + attachments
ports:
- "${NOCODB_PORT:-8080}:8080" # only the web UI is published
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://127.0.0.1:8080/api/v1/health >/dev/null 2>&1 || exit 1"]
carousell-monitor:
build: . # built locally from this repo — no registry involved
depends_on:
nocodb:
condition: service_healthy # do not start before the database answers
extra_hosts:
- "api.telegram.org:149.154.166.110"
volumes:
- carousell-data:/data # health.json, alert_state.json
networks: [carousell]
```
| Key | Why |
|---|---|
| `build: .` | the image is built from this directory; nothing is pulled from a registry |
| `depends_on: service_healthy` | the monitor bootstraps the schema at startup, so NocoDB must be up first |
| `extra_hosts` | pins `api.telegram.org` to its IPv4. On a Docker network without IPv6, the embedded DNS can hand back an AAAA record and the send hangs with no error — the alert is lost silently. Refresh the IP with `dig +short api.telegram.org` if sends start timing out; delete the two lines if your host has working IPv6 |
| named volumes | survive container recreation and `down`; no host-path/ACL problems |
| private network | only `nocodb` publishes a port, and the monitor resolves it as `http://nocodb:8080` — no IPs anywhere |
### `docker-compose.yml` (existing NocoDB)
The same monitor service, minus NocoDB, joining a network that already exists:
```yaml
networks:
- bridge_hoelee # ← your existing network name
networks:
bridge_hoelee:
external: true # created by the NocoDB stack, not by this one
```
Point `NOCODB_URL` at the existing container's name on that network (e.g. `http://nocodb:10380` if it listens on a non-default port).
## 5. The Dockerfile
```dockerfile
FROM python:3.11-alpine
WORKDIR /app
COPY monitor.py healthcheck.py /app/
RUN mkdir -p /data
VOLUME ["/data"]
HEALTHCHECK --interval=60s --timeout=15s --start-period=120s --retries=3 \
CMD python /app/healthcheck.py
CMD ["python", "-u", "/app/monitor.py"]
```
- **stdlib only** — no `pip install`, so the image is ~60 MB and builds in seconds.
- **Code is baked in** — shipping a change means rebuilding (`docker compose up -d --build`), not restarting.
- **`python -u`** — unbuffered output so logs appear immediately.
- **`start-period=120s`** — the first tick includes the schema bootstrap; give it room before Docker starts reporting health.
## 6. Extraction details (why it is fragile)
Carousell server-renders the search results into a `<script type="application/json">` blob; the monitor parses the largest one and reads `SearchListing.listingCards[]`. Each card carries `listingID`, `title`, `thumbnailURL`, `seller.username`, the condition as a paragraph, and timestamps under `aboveFold` (`time_created`, or `active_bump` for bumped listings).
Consequences worth knowing before you patch it:
- **HTTP 200 does not mean success.** Under soft rate limiting Carousell returns 200 with `listingCards: null`. That is a retryable condition, not an empty result set — the code raises a named error for it.
- **No `application/json` blob at all** means a challenge/blocked page. Same treatment.
- **Ad cards exist** (`listingID = 0`) and are skipped rather than archived.
- **Thumbnail URLs carry a `_progressive_thumbnail` suffix** that is stripped to get the full-size image.
## 7. Scaling and limits
One process polls everything. That is comfortable for tens of searches at 5-minute intervals. If you need more, raise `check_interval_minutes` rather than lowering `TICK_SECONDS`, and keep `FETCH_GAP_SECONDS ≥ 1` — the point of the gap is that your traffic never looks like a burst.
+101
View File
@@ -0,0 +1,101 @@
# NocoDB setup
The monitor stores everything in [NocoDB](https://nocodb.com) — an open-source Airtable alternative. It gives you the archive (with thumbnails), the watch-list UI, and the filter tables, and it is the only interface you need day to day.
The **base** and the **API token** are created by you; the **four tables** are created by the monitor on its first start.
---
## 1. First run of NocoDB
Start it (all-in-one stack):
```bash
docker compose -f docker-compose.allinone.yml up -d nocodb
```
Open `http://<host>:8080`, then:
1. **Create the admin account.** First click wins — pick a real password, there is no password reset without mail configured.
2. **Create a Base** (`+ New Base`), named e.g. `Carousell`. A base is a database; everything this project needs lives inside it.
3. Do not bother creating tables by hand. The monitor does that.
## 2. Collect the two values
**Base id** — open the base and read the browser URL:
```
http://localhost:8080/dashboard/#/nc/base/poqw1zjw3hnsk37/...
^^^^^^^^^^^^^^^ NOCODB_BASE_ID
```
**API token** — click your avatar (bottom-left) → **Account Settings** → **Tokens** → **Create token**.
- Give it a name (`carousell-monitor`), no expiry (or an expiry you will remember to renew).
- Copy the value immediately — it starts with `nc_pat_` and NocoDB will not show it again. That is `NOCODB_TOKEN`.
If your NocoDB is multi-workspace, create the token inside the workspace that owns the base.
## 3. The tables the monitor creates
Restart the monitor after filling in the token; `Listings`, `Settings`, `IgnoredSellers` and `IgnoredKeywords` appear in the base. Creating them is idempotent — restarting never duplicates a table or a row.
### `Settings` — your watch list (you edit this)
| Column | Type | Notes |
|---|---|---|
| `title` | text | label shown in alerts and used to link filters |
| `url` | URL | the full Carousell search URL, **must contain `sort_by=3`** |
| `enabled` | checkbox | untick to pause a search |
| `notify` | checkbox | untick to archive without alerting |
| `check_interval_minutes` | number | how often this search is polled (default 5) |
| `last_checked_at` | datetime | maintained by the monitor; empty = never seeded yet |
### `Listings` — the archive (written by the monitor)
| Column | Type | Notes |
|---|---|---|
| `product_url` | URL | dedupe key, `https://www.carousell.com.my/p/<id>/`, no query string |
| `title` | text | listing title |
| `price` | decimal | numeric, `RM` stripped (`85.00`) so you can sort and filter |
| `condition` | select | Brand new / Like new / Lightly used / Well used / Heavily used / Used |
| `image_url` | URL | the full-size photo URL |
| `image` | attachment | same photo as an attachment — renders as a thumbnail in grid view |
| `seller_name` / `seller_url` | text / URL | who is selling |
| `search_title` / `search_url` | text / URL | which watch found it |
| `listed_at` | datetime (UTC) | when the seller posted it |
| `first_seen_at` | datetime (UTC) | when the monitor first saw it |
| `notified` | checkbox | `false` = still queued for an alert; `true` = sent or deliberately silenced |
| `skip_notify` | checkbox | set when a filter matched (see below) |
### `IgnoredSellers` — mute a seller everywhere
One row per seller, column `seller_name` = the Carousell username. Their listings stay in the archive (`skip_notify = true`) but never reach Telegram.
### `IgnoredKeywords` — mute words for one search
| Column | Type | Notes |
|---|---|---|
| `watch` | link → `Settings` | **pick the search from the dropdown**; keywords only apply to it |
| `keyword` | text | case-insensitive substring match against the listing **title** |
Both filter tables are reloaded every tick, so edits take effect within a minute without a restart.
## 4. Getting a good search URL
1. Search on [carousell.com.my](https://www.carousell.com.my).
2. Set the sort to **Recent** (the newest listing must come first — the monitor only sees what the first page shows).
3. Copy the address bar into the `url` column. It should look like:
```
https://www.carousell.com.my/search/mechanical-keyboard?addRecent=true&canChangeKeyword=true&includeSuggestions=true&sort_by=3&t-search_query_source=direct_search
```
Tips: keep searches specific (a specific model, a narrow category). Each search returns roughly the first page of results; a very broad search whose first page turns over slowly will miss the churn below the fold.
## 5. Quality-of-life
- **Images**: switch `Listings` to **grid view** (`Fields` → include `image`) for a thumbnail wall of everything found.
- **Timezones**: stored timestamps are UTC on purpose. Change your display timezone in NocoDB's account settings if you want local times in the UI.
- **Filters/sorts**: `price` is numeric, so `price < 200 AND condition = 'Like new'` works.
- **Backups**: everything is in the `nocodb-data` volume. NocoDB also has its own export (base → `…` → Export), which is the friendlier thing to keep off-box.
+124
View File
@@ -0,0 +1,124 @@
# Operations
Running it, watching it, backing it up, upgrading it.
---
## Health at a glance
The container writes `/data/health.json` at the end of every tick, and the Docker HEALTHCHECK (`healthcheck.py`) reads it every 60 s:
```bash
docker exec carousell-monitor cat /data/health.json
```
```json
{
"last_run_epoch": 1791275045,
"ok": true,
"error": "",
"watch_count": 3,
"new_this_tick": 0,
"failed_watches": 0
}
```
| Field | Meaning |
|---|---|
| `ok` | the last tick completed without exceeding the failure threshold |
| `error` | empty, or which watches failed and why (`[partial 1/3] …` when the ratio threshold absorbed it) |
| `last_run_epoch` | when the tick ended (Unix seconds, UTC) |
| `watch_count` | how many enabled watches were read from `Settings` |
| `new_this_tick` | listings fetched this tick that were not already known — **not** how many alerts were sent |
Two traps worth internalising:
- **`ok: true` only proves the fetch loop ran.** Notification failures are retried on the next tick rather than reported, so an archive that fills up happily can still be silent in Telegram. To check the notify path, untick `notified` on a row and watch it flip back to `true` (that flip only happens after Telegram answered 200).
- **`new_this_tick: 0` is not evidence of anything** — it counts fetched rows, not delivered messages.
The container's own health is the other half:
```bash
docker compose -f docker-compose.allinone.yml ps # State / Health column
docker inspect --format '{{.State.Health.Status}}' carousell-monitor
```
`unhealthy` = the last tick failed or is older than `HEALTH_STALE_SECONDS` (600 s).
## Logs
```bash
docker compose -f docker-compose.allinone.yml logs -f --tail 100 carousell-monitor
```
The loop is deliberately quiet: one `ready:` line at startup, then nothing unless something fails. Per-watch failures are logged to stderr, and hard failures appear in `health.json`. If you are debugging "why no alert", logs are the wrong place — use the decision tree in [TROUBLESHOOTING.md](TROUBLESHOOTING.md).
## Alerting on the monitor itself
`ERROR_ALERT_AFTER` (default 3) consecutive failed ticks trigger one Telegram message:
```
🚨 carousell-monitor 故障
连续失败 3 次
错误: Uniform: carousell fetch HTTP 403
容器将标记为 unhealthy
```
and one recovery message when the next good tick arrives. The debounce state lives in `/data/alert_state.json`, so a container restart does not re-fire an alert you already saw.
The alert strings are Chinese in the current code (`monitor.py` → `alert_on_health()`); change the two `msg = (…)` literals if you want English.
## Backups
| What | Where | How |
|---|---|---|
| All listings, settings and filters | volume `nocodb-data` | stop the stack, tar the volume; or use NocoDB's own **Export base** |
| Health + alert state | volume `carousell-data` | disposable — do not bother |
| This repo's config | `.env` / the Portainer stack file | keep a copy in your password manager |
Nothing else is stateful. NocoDB is the single source of truth.
## Upgrading
**The monitor** (code change in this repo):
```bash
git pull
docker compose -f docker-compose.allinone.yml up -d --build
```
The rebuild re-creates the container. The schema bootstrap and the seen-set are idempotent, so nothing is duplicated.
**NocoDB** — change the image tag in the compose file and:
```bash
docker compose -f docker-compose.allinone.yml up -d
```
NocoDB migrates its own database on start. Back up `nocodb-data` first. The monitor's schema bootstrap talks to NocoDB's **meta API**, which does change between majors: the tag in `docker-compose.allinone.yml` is the version range this code is verified against, so upgrade it deliberately (and check the table columns in the UI afterwards).
**Portainer users:** same thing through the UI — edit the stack file, *Update the stack*. Be careful with `Re-pull image` on a stack whose env values were entered through Portainer's panel (see the README's masking warning).
## Tests
```bash
python test_pagination.py # stdlib only, no network, exit 0 = pass
```
Covers the things that have actually broken: NocoDB paging past 1000 rows, the Telegram HTTP verb bug, the failure-ratio threshold maths, and the fetch-gap timing.
## Tuning notes
| Goal | Change |
|---|---|
| More search coverage | add rows to `Settings`, keep `check_interval_minutes` ≥ 5 |
| React faster | lower `check_interval_minutes` (not `TICK_SECONDS` below ~30 s) |
| Be gentler on Carousell | raise `check_interval_minutes`, keep `FETCH_GAP_SECONDS` ≥ 1 |
| Alert inbox too noisy | set `notify = false` on a watch, or add entries to `IgnoredKeywords` / `IgnoredSellers` |
| Fewer failure alerts | raise `ERROR_ALERT_AFTER`, or let `FAILURE_RATIO_THRESHOLD` stay at `1.0` |
## Data safety rules
- `.env` and your NocoDB token are credentials. Never commit them, never paste them into an issue.
- The NocoDB token is scoped to a workspace: regenerate it in NocoDB and update the stack if it leaks.
- The monitor never deletes or modifies archived listings, apart from the `notified` / `skip_notify` flags. Cleaning up the archive is your job — NocoDB's grid view deletes rows fine.
+106
View File
@@ -0,0 +1,106 @@
# Quick start on Portainer (click-by-click)
For a fresh machine, no shell needed. Portainer runs the same compose file as the CLI flow — see the [README](../README.md) for that version.
**Before you start:** Portainer must already be up and connected to a Docker endpoint, and Portainer's own container needs access to the Docker socket (the standard install does). NocoDB will need one free host port — `8080` by default.
---
## 1. Create the NocoDB container first
Because you cannot paste the base id and token until NocoDB exists, do this in two deploys.
1. **Stacks → Add stack**
2. **Name:** `carousell-monitor`
3. **Build method:** *Web editor*
4. Paste the **first service only** for now — the `nocodb:` block from [`docker-compose.allinone.yml`](../docker-compose.allinone.yml) plus the closing `volumes:` / `networks:` sections:
```yaml
services:
nocodb:
image: nocodb/nocodb:2026.09.0
container_name: carousell-nocodb
restart: unless-stopped
environment:
PORT: "8080"
NC_AUTH_JWT_SECRET: <paste output of: openssl rand -hex 32>
NC_DISABLE_TELE: "true"
NC_ALLOW_LOCAL_HOOKS: "false"
volumes:
- nocodb-data:/usr/app/data
ports:
- "8080:8080"
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://127.0.0.1:8080/api/v1/health >/dev/null 2>&1 || exit 1"]
interval: 30s
timeout: 10s
retries: 5
start_period: 60s
networks:
- carousell
volumes:
nocodb-data:
networks:
carousell:
```
5. **Deploy the stack.** Wait for `carousell-nocodb` to show as running/healthy.
## 2. Create the base and collect the two values
Open `http://<your-host>:8080`:
1. Create the admin account.
2. Create a base (`+ New Base`) — call it `Carousell`.
3. Copy the **base id** from the URL (`…/nc/base/<BASE_ID>/…`).
4. Avatar (bottom-left) → **Account Settings → Tokens → Create token**, and copy the `nc_pat_…` value. It is shown once.
## 3. Add the monitor to the same stack
1. Get the Telegram values first ([docs/TELEGRAM-SETUP.md](TELEGRAM-SETUP.md)) — bot token from @BotFather, chat id from @userinfobot.
2. Back in Portainer: **Stacks → carousell-monitor → Editor** tab.
3. Replace the whole file with the content of [`docker-compose.allinone.yml`](../docker-compose.allinone.yml), **with the real values written in directly** instead of `${...}` placeholders:
```yaml
environment:
NOCODB_URL: http://nocodb:8080
NOCODB_TOKEN: nc_pat_your_token_here
NOCODB_BASE_ID: your_base_id_here
TELEGRAM_BOT_TOKEN: 8123456789:AAF...
TELEGRAM_CHAT_ID: 123456789
```
and the same for `NC_AUTH_JWT_SECRET` / `NC_DISABLE_TELE` in the `nocodb` service.
> **Why inline values instead of Portainer's environment panel?** Portainer CE masks secret-looking values that are entered through the UI and stores the mask with the stack. On the next stack update the container receives `***` and silently stops authenticating. Editing the YAML keeps the real value on your Portainer host, which is where a stack file's secrets live anyway.
4. **Update the stack.** Portainer recreates the services that changed and adds `carousell-monitor`.
5. **Containers → carousell-monitor → Logs**: expect one line starting with `ready: listings=… settings=… seen=0`.
## 4. Start using it
In NocoDB, open the base — four tables are now present. Add your first watch to **`Settings`** (see [docs/NOCODB-SETUP.md](NOCODB-SETUP.md#3-the-tables-the-monitor-creates)) and give it a minute. Then untick `notified` on any `Listings` row to force a test alert.
## 5. Everyday maintenance in Portainer
| Task | Where |
|---|---|
| Change a search / filters | NocoDB UI — nothing to redeploy |
| Change credentials or tuning | **Stacks → carousell-monitor → Editor** → edit YAML → **Update the stack** |
| See why it is unhealthy | **Containers → carousell-monitor → Logs**, then the `Listings`/`Settings` tables for real progress |
| Restart | **Containers → carousell-monitor → Restart** (state is in NocoDB, nothing is lost) |
| Update this project | pull the new code on the host, then **Editor** → *Update the stack* with `Re-pull image`/rebuild enabled, or `docker compose up -d --build` on the CLI |
| Back up | **Volumes → nocodb-data** (all data) — or NocoDB's own base export |
## 6. Optional: deploy from the repository instead
**Stacks → Add stack → Repository**, with:
| Field | Value |
|---|---|
| Repository URL | `https://github.com/hoelee/carousell-monitor` |
| Reference | `refs/heads/main` |
| Compose path | `docker-compose.allinone.yml` |
Set `NOCODB_BASE_ID` / `NOCODB_TOKEN` / `TELEGRAM_*` / `NC_AUTH_JWT_SECRET` in the environment panel for this one (they are not in git). Be aware of the masking caveat above, and that Portainer clones the repository on every deploy.
+71
View File
@@ -0,0 +1,71 @@
# Telegram setup
Two values are needed: a **bot token** (who sends) and a **chat id** (where it sends). Both are free and take about a minute.
---
## 1. Create the bot
1. Open Telegram and search for **@BotFather** (the one with a blue checkmark).
2. Send `/newbot`.
3. It asks for a **name** — anything, e.g. `My Carousell Alerts`.
4. It asks for a **username** — must be unique and end in `bot`, e.g. `mycarousell_alerts_bot`.
5. It replies with a token like:
```
Use this token to access the HTTP API:
8123456789:AAF7xK3nQw8_your_token_here_9dZ
```
That whole string is `TELEGRAM_BOT_TOKEN`. Treat it like a password — anyone holding it can send and read messages as your bot.
Optional, but nice: `/setdescription` and `/setuserpic` to make the alert messages look intentional.
## 2. Get your chat id (alert yourself)
1. **Send your new bot a message** — click the link BotFather gave you and say `hi`. This step matters: until you have messaged the bot, it is not allowed to message you, and sends fail with `chat not found`.
2. Message **@userinfobot**. It replies with your id:
```
Id: 123456789
```
That number is `TELEGRAM_CHAT_ID`.
## 3. Or alert a group
1. Create the group (or use an existing one) and **add your bot** to it as a member.
2. Send a message in the group (any text).
3. Read the group's chat id:
```bash
curl -s "https://api.telegram.org/bot<YOUR_TOKEN>/getUpdates"
```
Look for `"chat":{"id":-1001234567890,"title":"..."}`. Group ids are **negative** and usually start with `-100`. Use that number as `TELEGRAM_CHAT_ID`.
If `getUpdates` returns `{"ok":true,"result":[]}`, the bot has not seen any message yet — send another one in the group and retry.
## 4. Verify before you blame the monitor
```bash
# who am I?
curl -s "https://api.telegram.org/bot<YOUR_TOKEN>/getMe"
# send a test message
curl -s -X POST "https://api.telegram.org/bot<YOUR_TOKEN>/sendMessage" \
-d chat_id=<YOUR_CHAT_ID> -d text="hello from carousell-monitor"
```
Expected: `{"ok":true,...}` and a message on your phone. If this works, the monitor's credentials are right and anything missing is a monitor-side issue (filters, `notify` checkbox, pending queue).
| Error | Meaning |
|---|---|
| `401 Unauthorized` | token is wrong, or it was revoked in BotFather |
| `400 chat not found` | wrong chat id, or you never messaged the bot / the bot is not in the group |
| `403 bot was blocked by the user` | you blocked the bot — unblock it |
| the `curl` hangs forever | DNS/IPv6 trouble. The compose files pin `api.telegram.org` to its IPv4 address in `extra_hosts`; refresh that IP with `dig +short api.telegram.org` |
## 5. Rotate it later
If the token leaks: @BotFather → `/revoke` → pick the bot → you get a new token. Update `.env` (or the Portainer stack environment) and restart the monitor.
+116
View File
@@ -0,0 +1,116 @@
# Troubleshooting
A symptom, a cause, a fix — plus the decision tree to run when the symptom is the vague one: *"it stopped alerting me"*.
---
## 1. The decision tree for "no alert"
Four different failures look identical from the outside. Find which stage breaks before changing anything.
**Stage 0 — is it running at all?**
```bash
docker compose -f docker-compose.allinone.yml ps
docker exec carousell-monitor cat /data/health.json
docker exec carousell-monitor cat /data/alert_state.json # fail_streak, alerted
```
- `ok: false` → the fetch loop is dead. Read `error` and jump to §2.
- `last_run_epoch` older than a couple of ticks → the loop is stuck; check the logs for a traceback.
- `No such file` → the container never completed a tick (bad credentials most likely — see `NOCODB_TOKEN not set`, §2).
**Stage 1 — is it fetching?**
Open `Settings` and check `last_checked_at` on your watch. It should advance every `check_interval_minutes`. If it never advances, the watch is disabled (`enabled` unticked), the interval is huge, or every fetch is failing.
**Stage 2 — is it archiving?**
Open `Listings`, sort by `first_seen_at` descending. New rows appearing means fetch + parse + NocoDB writes all work. If `last_checked_at` advances but `Listings` stays empty, you are the victim of an over-narrow search, not a bug — nothing new has appeared.
**Stage 3 — is it notifying?**
This is where the archive can look healthy while Telegram is silent. Sort `Listings` by `notified`:
- Rows with `notified = false` **and** `skip_notify = false` that stay false → the send is failing (see §2 transport rows) or the monitor is stuck.
- Rows with `skip_notify = true` → a filter matched, by design. Check `IgnoredSellers` / `IgnoredKeywords` and the watch's `notify` checkbox.
**Stage 4 — is the transport sane?**
From inside the container, prove the credentials and egress in one shot:
```bash
docker exec carousell-monitor sh -c '
wget -qO- "https://api.telegram.org/bot$TELEGRAM_BOT_TOKEN/getMe"'
```
`{"ok":true,...}` means the token is good and the container can reach Telegram. Then send a real message:
```bash
docker exec carousell-monitor sh -c '
wget -qO- --post-data="chat_id=$TELEGRAM_CHAT_ID&text=test" \
"https://api.telegram.org/bot$TELEGRAM_BOT_TOKEN/sendMessage"'
```
If `getMe` works but the send does not, it is the chat id (a bot cannot open a conversation with you before you have messaged it — see [TELEGRAM-SETUP.md](TELEGRAM-SETUP.md)).
**Stage 5 — force a notification.**
In NocoDB, untick `notified` on any listing row. Within one tick the monitor re-sends it and flips the flag back to `true`. The flip **is** the proof of delivery: it only happens after Telegram returns HTTP 200.
---
## 2. Symptom → cause → fix
| Symptom | Likely cause | Fix |
|---|---|---|
| `NOCODB_TOKEN not set` in the logs, container exits | `.env` missing the token, or the stack did not pick it up | fill it in and re-create the container (`up -d`, or Portainer *Update the stack*) |
| `list tables failed: HTTP 401/403` | token wrong, expired, or from another workspace | create a token inside the workspace that owns the base |
| `list tables failed: HTTP 404` | wrong `NOCODB_BASE_ID` | re-copy the id from the base URL |
| `create table … failed` / `add column … failed` | the token can read but not write, or the base was deleted | verify with a write test (create a scratch table in the UI), check the base exists |
| `connection refused` to `nocodb` | the monitor is not on the same Docker network, or the host name is wrong | both services must share the network in the compose file; `NOCODB_URL` must use the service/container name, not `localhost` |
| `no application/json state found (blocked/ratelimited?)` | Carousell returned a challenge or error page | raise `check_interval_minutes`, keep `FETCH_GAP_SECONDS ≥ 1`, reduce the number of watches; the watch retries next interval |
| `listingCards null (soft-block/ratelimit?)` | Carousell answered 200 with an empty state blob | same as above — this is rate limiting, not a bug |
| `tick error: …` and the container goes unhealthy | an exception escaped the tick (NocoDB error, unexpected page shape) | read the message; the loop keeps running and retries |
| `telegram sendMessage failed: 400` in the logs | the request was malformed — historically a code bug where the Telegram **method name** was used as the HTTP verb. Fixed; keep `test_pagination.py` passing | run the test suite |
| Telegram sends hang and time out | Docker DNS handing back an IPv6 (AAAA) address on a network with no IPv6 | keep/refresh the `extra_hosts` pin (`dig +short api.telegram.org`) |
| `chat not found` (400) | you never started the bot; or the group id is wrong | message the bot once; re-read the id from `getUpdates` |
| Archive filling up, Telegram silent | see the tree in §1 — usually the `notify` checkbox, a filter match, or a failed send that is being retried |
| Every listing alerted twice | you reset `notified`, or two monitors point at the same base | untick only once; run one monitor per base |
| Hundreds of messages right after setup | the watch's `last_checked_at` was not empty (it was seeded before) | expected on a re-seed; delete `last_checked_at` only when you *want* a silent re-seed |
| Thumbnails broken in the NocoDB grid | the `image` column is not an Attachment column | restart the monitor — the bootstrap recreates missing columns |
| Timestamps 8 hours off | stored as UTC by design | change the display timezone in NocoDB |
| Container healthy but nothing happens | no enabled watches, or all of them fall outside their interval | add/enable a row in `Settings`, lower `check_interval_minutes` |
---
## 3. Useful commands
```bash
# status
docker compose -f docker-compose.allinone.yml ps
# follow the logs
docker compose -f docker-compose.allinone.yml logs -f --tail 100 carousell-monitor
# health + alert state as the container sees them
docker exec carousell-monitor cat /data/health.json
docker exec carousell-monitor cat /data/alert_state.json
# what credentials did it actually receive?
docker exec carousell-monitor sh -c 'env | grep -E "NOCODB|TELEGRAM|TICK|FAILURE"'
# is NocoDB reachable from the monitor's network?
docker exec carousell-monitor sh -c 'wget -qO- http://nocodb:8080/api/v1/health'
# restart / rebuild / stop
docker compose -f docker-compose.allinone.yml restart carousell-monitor
docker compose -f docker-compose.allinone.yml up -d --build
docker compose -f docker-compose.allinone.yml down
```
---
## 4. Reporting a bug
Include: the `health.json` contents, the last ~50 log lines, the compose file with every credential replaced by `***`, and your NocoDB version (`…/api/v1/version`). Do not paste tokens, chat ids or your base id.