- README rewritten around a copy-paste quick start (CLI and Portainer),
verified alert test, day-to-day NocoDB operations, troubleshooting table, FAQ
- new docker-compose.allinone.yml: NocoDB (pinned 2026.09.0, SQLite) + monitor
on a private network, with healthchecks and the Telegram IPv4 pin documented
- docs/: QUICKSTART-PORTAINER, TELEGRAM-SETUP, NOCODB-SETUP, ARCHITECTURE,
OPERATIONS, TROUBLESHOOTING (replace DOCUMENTATION.md + COMPOSE-SETUP.md)
- secrets: SECRETS.md is gitignored and untracked; tracked template is
SECRETS.example.md; real base id / chat id removed from .env.example
- LICENSE (MIT), .gitignore/.dockerignore tidied
- AGENTS.md: layout, iron rules, verification gates; host-specific deploy
details moved to the gitignored OPS-INTERNAL.md
Two problems found while verifying the pagination fix on the live stack:
1. fetch_listings() did state['SearchListing']['listingCards'] with no guard.
Carousell serves HTTP 200 with listingCards=null under soft rate limiting,
so the tick died with TypeError: 'NoneType' object is not iterable ->
ok:false -> container unhealthy, repeatedly.
2. run_tick() set ok = (no failures at all), so a single soft-blocked watch
out of 7 marked the whole monitor failed. That flaps on transient blocks
and (now that failures alert) would spam Telegram.
- fetch_listings: explicit null check -> clear retryable RuntimeError
- add FAILURE_RATIO_THRESHOLD (default 1.0 = all watches must fail); partial
failures are reported as '[partial n/N] ...' without failing the tick
- health gains failed_watches
- wire the knob into compose/.env.example/DOCUMENTATION/COMPOSE-SETUP
- 15 new checks: null vs empty cards, real card still parses, threshold edges
send_pending_notifications() read Listings with ?limit=1000 and no paging.
Once the table passed 1000 rows the newest records (highest Id, at the tail)
fell outside page 1, so they were never sent and never marked notified ->
notifications silently dead while health.json stayed ok:true. Found live
2026-09-22: Id 1007-1041 (35 rows, ~29h of listings) never alerted.
- add nc_list_all(): offset-paged full-table read
- use it for listings pending, seen, watches, settings, both ignore lists
- add alert_on_health(): Telegram failure alert debounced over
ERROR_ALERT_AFTER consecutive failed ticks, plus a recovery notice
(container already reported unhealthy via healthcheck.py on ok:false)
- new env knob ERROR_ALERT_AFTER wired into compose/.env.example/docs
- test_pagination.py: 30 checks incl. the page-2 regression and alert edges
Same tick previously fired every due watch back-to-back (burst of N requests,
worst right after a restart when all watches are due at once). Now each watch
URL fetch is followed by a pause of FETCH_GAP_SECONDS (float, default 1, env
tunable; 0 disables) on both success and failure paths.
Also thread kw_fk_col into load_ignored_keywords() inside
send_pending_notifications() so the physical-FK lookup (445040a) actually takes
effect instead of silently falling back to the watch Link column.