Replaces UV_INDEX_URL with UV_DEFAULT_INDEX which is the canonical env var for uv's default package index. Applied to builder, admin, and worker stages. Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Seed
Stealth web crawler with a live admin panel.
Bypasses Cloudflare / ArvanCloud bot-detection, executes configurable user flows, and tracks every run in a SQLite-backed dashboard.
Table of Contents
- Overview
- Architecture
- Project Structure
- Prerequisites
- Quick Start — Local
- Quick Start — Docker
- Configuration
- Scenarios
- Dynamic Flow System
- Admin Panel
- Makefile Reference
- Extending the Crawler
Overview
Seed runs two independent entry scenarios on a target website and then executes a fully configurable multi-step flow after each login:
| Scenario | Description |
|---|---|
| OTP Login | Types a phone number into the site's login form, waits for an SMS-OTP, submits it, then runs the flow |
| Fresh Session | Opens the browser with a clean cookie jar / localStorage so the site treats the visitor as a brand-new user, then runs the flow |
The flow consists of three automatic steps configured live from the admin panel:
- Search — types one or more search texts into a configurable search box
- Infinite Scroll — scrolls down (window or a specific container) until a target item appears or a scroll limit is hit
- Click — clicks each target item by its ID
Every run — success or failure — is recorded in SQLite and visible on the dashboard.
Architecture
Browser (undetected-chromedriver + selenium-stealth)
│
├── Scenario 1: OTP Login ──┐
└── Scenario 2: Fresh Session ┘
│
DynamicFlow
(reads active config
from SQLite at runtime)
│
┌──────────────┼──────────────┐
search_step scroll_step click_step
│
SQLite (data/tracker.db)
│
FastAPI Admin Panel
(dashboard + flow config CRUD)
- Anti-detection:
undetected-chromedriverpatches Chrome's automation flags;selenium-stealthmasksnavigator.webdriver, plugins, and language headers. A random user-agent is chosen per session. - Flow config is hot: the crawler reads the active config from the DB on every run — no restart needed after editing a config in the admin panel.
- SQLite: zero-infrastructure database. In Docker it lives on a named volume (
seed-data) so it survives container rebuilds.
Project Structure
seed/
├── crawler/
│ ├── driver.py # stealth Chrome factory + human_delay helper
│ ├── dynamic_flow.py # search → scroll → click executor
│ ├── flow_runner.py # legacy step-registry runner (extendable)
│ └── scenarios/
│ ├── otp_login.py # Scenario 1: SMS-OTP login
│ └── fresh_session.py # Scenario 2: fresh cookie session
├── admin/
│ ├── main.py # FastAPI app wiring
│ ├── models.py # SQLite DDL + async CRUD helpers
│ ├── routes.py # REST endpoints (stats + flow config CRUD)
│ └── static/
│ └── dashboard.html # Single-file SPA (no build step)
├── config.py # Typed config from environment / .env
├── main.py # CLI entry point
├── pyproject.toml # deps + Pyright config
├── uv.toml # uv index override (forces PyPI)
├── Makefile # developer shortcuts
├── Dockerfile # multi-stage: builder + runtime (Chrome included)
├── docker-compose.yml # admin service + on-demand crawler service
└── .env.example # all configurable variables with defaults
Prerequisites
Local
| Tool | Version | Notes |
|---|---|---|
| Python | ≥ 3.11 | |
| uv | latest | curl -LsSf https://astral.sh/uv/install.sh | sh |
| Google Chrome | any recent | must be installed on the host |
Docker
| Tool | Notes |
|---|---|
| Docker ≥ 24 | |
| Docker Compose v2 | bundled with Docker Desktop |
Chrome is installed inside the Docker image — no host Chrome needed.
Quick Start — Local
# 1. Clone and enter the project
git clone <repo-url> seed && cd seed
# 2. Install dependencies (creates .venv automatically)
make install
# 3. Configure
cp .env.example .env
$EDITOR .env # set TARGET_URL, PHONE_NUMBERS, ADMIN_PASSWORD, …
# 4. Start the admin panel
make admin
# → http://localhost:8000
# 5. In another terminal — run a scenario
make fresh # fresh-session scenario
make otp # all phones from PHONE_NUMBERS in .env
make otp PHONE=+98910... # single phone
Quick Start — Docker
# 1. Copy and fill in the env file
cp .env.example .env
$EDITOR .env
# 2. Build the image
make build
# 3. Start the admin panel (detached)
make up
# → http://localhost:8000
# 4. Run scenarios inside Docker
make docker-fresh
make docker-otp PHONE=+98910...
# 5. Watch logs
make logs
# 6. Stop everything
make down
The SQLite database is stored in the seed-data Docker volume and persists across rebuilds.
Worker browser sessions are recorded at 12 FPS into the shared seed-recordings
volume. The admin panel's Recordings tab lists and downloads the newest three
completed sessions.
Configuration
Copy .env.example to .env and set values before running anything.
| Variable | Default | Description |
|---|---|---|
TARGET_URL |
https://example.com |
The website to crawl |
ADMIN_USERNAME |
admin |
Admin panel username |
ADMIN_PASSWORD |
changeme |
Admin panel password — change this |
ADMIN_HOST |
127.0.0.1 |
Bind address for FastAPI (Docker overrides to 0.0.0.0) |
ADMIN_PORT |
8000 |
Port for the admin panel |
DB_PATH |
data/tracker.db |
Path to the SQLite file |
HEADLESS |
true |
Set to false to watch the browser (local only) |
CHROME_BINARY |
(auto) | Explicit path to Chrome, e.g. /usr/bin/google-chrome-stable |
RECORD_SESSIONS |
false |
Record browser sessions; Docker workers override this to true |
RECORDING_FPS |
12 |
Recording frame rate |
RECORDING_MAX_FILES |
3 |
Number of completed recordings retained across all workers |
RECORDINGS_DIR |
recordings |
Recording directory; Docker uses the shared /app/recordings volume |
PHONE_NUMBERS |
(empty) | Comma-separated list of phone numbers for Scenario 1 |
FLOW_STEPS |
step_home |
Comma-separated step names for the legacy step-registry runner |
Scenarios
Scenario 1 — SMS-OTP Login (crawler/scenarios/otp_login.py)
- Navigates to
TARGET_URL - Finds the phone-number input (CSS / XPath selector, configurable at top of file)
- Types the phone number character-by-character (human pacing)
- Clicks the "Send OTP" button
- Calls
otp_resolver(phone)in a polling loop (up to 12 × 5–8 s) — you must implement this to fetch the OTP from your SMS gateway / webhook - Types the OTP and confirms login
- Waits for a success indicator on the page
To wire up your SMS gateway, edit otp_resolver inside main.py → run_otp():
def otp_resolver(phone: str) -> str:
# Example: poll your own API
resp = httpx.get(f"https://sms-api.example.com/latest?phone={phone}", timeout=5)
return resp.json().get("code", "")
To update selectors to match the target site, edit the class-level constants at the top of OTPLoginScenario:
_PHONE_FIELD_SELECTOR = (By.CSS_SELECTOR, "input[name='mobile']")
_OTP_FIELD_SELECTOR = (By.CSS_SELECTOR, "input[name='code']")
_SUCCESS_INDICATOR = (By.CSS_SELECTOR, ".home-feed")
Scenario 2 — Fresh Session (crawler/scenarios/fresh_session.py)
- Calls
driver.delete_all_cookies() - Clears
localStorageandsessionStoragevia JavaScript - Navigates to
TARGET_URL - Waits for
document.readyState === 'complete'
Each invocation of this scenario gets a brand-new Chrome profile via undetected-chromedriver, so no state leaks between runs even without explicit clearing.
Dynamic Flow System
After either scenario completes login / session setup, DynamicFlow runs the configured steps.
All settings are edited live from the admin panel — no code changes or restarts needed.
How it works
For each search_text in config:
1. search_step → find search box → type text → submit → wait for results
2. For each target_item_id in config:
scroll_step → scroll until item selector matches (or max_scrolls hit)
click_step → scroll element into view → click it
Selector templates
The Item Selector Template field supports a {item_id} placeholder:
| Template | Item ID | Resolved selector |
|---|---|---|
[data-id='{item_id}'] |
prod-123 |
[data-id='prod-123'] |
.result[data-sku="{item_id}"] |
SKU-99 |
.result[data-sku="SKU-99"] |
.product-card |
(any) | .product-card (first visible match) |
Both CSS and XPath are supported for every selector field.
Infinite scroll behaviour
| Config field | Effect |
|---|---|
Scroll Container Selector |
Scrolls this element; blank → scrolls window |
Max Scrolls |
Hard limit — aborts after this many scroll iterations |
Scroll Pause (ms) |
Sleep between each scroll to let new content load |
No New Content Timeout (ms) |
If the DOM height hasn't grown for this long → stop and mark as failed |
Admin Panel
Accessible at http://localhost:8000 (credentials from .env).
Dashboard tab
| Widget | Shows |
|---|---|
| Stat cards | Total runs · Unique identifiers · Successes · Failures · Success rate |
| Scenario Breakdown | Bar chart: success rate per scenario (otp / fresh) |
| Top Failure Reasons | Most frequent error strings with counts |
| Recent Runs table | Last 50 runs with ID, time, scenario, identifier, status, duration, failure reason |
Auto-refreshes every 15 seconds.
Flow Configs tab
Create, edit, activate, and delete flow configurations.
Only one config is active at a time; the crawler reads it on every run.
Each config stores:
- Search texts (one per line)
- Search box selector + type (CSS / XPath)
- Optional submit button selector (falls back to Enter key)
- Results container selector (waited for after search)
- Scroll container selector, max scrolls, pause, no-new-content timeout
- Item selector template + type
- Target item IDs (one per line)
Makefile Reference
make install Bootstrap: create venv, install all deps
make sync Sync venv with lockfile (fast)
make lock Re-resolve and update uv.lock
make typecheck Run Pyright
make lint Alias for typecheck
make admin Start admin panel locally → http://localhost:8000
make otp Run OTP scenario (all phones from .env)
make otp PHONE=+98... Run OTP scenario for one phone
make fresh Run fresh-session scenario
make build Build Docker image
make up Start admin panel in Docker (detached)
make down Stop containers
make logs Follow container logs
make shell Shell into the running admin container
make docker-otp Run OTP scenario inside Docker
make docker-fresh Run fresh-session scenario inside Docker
make clean Remove .venv, __pycache__, .pyc files
Extending the Crawler
Add a new flow step (legacy registry)
# crawler/flow_runner.py
@register_step("step_add_to_cart")
def step_add_to_cart(driver: uc.Chrome) -> dict[str, object]:
btn = driver.find_element(By.CSS_SELECTOR, "button.add-to-cart")
btn.click()
return {"added": True}
Then set FLOW_STEPS=step_home,step_add_to_cart in .env.
These static steps run in addition to (after) the dynamic flow.
Add a new scenario
- Create
crawler/scenarios/my_scenario.pywith a class following the same pattern asOTPLoginScenario - Add a branch in
main.pyandMakefile
Use a remote WebDriver (Selenium Grid)
Replace make_driver() in crawler/driver.py with a webdriver.Remote(...) call pointing at your Grid hub.