Skip to content
On this page

Run a scrape

Step 2 of the two-step workflow. Loads a blueprint, follows pagination, extracts each item (and its detail page when configured), and prints/saves the items as JSON.

bash
php artisan datahelm:scrap:run <blueprint-path-or-host> [--limit=N] [--output=PATH]

The blueprint argument is either a file path, or a host name previously saved with --blueprint during generation (e.g. books.toscrape.com). --limit=N stops after N items.

Where output goes

By default the JSON is saved automatically to storage/app/scrapes/<name>.json, where <name> is derived from the command — e.g. datahelm:robot:exampleauctions writes storage/app/scrapes/exampleauctions.json (no flag needed). Progress (fetching: …) and the saved-path message go to stderr, so stdout stays clean for piping.

bash
# auto-saves to storage/app/scrapes/exampleauctions.json
php artisan datahelm:robot:exampleauctions --limit=20

# choose a different path
php artisan datahelm:robot:exampleauctions --limit=20 --output=storage/app/scrapes/today.json

# print to stdout instead of saving (e.g. to pipe)
php artisan datahelm:robot:exampleauctions --limit=20 --output=- > exampleauctions.json

--limit and --output are available on datahelm:scrap:run and on every generated robot command.

Full example

bash
php artisan datahelm:scrap:generate https://books.toscrape.com/ --get-detail=true --blueprint
php artisan datahelm:scrap:run books.toscrape.com --limit=20

Real, verified output for this exact example is in the books.toscrape.com tutorial.

Output formats

DataHelm architecture: databases, REST APIs, CSV files and website scraping feed the Navigation Engine, which passes unified data through the Data Refinement Engine into a clean data warehouse, structured API output, or dynamic dashboards

Set in the blueprint's output_config.format:

FormatDescription
jsonPretty-printed JSON array (default)
jsonlOne JSON object per line — better for large crawls, easy to stream
csvComma-separated; first row is headers; array values are JSON-encoded in their cell
markdownOne Markdown section per item — LLM ingestion / RAG. See Markdown output

flatten: true collapses nested arrays: saved_images[0]saved_images_0. You can also exclude_fields and rename_fields. See the Blueprint reference.

Streaming output

Write items to disk as they are scraped instead of buffering everything in memory — useful for large crawls (thousands of items):

json
"output_config": {
  "format": "jsonl",
  "stream": true
}

Works with all formats. When streaming is active, the robot writes each item immediately after it is scraped; the output file grows in real time and you can tail it while the crawl runs.

Crawl stats

After every crawl a summary is printed to stderr automatically — no configuration needed:

--- Crawl stats ---
  Items scraped : 24 in 8s (3.0/s)
  Pages         : 3 fetched, 0 failed
  Detail pages  : 24 fetched, 0 failed
  Images        : 48 saved, 0 failed
  Cache         : 0 hits, 27 misses

The --limit flag and streaming mode are both reflected correctly.

Throttling and caps

These keep crawls polite and bounded (all set in the blueprint):

  • crawl_config.delay_between_pages_ms / delay_between_items_ms — fixed delays.
  • crawl_config.max_items — blueprint-level hard cap (0 = no cap). The per-run --limit flag takes precedence when non-zero.
  • http_config.delay_ms — delay between page requests; for auction-style sites, 300–500 ms avoids rate limiting.
  • auto_throttle — dynamically adjusts the inter-page delay based on actual server latency (Scrapy's AutoThrottle). See the Blueprint reference.
  • cache — caches raw page HTML to disk so re-runs replay from cache instead of hitting the live site, ideal while iterating on selectors. See the Blueprint reference.

Resumable crawls

For a listing you crawl repeatedly (a daily robot over an auction site, say), most items were already scraped last time. Mark the blueprint resumable and the engine persists its dedup state between runs, so a re-run with --resume only processes new items:

bash
# 1. bake it in at generate time (sets "resumable": true in the blueprint)
php artisan datahelm:scrap:generate "https://example.com/listing" --resumable --robot-name=example

# 2. first run scrapes everything and saves the seen-keys state
php artisan datahelm:robot:example

# 3. subsequent runs skip every item whose dedup key was already seen
php artisan datahelm:robot:example --resume

How it works:

  • After each run of a resumable blueprint, the set of already-seen dedup key values (plus a total item count and timestamp) is saved as JSON to storage/app/scrapes/.state/{name}.json.
  • --resume loads that state before crawling, so previously-seen keys are skipped. Without the flag, the run starts fresh (but still re-saves state at the end).
  • Every scaffolded robot already carries the --resume flag in its signature.

Because resumption keys off dedup.key_field, make sure dedup is enabled with a key that is stable across runs (link usually is; a position-based field is not).

Reset by deleting the state file:

bash
rm storage/app/scrapes/.state/example.json

Tail the application log

In the Docker stack, use the container name from docker ps (php_datahelm), not the image name:

bash
docker exec -it php_datahelm tail -f /var/www/html/storage/logs/laravel.log

Next: Scaffold a robot →

Released under the MIT License.