Local Crawler
The ai12z Local Crawler is a command-line tool you run on your own computer. It crawls a website with a real (headless) Chrome browser, extracts page text, JSON-LD, and PDF content, and packages everything into a JSON file that you upload to ai12z with Add File.
It's built on the open-source Crawlee framework and runs on Mac, Windows, or in Docker.
The Local Crawler is distributed as installation ZIP files for Mac and Windows. To get them, contact support@ai12z.com.
When to Use the Local Crawler
For most public websites, Website ingestion is simpler: ai12z crawls the site for you. Use the Local Crawler when ai12z's cloud crawler can't reach the content:
- Sites that block crawlers: the site's bot protection blocks the ai12z crawler. Whitelisting or a connector solves this, but both need the site owner's IT team. When you're building a demo or proof of concept, you usually can't get IT involved until you have something to show.
- Vercel preview deployments: preview sites behind Vercel deployment protection. See Vercel Preview Deployments.
- Sites only your network can reach: the crawler runs on your computer, so it can crawl anything your browser can open, such as sites behind a VPN.
- Login-only content: pages that need a logged-in session, using cookies exported from your browser.
- Reviewing content first: you get a file you can inspect, filter, and clean before anything reaches ai12z.
Requirements
- Node.js 22.13 or newer (the 24 LTS version is recommended). To check, run
node --versionin Terminal (Mac) or Command Prompt (Windows). If it shows an error or an older version, install the LTS version from nodejs.org. - macOS, or Windows 10 or 11. Docker is also supported (see Docker).
- About 500 MB of downloads during installation, including the browser the crawler uses. Google Chrome does not need to be installed.
Installation
Extract the ZIP file you received, then run the installer for your platform.
Mac
- Open Terminal (press
Cmd + Space, type "Terminal", and press Enter). - Go to the extracted folder and run the installer:
cd ~/Downloads/crawlee-scraper-mac
./install-mac.sh
Windows
- Double-click
install-windows.batin the extracted folder. - If Windows shows an "Unknown Publisher" warning, click More info, then Run anyway.
The installer checks your Node.js version, installs the exact tested versions of the crawler's dependencies, and downloads the browser it uses. Each package also includes a README-FIRST.txt and a quick start guide.
Run all crawler commands from inside the extracted folder.
Crawl a Website
Start from one or more URLs, a sitemap, or a CSV file of URLs. Use exactly one of these:
# Start from a URL and follow links on the same site
node cli-crawler.mjs --urls "https://www.example.com/"
# Crawl every page in a sitemap
node cli-crawler.mjs --sitemap "https://www.example.com/sitemap.xml"
# Crawl a list of URLs from a CSV file (one URL per row)
node cli-crawler.mjs --csv urls.csv
Results are saved as JSON files in storage/datasets/default/. URLs that fail after all retries are listed in failed-urls.txt.
Common variations:
# Only English pages, skipping blog and discussion pages
node cli-crawler.mjs --urls "https://www.example.com/" --include /en/ --exclude blog discussion
# Only the listed pages, without following links
node cli-crawler.mjs --csv urls.csv --discover false
# Slow down for sites that limit request rates (milliseconds between pages)
node cli-crawler.mjs --urls "https://www.example.com/" --delay 5000
Each new crawl replaces the previous crawl's results. Create the upload file before you start the next crawl. If the crawler asks whether to clear an existing request queue, answer Y for a fresh crawl.
Password-Protected Sites
If the site shows the browser's "Sign in" dialog (HTTP Basic authentication), set the username and password before you run the crawler, in the same window.
Mac (Terminal):
export CRAWL_AUTH="username:password"
node cli-crawler.mjs --urls "https://staging.example.com/"
Windows (Command Prompt):
set "CRAWL_AUTH=username:password"
node cli-crawler.mjs --urls "https://staging.example.com/"
Windows (PowerShell):
$env:CRAWL_AUTH = "username:password"
node cli-crawler.mjs --urls "https://staging.example.com/"
The setting lasts until you close the window. To keep it, add the line CRAWL_AUTH=username:password to the .env file in the crawler folder instead, or pass --auth "username:password" with each command. When it's working, the crawler prints Using HTTP Basic auth as "username" at startup.
The credentials are only sent to the site being crawled, never to other sites it links to.
For a password-protected site, start from --urls or --csv and let the crawler follow links. A --sitemap is loaded without credentials.
Sites That Block the Crawler
If pages fail with Request blocked - received 403 status code, the site's bot protection is rejecting the crawler. Try identifying it as the ai12z crawler:
node cli-crawler.mjs --urls "https://www.example.com/" --ua "ai12zCopilot/1.0"
If the site's team has allowed the ai12zCopilot/1.0 user agent (see Whitelisting the ai12z Crawler), this matches that rule. Rules based on ai12z's IP address don't apply, because the Local Crawler connects from your computer's IP address.
PDFs
PDFs are included automatically, whether they're listed as start URLs or linked from crawled pages. The crawler extracts their text and keeps tables intact.
- To skip PDFs, add
--pdf false. - For PDF download links that don't end in
.pdf, put them in a CSV file and add--force-pdf. - PDFs with no extractable text, such as scanned images, are listed in
failed-urls.txtinstead of being uploaded empty.
Vercel Preview Deployments
For a site behind Vercel deployment protection, pass the project's protection bypass secret:
node cli-crawler.mjs --urls "https://your-preview.vercel.app/" --vercel-bypass "your-bypass-secret"
All Crawler Options
Separate several values with spaces, for example --include "/en/" "/blog/". Run node cli-crawler.mjs --help to see the same list.
| Option | What it does | Example |
|---|---|---|
--urls (-u) | One or more start URLs | --urls "https://example.com" |
--sitemap (-s) | Load start URLs from a sitemap | --sitemap "https://example.com/sitemap.xml" |
--csv | File with one URL per row | --csv "urls.csv" |
--include | Only crawl URLs containing one of these texts | --include "/blog/" |
--exclude | Skip URLs containing any of these texts | --exclude "/admin/" "/login" |
--discover | Follow links found on pages (default true) | --discover false |
--delay | Milliseconds to wait between pages (default 2000) | --delay 5000 |
--pdf | Extract PDFs, both start URLs and PDFs linked from pages (default true) | --pdf false |
--force-pdf | Treat every URL as a PDF download (for links without .pdf) | --force-pdf |
--auth | Username and password for a "Sign in" dialog (or set CRAWL_AUTH) | --auth "user:password" |
--user-agent (--ua) | Identify the crawler differently (helps with 403 blocks) | --ua "ai12zCopilot/1.0" |
--iquery | Ignore ?query strings when de-duplicating pages (default true) | --iquery false |
--cookies | Cookie file for login-protected sites (default cookie.json) | --cookies "my-cookies.json" |
--cdomain | Domain for cookies that don't specify one | --cdomain ".example.com" |
--vercel-bypass | Bypass secret for Vercel-protected preview sites | --vercel-bypass "secret" |
--vercel-bypass-cookie | Vercel bypass cookie mode (default true) | --vercel-bypass-cookie samesitenone |
--headless (-h) | Hide the browser window (default true) | --headless false |
Create the Upload File
When the crawl finishes, combine the results into a file for ai12z:
node merge-crawlee-dataset.js
This writes bulk-upload.json. If the content is larger than ai12z's 20 MB upload limit, it writes bulk-upload-part1.json, bulk-upload-part2.json, and so on instead. Each run first removes the upload files from the previous run, so the folder only holds files from the latest crawl.
When it asks whether to delete the storage directory, answer N unless you're finished with this crawl.
You can filter and clean the content as you merge:
| Option | What it does |
|---|---|
--include | Only include URLs containing one of these texts |
--exclude | Skip URLs containing any of these texts |
--remove | Remove these exact texts from every page, e.g. a repeated notice |
--substrings | Remove every line containing one of these texts, e.g. a cookie banner |
--include-jsonld | Include JSON-LD structured data (default true) |
node merge-crawlee-dataset.js --exclude /privacy /terms --substrings "We use cookies"
Upload to ai12z
- In ai12z, go to Documents and click Add Document.
- In the dialog that opens, choose Add File from the drop-down.
- Drag in
bulk-upload.json. If the merge created part files, upload eachbulk-upload-partN.jsonfile in turn.
Upload the .json file itself; ZIP files aren't accepted for JSON uploads.
The file uses the crawler's own format ("type": "ai12zCrawlee"), with the content, title, url, imageLinks, and jsonld of each page. See Bulk Content Upload JSON for how ai12z uses these fields.
Docker
For teams that prefer containers, each package includes a Dockerfile. From the extracted folder:
# Build the image
docker build -t crawlee-scraper .
# Crawl: results are written to ./storage
docker run --rm -it -v "$(pwd)/storage:/app/storage" crawlee-scraper --urls "https://www.example.com/"
# Create bulk-upload.json in the current folder
docker run --rm -it -v "$(pwd):/work" -w /work --entrypoint node crawlee-scraper /app/merge-crawlee-dataset.js
For a password-protected site, add -e CRAWL_AUTH="username:password" to the crawl command. On Windows, replace $(pwd) with the full path of the folder, for example C:\crawler.
Troubleshooting
- "node is not recognized" or "command not found: node": install Node.js 22.13 or newer from nodejs.org, then open a new window.
- "Could not find Chrome": the browser download during installation didn't complete. In the crawler folder, run
npx puppeteer browsers install chrome. - Every page fails with a 401 error: the site needs a username and password. See Password-Protected Sites.
- "Request blocked - received 403 status code": see Sites That Block the Crawler.
net::ERR_ABORTEDon a URL: the URL is a file download, usually a PDF link without.pdfin it. Use--force-pdffor those URLs.- Running out of memory on a large site: run
node --max-old-space-size=4096 cli-crawler.mjswith your usual options.
For anything else, contact support@ai12z.com.
License
The Local Crawler is licensed for use with the ai12z platform only. See the LICENSE file in the package for the full terms.