> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hydrafetch.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Crawl

> Discover and scrape a whole site as one asynchronous job.

Crawl points Hydrafetch at a starting URL, discovers the site's pages for you, and scrapes each one. It runs as a single asynchronous job: you get a `crawlId` back immediately, then poll it or register a webhook to collect the per-page results. One credit per page scraped.

## When to use

* You want every page of a site (or a section of it) as clean data, and you do not have the URL list yourself.
* The job is large enough that waiting on a single request is impractical.

If you already have the exact URLs, use [Batch](/endpoints/batch) — no discovery needed. To preview which URLs a crawl would reach without scraping them, use [Map](/endpoints/map).

## Example request

Send a `POST` to `/v1/web/crawl`. Scope the crawl with `limit`, `maxDepth`, and path filters, and control how each page is scraped with `scrapeOptions`.

<CodeGroup>
  ```bash cURL theme={null}
  curl -X POST https://api.hydrafetch.com/v1/web/crawl \
    -H "X-API-Key: hf_your_key_here" \
    -H "Content-Type: application/json" \
    -d '{
      "url": "https://example.com/blog",
      "limit": 100,
      "maxDepth": 2,
      "includePaths": ["^/blog"],
      "excludePaths": ["^/tag"],
      "scrapeOptions": { "formats": ["markdown", "links"] }
    }'
  ```

  ```javascript Node theme={null}
  const res = await fetch("https://api.hydrafetch.com/v1/web/crawl", {
    method: "POST",
    headers: {
      "X-API-Key": "hf_your_key_here",
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      url: "https://example.com/blog",
      limit: 100,
      maxDepth: 2,
      includePaths: ["^/blog"],
      excludePaths: ["^/tag"],
      scrapeOptions: { formats: ["markdown", "links"] },
    }),
  });
  const { crawlId } = await res.json();
  ```

  ```python Python theme={null}
  import requests

  res = requests.post(
      "https://api.hydrafetch.com/v1/web/crawl",
      headers={"X-API-Key": "hf_your_key_here"},
      json={
          "url": "https://example.com/blog",
          "limit": 100,
          "maxDepth": 2,
          "includePaths": ["^/blog"],
          "excludePaths": ["^/tag"],
          "scrapeOptions": {"formats": ["markdown", "links"]},
      },
  )
  crawl_id = res.json()["crawlId"]
  ```
</CodeGroup>

## Example response

The crawl is accepted right away:

```json theme={null}
{ "crawlId": "019f3c09-6fae-740f-9257-10c2b6af7f43", "status": "queued" }
```

## Keeping a crawl inside one section

A starting URL says where to begin, not where to stay. Discovery also uses the site's own published page list, which covers everything, so pointing a crawl at `https://example.com/blog` and setting nothing else will happily return pages from the rest of the site.

Add `includePaths` to hold it inside the section, and write the pattern so that the starting page matches it too:

```json theme={null}
{ "url": "https://example.com/blog", "includePaths": ["^/blog"] }
```

`["^/blog/.*"]` looks equivalent and is not. A starting URL loses its trailing slash, so `https://example.com/blog/` starts at the path `/blog`, which that pattern does not match. The crawl would then have no page to begin from, so we refuse the request instead of handing back a job that finishes with nothing in it.

## Request options

<ParamField body="url" type="string" required>
  The site to start from. Must be `http(s)`.
</ParamField>

<ParamField body="limit" type="number" default="100">
  Maximum number of pages to scrape. 1–500.
</ParamField>

<ParamField body="maxDepth" type="number">
  How many links deep from the starting page to follow. 0–10.
</ParamField>

<ParamField body="includePaths" type="string[]">
  Only follow URLs whose path matches every one of these patterns (e.g. `["^/blog"]`). Up to 50. **The patterns apply to the starting URL as well.** If they exclude it, the crawl has nowhere to begin, so the request is refused rather than returning an empty job.
</ParamField>

<ParamField body="excludePaths" type="string[]">
  Skip URLs whose path matches any of these patterns (e.g. `["^/tag"]`). Up to 50. As with `includePaths`, these apply to the starting URL too.
</ParamField>

<ParamField body="allowSubdomains" type="boolean" default="false">
  Also follow links into subdomains of the starting site.
</ParamField>

<ParamField body="allowExternalLinks" type="boolean" default="false">
  Also follow links that lead off the starting site.
</ParamField>

<ParamField body="ignoreQueryParameters" type="boolean" default="false">
  Treat URLs that differ only by query string as the same page.
</ParamField>

<ParamField body="sitemap" type="string" default="include">
  Whether to seed discovery from the site's published page list: `skip` or `include`. Because that list covers the whole site, **starting from a section does not by itself limit the crawl to that section.** Use `includePaths` to do that.
</ParamField>

<ParamField body="scrapeOptions" type="object">
  How to scrape each page — `formats`, `onlyMainContent`, `includeTags`, `excludeTags`, `removeBase64Images`, `blockAds`, `waitFor`, `timeout`, `location`, `headers`, `preferStructure`, `maxAge`. Same options as a single [Scrape](/endpoints/scrape), with per-page formats limited to `markdown`, `html`, `rawHtml`, `links`, and `structured`.
</ParamField>

<ParamField body="webhook" type="object">
  Register a callback instead of polling. `webhook.url` receives progress and completion events; `webhook.headers` are extra headers sent with each callback (e.g. for authentication).
</ParamField>

## Poll for results

Poll `GET /v1/web/crawl/{id}` for progress and per-page results.

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.hydrafetch.com/v1/web/crawl/019f3c09-6fae-740f-9257-10c2b6af7f43 \
    -H "X-API-Key: hf_your_key_here"
  ```

  ```javascript Node theme={null}
  const res = await fetch(
    `https://api.hydrafetch.com/v1/web/crawl/${crawlId}`,
    { headers: { "X-API-Key": "hf_your_key_here" } },
  );
  const status = await res.json();
  ```

  ```python Python theme={null}
  import requests

  res = requests.get(
      f"https://api.hydrafetch.com/v1/web/crawl/{crawl_id}",
      headers={"X-API-Key": "hf_your_key_here"},
  )
  status = res.json()
  ```
</CodeGroup>

```json theme={null}
{
  "id": "019f3c09-6fae-740f-9257-10c2b6af7f43",
  "kind": "crawl",
  "status": "running",
  "seedUrl": "https://example.com",
  "total": 100,
  "completed": 42,
  "failed": 1,
  "creditsUsed": 42,
  "pages": [
    {
      "url": "https://example.com/blog/post",
      "status": "completed",
      "depth": 1,
      "discoveredVia": "sitemap",
      "error": null,
      "data": {
        "url": "https://example.com/blog/post",
        "finalUrl": "https://example.com/blog/post",
        "status": 200,
        "cached": false,
        "metadata": { "title": "A post", "pageType": "article", "wordCount": 812 },
        "usage": { "creditsUsed": 1, "creditsRemaining": 4993, "freshness": "fresh" },
        "markdown": "# A post\n\n..."
      }
    }
  ]
}
```

## Response fields

<ResponseField name="id" type="string">The crawl id.</ResponseField>
<ResponseField name="kind" type="string">`crawl` or `batch`.</ResponseField>
<ResponseField name="status" type="string">Overall job state: `running`, `completed`, `failed`, or `cancelled`.</ResponseField>
<ResponseField name="seedUrl" type="string">The starting URL.</ResponseField>
<ResponseField name="total" type="number">Total pages in this job.</ResponseField>
<ResponseField name="completed" type="number">Pages scraped so far.</ResponseField>
<ResponseField name="failed" type="number">Pages that failed.</ResponseField>
<ResponseField name="creditsUsed" type="number">Credits consumed so far. One per scraped page.</ResponseField>

<ResponseField name="pages" type="object[]">
  Per-page results, each with `url`, `status` (`queued`, `running`, `completed`, `failed`), `depth`, `discoveredVia`, `error`, and `data` (the scraped page, once completed).

  `discoveredVia` tells you how a URL entered the job: `requested` for one you gave us (the crawl seed, or any batch entry), `sitemap` for one read from the site's sitemap, `link` for one followed from a page we had already fetched. It is the quickest way to see why a crawl returned the set it did, and whether raising the limit or the depth would find more.
</ResponseField>

<Note>
  A crawl costs one credit per page scraped, reflected in `creditsUsed`. Scope the job with `limit`, `maxDepth`, and path filters to keep spend predictable, and preview reach first with [Map](/endpoints/map).
</Note>

## Next steps

<CardGroup cols={2}>
  <Card title="Crawl API reference" icon="https://mintcdn.com/hydrafetch/koPVMLXM3S4OTC3p/icons/code.svg?fit=max&auto=format&n=koPVMLXM3S4OTC3p&q=85&s=ce116045b848b9f082595f6161ed7383" href="/api-reference" width="18" height="18" data-path="icons/code.svg">
    Full request and response schema with a live playground.
  </Card>

  <Card title="Map a site" icon="https://mintcdn.com/hydrafetch/DnAi7n_kype0jB2E/icons/sitemap.svg?fit=max&auto=format&n=DnAi7n_kype0jB2E&q=85&s=cffc277c9028dc4743694521f28f8088" href="/endpoints/map" width="18" height="18" data-path="icons/sitemap.svg">
    Preview a crawl's scope without scraping.
  </Card>

  <Card title="Scrape a list of URLs" icon="https://mintcdn.com/hydrafetch/DnAi7n_kype0jB2E/icons/bullet-list.svg?fit=max&auto=format&n=DnAi7n_kype0jB2E&q=85&s=879895a11dd7930d967e781b595fd992" href="/endpoints/batch" width="18" height="18" data-path="icons/bullet-list.svg">
    When you already have the URLs.
  </Card>

  <Card title="Scrape one URL" icon="https://mintcdn.com/hydrafetch/DnAi7n_kype0jB2E/icons/file-content.svg?fit=max&auto=format&n=DnAi7n_kype0jB2E&q=85&s=478eea6173d8c82bf8c41ab7277656bf" href="/endpoints/scrape" width="18" height="18" data-path="icons/file-content.svg">
    The per-page primitive behind a crawl.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.