> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hydrafetch.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Extract structured data from URLs

> Pull schema-shaped JSON out of one or many pages in a single call. An LLM maps each page onto your schema and/or prompt. Point at a single page, a list, or a crawl scope with a trailing `/*` wildcard; optionally let web search find extra source pages, return per-field confidence with the source passage behind each value, and merge everything into one deduplicated collection of entities. Charged per page that returns data.



## OpenAPI

````yaml https://api.hydrafetch.com/openapi.json post /v1/web/extract
openapi: 3.0.0
info:
  title: Hydrafetch API
  description: >-
    Hydrafetch is a web data API for developers and agents. Send a URL and get
    back clean Markdown, the page's own structured data, schema-shaped JSON,
    links, or a summary, with the navigation, banners and boilerplate stripped
    out. Scrape one page, crawl a whole site, run a search, or extract to a
    schema, all through one API with one response shape. Every call costs one
    credit a page whatever it took to fetch, and failures are never billed.


    Point us at a whole site and get every page. Ask a question and get answers
    with per-field

    confidence and the passage each value came from. You describe the outcome
    you want — the

    pipeline decides how to get it.


    ## Authentication


    Every request is authenticated with your API key in the `X-API-Key` header.
    Keys are scoped to a

    workspace and carry its credit balance.


    ## Credits


    Calls are billed in credits and charged only on success. A standard scrape
    is one credit; richer

    formats and the extraction tier cost more. Each response reports what it
    consumed.


    ## Conventions


    All timestamps are UTC ISO 8601. Long-running jobs (crawl, batch) return a
    job id you poll, or a

    webhook you register.


    ## Errors


    Every failure returns the same shape, whatever the status:


    ```json

    { "success": false, "error": { "code": "INSUFFICIENT_CREDITS", "message":
    "Insufficient credits" },
      "meta": { "requestId": "019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c" } }
    ```


    Branch on `error.code`, which is stable. `error.message` is written for a
    person and may be

    reworded without notice. Validation failures add `error.details`. Quote
    `meta.requestId` when

    asking us about a specific failure. Every operation documents the codes it
    can return.


    ## Rate limits


    Limits are per workspace, not per key, over a 60 second window, and the
    ceiling comes from your

    plan. Every response carries the IETF RateLimit header fields so you can
    self-throttle rather

    than discovering the limit by hitting it:


    ```http

    RateLimit-Policy: "workspace";q=600;w=60

    RateLimit: "workspace";r=599;t=42

    ```


    `q` is the quota, `w` the window in seconds, `r` the requests remaining and
    `t` the seconds until

    the window resets. A 429 also carries `Retry-After` in seconds; wait that
    long rather than

    retrying immediately. The older `X-RateLimit-Limit`, `X-RateLimit-Remaining`
    and

    `X-RateLimit-Reset` headers are still sent and mean the same thing.


    ## Versioning and deprecation


    The version is in the path: every endpoint on this API lives under `/v1/`.
    Within a version we

    only make additive changes — new endpoints, new optional parameters, new
    fields in a response.

    Adding a field to a response is not a breaking change, so parse defensively
    and ignore what you

    do not recognise.


    Anything that would break an existing integration ships under a new version
    path instead. When an

    endpoint or a version is retired, the responses say so before it stops
    working:


    ```http

    Deprecation: @1780272000

    Sunset: Sat, 01 May 2027 00:00:00 GMT

    Link: <https://docs.hydrafetch.com/changelog>; rel="deprecation"

    ```


    `Deprecation` (RFC 9745) marks when the endpoint became deprecated, `Sunset`
    (RFC 8594) when it

    stops responding, and the `deprecation` link relation points at what to move
    to. Nothing on this

    API is deprecated today, so you will not see these headers yet. Watch for
    them rather than for an

    announcement: the headers are the notice, and the linked changelog carries
    what to move to.
  version: '1.0'
  contact:
    name: Hydrafetch
    url: https://hydrafetch.com
    email: support@hydrafetch.com
servers:
  - url: https://api.hydrafetch.com
    description: Production
security:
  - apiKey: []
tags:
  - name: Scrape
    description: >-
      Turn one URL into clean, LLM-ready content. Ask for Markdown, HTML, links,
      or the page’s own structured data, and poll a job id when a fetch runs
      long.
  - name: Crawl & batch
    description: >-
      Whole sites rather than single pages. Map every URL, crawl with depth and
      path rules, batch a list you already have, and read the webhook deliveries
      for either.
  - name: Search
    description: >-
      Search the web and get the ranked results back already scraped, so an
      agent has content to cite rather than links to fetch.
  - name: Extract
    description: >-
      Pull schema-shaped JSON out of one or many pages, with optional per-field
      confidence and the source passage behind each value.
  - name: Brand
    description: >-
      Resolve a domain into the company ready to render: logos for light and
      dark, the palette ranked by how the site uses it, fonts, socials, and the
      design system behind them.
  - name: Media
    description: >-
      Images and screenshots from a page, with source and alt text, or a
      rendered capture of the page as it appears.
paths:
  /v1/web/extract:
    post:
      tags:
        - Extract
      summary: Extract structured data from URLs
      description: >-
        Pull schema-shaped JSON out of one or many pages in a single call. An
        LLM maps each page onto your schema and/or prompt. Point at a single
        page, a list, or a crawl scope with a trailing `/*` wildcard; optionally
        let web search find extra source pages, return per-field confidence with
        the source passage behind each value, and merge everything into one
        deduplicated collection of entities. Charged per page that returns data.
      operationId: run
      parameters: []
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ExtractRequestDto'
      responses:
        '200':
          description: The extracted data.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ExtractResponseDto'
        '400':
          description: The request body or query failed validation.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: VALIDATION_ERROR
                  message: url must be a valid URL
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '401':
          description: >-
            The X-API-Key header is missing, malformed, or names a key that no
            longer exists.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: UNAUTHORIZED
                  message: >-
                    Missing X-API-Key header. Agents:
                    https://hydrafetch.com/auth.md
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '402':
          description: >-
            The workspace has no credits left for this request. Nothing was
            charged and nothing was fetched.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: INSUFFICIENT_CREDITS
                  message: Insufficient credits
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '403':
          description: The key is valid but the workspace may not make this request.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: WORKSPACE_BANNED
                  message: >-
                    This workspace has been suspended and cannot make API
                    requests.
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '404':
          description: No such job, or the id belongs to another workspace.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: NOT_FOUND
                  message: Job not found
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '422':
          description: >-
            The target site did not respond, so there was nothing to return.
            This is a fact about that URL rather than a fault on our side, and
            retrying it will not change the answer. Nothing was charged.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: UPSTREAM_UNREACHABLE
                  message: The origin did not respond to any attempt we made.
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '429':
          description: >-
            The workspace exceeded its per-minute request limit. Read the
            RateLimit headers, or Retry-After on this response, and retry after
            the window resets.
          headers:
            Retry-After:
              description: >-
                Seconds to wait before retrying. Prefer this over a fixed
                backoff.
              schema:
                type: integer
                example: 17
            RateLimit:
              description: Requests remaining (r) and seconds until the window resets (t).
              schema:
                type: string
                example: '"workspace";r=0;t=17'
            RateLimit-Policy:
              description: >-
                The quota (q) for your plan and the window it applies over (w),
                in seconds.
              schema:
                type: string
                example: '"workspace";q=600;w=60'
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: TOO_MANY_REQUESTS
                  message: Rate limit exceeded for this workspace and plan.
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '500':
          description: Unexpected server error. Nothing was charged.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: INTERNAL_SERVER_ERROR
                  message: Internal server error
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '503':
          description: >-
            A dependency needed for the requested formats is unavailable. Safe
            to retry.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: SERVICE_UNAVAILABLE
                  message: The summary and json formats are not available.
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
components:
  schemas:
    ExtractRequestDto:
      type: object
      properties:
        urls:
          maxItems: 10
          example:
            - https://example.com/products/widget
            - https://example.com/products/*
          description: >-
            The pages to extract from. Each must be an http(s) URL. A trailing
            `/*` marks a crawl scope: every page discovered under that path is
            extracted and merged into the result.
          type: array
          items:
            type: string
        schema:
          type: object
          additionalProperties: true
          description: >-
            JSON Schema describing the shape you want back. Optional if `prompt`
            is given; when both are present the schema fixes the field names and
            types while the prompt guides what to pull.
        prompt:
          type: string
          maxLength: 2000
          example: Pull the product name, price in USD, and whether it is in stock.
          description: >-
            Natural-language instruction for what to extract. Use with or
            instead of a schema.
        preferStructure:
          type: boolean
          description: >-
            Preserve document structure (headings, lists, tables) over prose
            density when reading the page — good for listing and catalog pages.
            Default off.
        enableWebSearch:
          type: boolean
          description: >-
            Pull in extra source pages by web-searching your prompt, to fill
            fields your URLs do not cover. Requires a `prompt`.
        showSources:
          type: boolean
          description: >-
            Return the concrete list of URLs that were actually extracted, after
            any wildcard and web-search expansion. Default off.
        showConfidence:
          type: boolean
          description: >-
            For each field, return a confidence score and the exact source
            passage the value was drawn from. Default off.
        mergeEntities:
          type: boolean
          description: >-
            Merge the per-page results into one deduplicated collection — one
            row per entity, with its contributing source URLs — instead of a
            separate result per page. Default off.
        maxAge:
          type: number
          minimum: 0
          maximum: 604800000
          description: >-
            Reuse a recent capture of each page if it is younger than this many
            milliseconds. Omit or 0 to always fetch fresh. Capped at 7 days.
      required:
        - urls
    ExtractResponseDto:
      type: object
      properties:
        data:
          $ref: '#/components/schemas/ExtractResultDto'
      required:
        - data
    WebError:
      type: object
      required:
        - success
        - error
        - meta
      description: >-
        Every failure on this API returns this shape, whatever the status.
        Successful responses never carry it.
      properties:
        success:
          type: boolean
          enum:
            - false
          example: false
        error:
          $ref: '#/components/schemas/WebErrorBody'
        meta:
          $ref: '#/components/schemas/WebErrorMeta'
    ExtractResultDto:
      type: object
      properties:
        results:
          description: One result per extracted page.
          type: array
          items:
            $ref: '#/components/schemas/ExtractItemResultDto'
        sources:
          description: >-
            The concrete URLs actually extracted, after wildcard and web-search
            expansion. Present only when `showSources` is set.
          type: array
          items:
            type: string
        collection:
          description: >-
            The deduplicated collection, one row per entity. Present only when
            `mergeEntities` is set.
          type: array
          items:
            $ref: '#/components/schemas/MergedEntityDto'
      required:
        - results
    WebErrorBody:
      type: object
      required:
        - code
        - message
      properties:
        code:
          type: string
          enum:
            - VALIDATION_ERROR
            - INVALID_INPUT
            - UNAUTHORIZED
            - INSUFFICIENT_CREDITS
            - FORBIDDEN
            - WORKSPACE_BANNED
            - NOT_FOUND
            - TOO_MANY_REQUESTS
            - UPSTREAM_UNREACHABLE
            - SERVICE_UNAVAILABLE
            - INTERNAL_SERVER_ERROR
          description: >-
            Stable, machine-readable reason. Branch on this rather than on the
            message, which may be reworded.
          example: INSUFFICIENT_CREDITS
        message:
          type: string
          description: Human-readable explanation. Safe to show a user, not to parse.
          example: Insufficient credits
        details:
          type: object
          additionalProperties: true
          description: Optional machine-readable context, present on validation failures.
    WebErrorMeta:
      type: object
      properties:
        requestId:
          type: string
          description: Quote this when contacting support about a specific failure.
          example: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
    ExtractItemResultDto:
      type: object
      properties:
        url:
          type: string
          example: https://example.com/products/widget
          description: The page this result came from.
        data:
          type: object
          additionalProperties: true
          nullable: true
          description: >-
            The extracted data, shaped by your schema and/or prompt. Null when
            nothing matched.
        fields:
          type: object
          additionalProperties:
            $ref: '#/components/schemas/FieldProvenanceDto'
          description: >-
            Per-field confidence and source passage, keyed by field name.
            Present only when `showConfidence` is set.
        error:
          type: object
          nullable: true
          example: null
          description: Set when this page could not be extracted; null on success.
      required:
        - url
        - data
        - error
    MergedEntityDto:
      type: object
      properties:
        data:
          type: object
          additionalProperties: true
          description: >-
            The unioned data for one entity, filled across every contributing
            page.
        sources:
          description: Every source URL that contributed to this entity.
          type: array
          items:
            type: string
      required:
        - data
        - sources
    FieldProvenanceDto:
      type: object
      properties:
        confidence:
          type: number
          minimum: 0
          maximum: 1
          example: 0.92
          description: How certain the value is correct given the page, from 0 to 1.
        evidence:
          type: string
          example: Priced at $49.00 with free shipping.
          description: >-
            The exact short passage the value was drawn from, or an empty string
            if none.
      required:
        - confidence
        - evidence
  securitySchemes:
    apiKey:
      type: apiKey
      in: header
      name: X-API-Key

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.