> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hydrafetch.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Scrape to Markdown

> Fetch a URL and return clean Markdown of its main content.



## OpenAPI

````yaml https://api.hydrafetch.com/openapi.json post /v1/web/markdown
openapi: 3.0.0
info:
  title: Hydrafetch API
  description: >-
    Hydrafetch is a web data API for developers and agents. Send a URL and get
    back clean Markdown, the page's own structured data, schema-shaped JSON,
    links, or a summary, with the navigation, banners and boilerplate stripped
    out. Scrape one page, crawl a whole site, run a search, or extract to a
    schema, all through one API with one response shape. Every call costs one
    credit a page whatever it took to fetch, and failures are never billed.


    Point us at a whole site and get every page. Ask a question and get answers
    with per-field

    confidence and the passage each value came from. You describe the outcome
    you want — the

    pipeline decides how to get it.


    ## Authentication


    Every request is authenticated with your API key in the `X-API-Key` header.
    Keys are scoped to a

    workspace and carry its credit balance.


    ## Credits


    Calls are billed in credits and charged only on success. A standard scrape
    is one credit; richer

    formats and the extraction tier cost more. Each response reports what it
    consumed.


    ## Conventions


    All timestamps are UTC ISO 8601. Long-running jobs (crawl, batch) return a
    job id you poll, or a

    webhook you register.


    ## Errors


    Every failure returns the same shape, whatever the status:


    ```json

    { "success": false, "error": { "code": "INSUFFICIENT_CREDITS", "message":
    "Insufficient credits" },
      "meta": { "requestId": "019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c" } }
    ```


    Branch on `error.code`, which is stable. `error.message` is written for a
    person and may be

    reworded without notice. Validation failures add `error.details`. Quote
    `meta.requestId` when

    asking us about a specific failure. Every operation documents the codes it
    can return.


    ## Rate limits


    Limits are per workspace, not per key, over a 60 second window, and the
    ceiling comes from your

    plan. Every response carries the IETF RateLimit header fields so you can
    self-throttle rather

    than discovering the limit by hitting it:


    ```http

    RateLimit-Policy: "workspace";q=600;w=60

    RateLimit: "workspace";r=599;t=42

    ```


    `q` is the quota, `w` the window in seconds, `r` the requests remaining and
    `t` the seconds until

    the window resets. A 429 also carries `Retry-After` in seconds; wait that
    long rather than

    retrying immediately. The older `X-RateLimit-Limit`, `X-RateLimit-Remaining`
    and

    `X-RateLimit-Reset` headers are still sent and mean the same thing.


    ## Versioning and deprecation


    The version is in the path: every endpoint on this API lives under `/v1/`.
    Within a version we

    only make additive changes — new endpoints, new optional parameters, new
    fields in a response.

    Adding a field to a response is not a breaking change, so parse defensively
    and ignore what you

    do not recognise.


    Anything that would break an existing integration ships under a new version
    path instead. When an

    endpoint or a version is retired, the responses say so before it stops
    working:


    ```http

    Deprecation: @1780272000

    Sunset: Sat, 01 May 2027 00:00:00 GMT

    Link: <https://docs.hydrafetch.com/changelog>; rel="deprecation"

    ```


    `Deprecation` (RFC 9745) marks when the endpoint became deprecated, `Sunset`
    (RFC 8594) when it

    stops responding, and the `deprecation` link relation points at what to move
    to. Nothing on this

    API is deprecated today, so you will not see these headers yet. Watch for
    them rather than for an

    announcement: the headers are the notice, and the linked changelog carries
    what to move to.
  version: '1.0'
  contact:
    name: Hydrafetch
    url: https://hydrafetch.com
    email: support@hydrafetch.com
servers:
  - url: https://api.hydrafetch.com
    description: Production
security:
  - apiKey: []
tags:
  - name: Scrape
    description: >-
      Turn one URL into clean, LLM-ready content. Ask for Markdown, HTML, links,
      or the page’s own structured data, and poll a job id when a fetch runs
      long.
  - name: Crawl & batch
    description: >-
      Whole sites rather than single pages. Map every URL, crawl with depth and
      path rules, batch a list you already have, and read the webhook deliveries
      for either.
  - name: Search
    description: >-
      Search the web and get the ranked results back already scraped, so an
      agent has content to cite rather than links to fetch.
  - name: Extract
    description: >-
      Pull schema-shaped JSON out of one or many pages, with optional per-field
      confidence and the source passage behind each value.
  - name: Brand
    description: >-
      Resolve a domain into the company ready to render: logos for light and
      dark, the palette ranked by how the site uses it, fonts, socials, and the
      design system behind them.
  - name: Media
    description: >-
      Images and screenshots from a page, with source and alt text, or a
      rendered capture of the page as it appears.
paths:
  /v1/web/markdown:
    post:
      tags:
        - Scrape
      summary: Scrape to Markdown
      description: Fetch a URL and return clean Markdown of its main content.
      operationId: markdown
      parameters: []
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ScrapeFormatRequestDto'
      responses:
        '200':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ScrapeResponseDto'
        '201':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ScrapeAcceptedDto'
        '400':
          description: The request body or query failed validation.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: VALIDATION_ERROR
                  message: url must be a valid URL
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '401':
          description: >-
            The X-API-Key header is missing, malformed, or names a key that no
            longer exists.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: UNAUTHORIZED
                  message: >-
                    Missing X-API-Key header. Agents:
                    https://hydrafetch.com/auth.md
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '402':
          description: >-
            The workspace has no credits left for this request. Nothing was
            charged and nothing was fetched.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: INSUFFICIENT_CREDITS
                  message: Insufficient credits
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '403':
          description: The key is valid but the workspace may not make this request.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: WORKSPACE_BANNED
                  message: >-
                    This workspace has been suspended and cannot make API
                    requests.
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '404':
          description: No such job, or the id belongs to another workspace.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: NOT_FOUND
                  message: Job not found
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '422':
          description: >-
            The target site did not respond, so there was nothing to return.
            This is a fact about that URL rather than a fault on our side, and
            retrying it will not change the answer. Nothing was charged.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: UPSTREAM_UNREACHABLE
                  message: The origin did not respond to any attempt we made.
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '429':
          description: >-
            The workspace exceeded its per-minute request limit. Read the
            RateLimit headers, or Retry-After on this response, and retry after
            the window resets.
          headers:
            Retry-After:
              description: >-
                Seconds to wait before retrying. Prefer this over a fixed
                backoff.
              schema:
                type: integer
                example: 17
            RateLimit:
              description: Requests remaining (r) and seconds until the window resets (t).
              schema:
                type: string
                example: '"workspace";r=0;t=17'
            RateLimit-Policy:
              description: >-
                The quota (q) for your plan and the window it applies over (w),
                in seconds.
              schema:
                type: string
                example: '"workspace";q=600;w=60'
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: TOO_MANY_REQUESTS
                  message: Rate limit exceeded for this workspace and plan.
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '500':
          description: Unexpected server error. Nothing was charged.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: INTERNAL_SERVER_ERROR
                  message: Internal server error
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
        '503':
          description: >-
            A dependency needed for the requested formats is unavailable. Safe
            to retry.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WebError'
              example:
                success: false
                error:
                  code: SERVICE_UNAVAILABLE
                  message: The summary and json formats are not available.
                meta:
                  requestId: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
components:
  schemas:
    ScrapeFormatRequestDto:
      type: object
      properties:
        url:
          type: string
          example: https://example.com
          description: The URL to scrape. Must be http(s).
        maxAge:
          type: number
          minimum: 0
          maximum: 604800000
          description: >-
            Serve from cache if a capture of this URL is younger than this many
            milliseconds. Omit for the default 24h window; 0 always fetches
            fresh. Capped at 7 days.
        cacheOnly:
          type: boolean
          description: >-
            Only serve from cache. If there is no fresh cached copy, return 404
            instead of fetching.
        storeInCache:
          type: boolean
          description: Persist the capture for later re-extraction. Default true.
        preferStructure:
          type: boolean
          description: >-
            Preserve document structure (headings, lists, tables) over prose
            density — good for marketing and service pages. Default off.
        async:
          type: boolean
          description: >-
            Return a job id immediately instead of waiting for the result. Poll
            GET /v1/web/scrape/{id}. Opt in only: a synchronous call never
            degrades to a job id on its own, so you get a page or an error, not
            two response shapes.
        onlyMainContent:
          type: boolean
          description: >-
            Return only the main content, dropping nav/boilerplate. Default
            true.
        includeTags:
          maxItems: 50
          example:
            - article
            - main
          description: >-
            CSS selectors to keep. When set, only matching elements are
            considered.
          type: array
          items:
            type: string
        excludeTags:
          maxItems: 50
          example:
            - .ad
            - '#comments'
          description: CSS selectors to strip before extraction.
          type: array
          items:
            type: string
        removeBase64Images:
          type: boolean
          description: Strip inline base64 images from the output. Default true.
        blockAds:
          type: boolean
          description: Remove common ad and tracking elements. Default true.
        includeLinks:
          type: boolean
          description: >-
            Keep inline links in the markdown. Default true — a bare `Read the
            guide` is worth less to a model than one it can follow. Turn off for
            the densest possible prose.
        waitFor:
          type: number
          minimum: 0
          maximum: 30000
          description: Extra milliseconds to let the page settle before capture.
        timeout:
          type: number
          minimum: 1000
          maximum: 100000
          description: Overall time budget for the request, in milliseconds.
        location:
          $ref: '#/components/schemas/LocationDto'
        headers:
          type: object
          additionalProperties:
            type: string
          description: Extra request headers to send when fetching the page.
        method:
          type: string
          enum:
            - GET
            - POST
          default: GET
          description: >-
            HTTP method to fetch the URL with. Use POST for endpoints that only
            answer a POST, such as JSON and GraphQL APIs. Defaults to GET.
        body:
          type: string
          example: '{"limit":25,"offset":0}'
          maxLength: 65536
          description: >-
            Request body, sent verbatim. Only used when `method` is POST. Set
            the content type yourself via `headers`; nothing is assumed about
            the format.
      required:
        - url
    ScrapeResponseDto:
      type: object
      properties:
        data:
          $ref: '#/components/schemas/WebScrapeDataDto'
      required:
        - data
    ScrapeAcceptedDto:
      type: object
      properties:
        jobId:
          type: string
          example: 019f3c09-6fae-740f-9257-10c2b6af7f43
          description: Poll this job id at GET /v1/web/scrape/{id}.
        status:
          type: string
          enum:
            - queued
          example: queued
          description: 'Only returned when the request asked for `async: true`.'
      required:
        - jobId
        - status
    WebError:
      type: object
      required:
        - success
        - error
        - meta
      description: >-
        Every failure on this API returns this shape, whatever the status.
        Successful responses never carry it.
      properties:
        success:
          type: boolean
          enum:
            - false
          example: false
        error:
          $ref: '#/components/schemas/WebErrorBody'
        meta:
          $ref: '#/components/schemas/WebErrorMeta'
    LocationDto:
      type: object
      properties:
        country:
          type: string
          example: us
          description: ISO 3166 alpha-2 country to fetch the page as if from.
        languages:
          example:
            - en-US
            - en
          description: Preferred content languages, most-preferred first.
          type: array
          items:
            type: string
    WebScrapeDataDto:
      type: object
      properties:
        url:
          type: string
          example: https://example.com
          description: The URL you requested.
        finalUrl:
          type: string
          example: https://example.com/
          description: The final URL after any redirects.
        redirected:
          type: boolean
          example: false
          description: >-
            Whether the origin redirected: true when `finalUrl` differs from the
            URL you requested. Explicit so a page with no redirect is
            distinguishable from one whose redirect was not tracked.
        status:
          type: number
          example: 200
          description: HTTP status of the fetched page.
        cached:
          type: boolean
          example: false
          description: Whether this result was served from cache.
        warning:
          type: string
          description: Set when the page was returned with a caveat (e.g. partial content).
        metadata:
          $ref: '#/components/schemas/WebScrapeMetadataDto'
        usage:
          $ref: '#/components/schemas/WebUsageDto'
        quality:
          description: Per-page extraction quality signals.
          allOf:
            - $ref: '#/components/schemas/WebQualityDto'
        markdown:
          type: string
          description: >-
            Clean Markdown of the main content. Returned when `markdown` is
            requested.
        html:
          type: string
          description: Cleaned main-content HTML. Returned when `html` is requested.
        rawHtml:
          type: string
          description: The unmodified page HTML. Returned when `rawHtml` is requested.
        links:
          description: Returned when `links` is requested.
          allOf:
            - $ref: '#/components/schemas/WebLinksDto'
        structured:
          description: Returned when `structured` is requested.
          allOf:
            - $ref: '#/components/schemas/WebStructuredDataDto'
        summary:
          type: string
          description: >-
            A concise factual summary. Returned when `summary` is requested
            (LLM-backed).
        json:
          type: object
          additionalProperties: true
          description: Schema-shaped JSON. Returned when `json` is requested (LLM-backed).
      required:
        - url
        - finalUrl
        - redirected
        - status
        - cached
        - metadata
    WebErrorBody:
      type: object
      required:
        - code
        - message
      properties:
        code:
          type: string
          enum:
            - VALIDATION_ERROR
            - INVALID_INPUT
            - UNAUTHORIZED
            - INSUFFICIENT_CREDITS
            - FORBIDDEN
            - WORKSPACE_BANNED
            - NOT_FOUND
            - TOO_MANY_REQUESTS
            - UPSTREAM_UNREACHABLE
            - SERVICE_UNAVAILABLE
            - INTERNAL_SERVER_ERROR
          description: >-
            Stable, machine-readable reason. Branch on this rather than on the
            message, which may be reworded.
          example: INSUFFICIENT_CREDITS
        message:
          type: string
          description: Human-readable explanation. Safe to show a user, not to parse.
          example: Insufficient credits
        details:
          type: object
          additionalProperties: true
          description: Optional machine-readable context, present on validation failures.
    WebErrorMeta:
      type: object
      properties:
        requestId:
          type: string
          description: Quote this when contacting support about a specific failure.
          example: 019e8a3c-9f0b-7c12-88ab-1d2e3f4a5b6c
    WebScrapeMetadataDto:
      type: object
      properties:
        title:
          type: object
          nullable: true
          example: Example Domain
          description: The page title.
        structure:
          type: string
          enum:
            - markdown
            - plain
          example: markdown
          description: >-
            Whether the content carries markdown structure (headings, lists,
            emphasis) or is plain prose. `plain` is not an error — the default
            extractor optimises for capturing every word and wins on listing and
            forum pages. If you need the structure, re-request with
            `preferStructure: true`.
        pageType:
          type: string
          example: article
          description: Coarse page classification (e.g. article, listing, forum, docs).
        wordCount:
          type: number
          example: 214
          description: Word count of the extracted main content.
        description:
          type: object
          nullable: true
          example: A short summary of the page, as published by the page itself.
          description: The page's own description (meta description / og:description).
        language:
          type: object
          nullable: true
          example: en
          description: The language the page declares.
        author:
          type: object
          nullable: true
          example: Jane Doe
          description: The declared author.
        siteName:
          type: object
          nullable: true
          example: Example Blog
          description: The declared site name.
        publishedTime:
          type: object
          nullable: true
          example: '2026-01-05'
          description: When the page says it was published (ISO 8601).
        image:
          type: object
          nullable: true
          example: https://example.com/cover.png
          description: The page's lead image (og:image).
      required:
        - title
        - structure
        - pageType
        - wordCount
        - description
        - language
        - author
        - siteName
        - publishedTime
        - image
    WebUsageDto:
      type: object
      properties:
        creditsUsed:
          type: number
          example: 1
          description: Credits this call consumed. Charged only on success.
        creditsRemaining:
          type: number
          example: 4999
          description: Credits left in your workspace's balance after this call.
        freshness:
          type: string
          enum:
            - cache
            - fresh
          example: fresh
          description: Whether the result was served from cache or freshly fetched.
      required:
        - creditsUsed
        - creditsRemaining
        - freshness
    WebQualityDto:
      type: object
      properties:
        confidence:
          type: number
          example: 0.94
          minimum: 0
          maximum: 1
          description: >-
            How trustworthy the extraction is, from 0 to 1. High when
            independent checks corroborate substantial content; near zero when
            the page yielded almost nothing.
        complete:
          type: boolean
          example: true
          description: >-
            Whether the result captured the bulk of the content available on the
            page. False when the result looks thin or truncated.
        blocked:
          type: boolean
          example: false
          description: Whether the page appeared to be behind a challenge or bot wall.
      required:
        - confidence
        - complete
        - blocked
    WebLinksDto:
      type: object
      properties:
        internal:
          description: Links pointing to the same site.
          type: array
          items:
            type: string
        external:
          description: Links pointing to other sites.
          type: array
          items:
            type: string
      required:
        - internal
        - external
    WebStructuredDataDto:
      type: object
      properties:
        entities:
          description: >-
            The page's own structured data, normalised into one deduplicated
            list of typed entities.
          type: array
          items:
            $ref: '#/components/schemas/WebStructuredEntityDto'
        jsonLd:
          description: Raw JSON-LD blocks, as found on the page.
          type: array
          items:
            type: object
        microdata:
          description: Raw microdata items.
          type: array
          items:
            type: object
        opengraph:
          description: Raw OpenGraph/Twitter-card tags.
          type: array
          items:
            type: object
        rdfa:
          description: Raw RDFa items.
          type: array
          items:
            type: object
        appState:
          description: Names of embedded framework app-state blocks detected on the page.
          type: array
          items:
            type: string
      required:
        - entities
        - jsonLd
        - microdata
        - opengraph
        - rdfa
        - appState
    WebStructuredEntityDto:
      type: object
      properties:
        type:
          type: string
          example: Product
          description: The schema.org type of the entity.
        source:
          type: string
          enum:
            - json-ld
            - microdata
            - rdfa
            - opengraph
          description: Which structured-data syntax the entity came from.
        properties:
          type: object
          additionalProperties: true
          description: The entity's properties, as published.
      required:
        - type
        - source
        - properties
  securitySchemes:
    apiKey:
      type: apiKey
      in: header
      name: X-API-Key

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.