跳至内容

爬取 API

Crawl API for coordinated multi-page web collection.

Collect related eligible public pages as one bounded workflow, with source scope, collection rules, and results aligned to your pipeline.

  • Approved entry points The public source space to evaluate
  • Coverage criteria The page families that matter
  • Result requirements What usable means downstream
  • Operating cadence One-time or recurring business need

Collection fit

When one page request becomes a collection workflow.

Use Crawl API when related pages need to be treated as one bounded collection. Your team still defines what complete, correct, and usable means for its business decision.

01 · related page set

Start with the unit of work.

The requirement spans a related set of eligible public pages, not a queue of isolated URL requests your application wants to orchestrate itself.

02 · coordinated job

Give the collection one boundary.

Entry points, relevant page families, exclusions, and acceptance rules belong to one agreed collection brief.

03 · collection handoff

Review the result as a whole.

Judge coverage, source context, exceptions, and downstream usability against the agreed result contract—not page retrieval alone.

Crawl API is a fit whenthe bounded page collection is the unit your pipeline needs to receive and evaluate.Compare the alternatives

The crawl brief

Define the collection boundary before the first run.

Bring the source and business requirements that determine whether a collection is eligible, bounded, reviewable, and useful.

01

Source eligibility

Name the public sources, collection purpose, and policy requirements your team has reviewed.

02

Entry points

Bring representative domains, sections, or known pages that define where evaluation starts.

03

Coverage criteria

Describe the page families that count and the evidence your team uses to assess coverage.

04

Exclusions and stop conditions

Make off-limits areas, irrelevant paths, and business stop conditions explicit before collection.

05

Business cadence

Explain whether the need is one-time or recurring so the operating model can be evaluated honestly.

06

Quality checks

Define the checks your team will use to decide whether the collection is complete enough and usable.

Planning inputs, not a parameter list. The crawl brief turns source scope, boundaries, exclusions, and result needs into a workable collection plan.

From scope to collection

A crawl job should have a
clear beginning, boundary, and handoff.

The operating model begins with a buyer-approved brief and ends with a result your team can review against explicit acceptance criteria.

  1. 01 · define

    Frame the source space.

    Your team supplies the approved entry points, relevant page families, exclusions, intended use, and quality criteria.

    Crawl brief
  2. 02 · execute

    Run the bounded crawl.

    WebScrapingAPI operates the agreed collection path and returns results with source context.

    Managed collection
  3. 03 · review

    Validate the collection.

    Your team assesses coverage, exceptions, source context, and downstream usability against its quality criteria.

    Buyer review
Brief before build

A clear crawl brief keeps source scope, result shape, exception handling, and delivery aligned before your team builds around the workflow.

Define done

Specify the usable result before choosing a delivery format.

A multi-page collection is useful only when every result and exception can be interpreted inside the customer's downstream workflow.

Result unit

What is one page result?

Agree on the minimum page-level object your pipeline needs to inspect, process, or reject.

Provenance

How is source context retained?

Identify the source identity and collection context that must travel with each usable result.

Exceptions

How are skipped or failed pages represented?

Define which non-result states must remain visible so absence is not mistaken for evidence.

Handoff

Where must the collection go next?

Describe the system, team, and acceptance step that receive the collection after the job boundary.

Result planning

Align result structure, delivery route, and retention window with the workflow before the first production run.

Implementation plan

Define the job shape before your team builds around it.

Start with the operating decisions below so the integration is based on the source scope and result your workflow needs.

Access method and authenticationProvision securely

The production access surface and credential flow.

Input and scope schemaDefine scope

How entry points, boundaries, and collection requirements are represented.

Job lifecycleDefine states

The operating states and customer actions available at each state.

Result retrieval or deliveryChoose handoff

How the collection and its source context reach the receiving system.

Error and retry semanticsSet policy

Which exceptions are visible, who decides on another attempt, and how repeated work is treated.

Limits and retentionSet bounds

The scope, operating constraints, and result-availability window.

Operating ownership

WebScrapingAPI operates the crawl. Your team owns collection intent and review.

A written boundary prevents a web-access product from being mistaken for a fully operated data-delivery program.

网页抓取接口

管理爬执行

  • Collection and access execution for the scoped crawl
  • Job, result, and error surface for the workflow
  • Structured collection handoff
你的团队

意图,界限和审查

  • Source eligibility, purpose, entry points, exclusions, and stop criteria
  • Coverage and quality validation
  • Parsing and schema unless included in the crawl plan
  • Storage, retention, and downstream use
  • Source-change monitoring outside the crawl workflow
共同计划

爬行工作流程

  • Input and coverage behavior
  • Result and handoff behavior
  • Lifecycle, exception, and operating controls
  • Limits, cadence, pricing, support, and security requirements

Commercial fit

Size the collection you actually need.

Pricing depends on source scope, volume, cadence, result needs, and delivery model. Start with a representative collection so the plan matches the work.

你的爬行简报

带来工作量,而不是简单的页面数量.

Source set

Representative eligible entry points

Page families

What belongs in the collection

Expected breadth

The practical shape of the source space

Business cadence

One-time or recurring need

Result requirements

Context, exceptions, and acceptance

Receiving workflow

Where the collection goes next

FAQ

Crawl API questions for planning a collection.

These answers explain product fit, ownership, and how to prepare a useful crawl brief.

Bring us a representative source

What is Crawl API?

Crawl API is designed as the multi-page collection layer for a bounded set of related eligible public pages. Start by sharing the source space, collection boundary, and result your pipeline needs.

When should I use Crawl API instead of Scraper API?

Evaluate Crawl API when the unit of work is a related page set that should be handled as one bounded collection. Use Scraper API when your application already knows each URL and needs one managed page request at a time.

How is Crawl API different from Browser API?

Crawl API is evaluated around a scoped collection spanning related pages. Browser API is one documented browser-backed request where your application specifies the page state, supported interactions, context, and response.

How is Crawl API different from Data API?

Crawl API starts from an approved source space and a bounded multi-page collection need. Data API returns maintained structured records only for supported sources and schemas.

How is Crawl API different from Managed Data?

With Crawl API, your team owns collection intent, quality review, downstream validation, and the operating work outside the crawl job. Managed Data moves recurring collection, extraction, quality, maintenance, and delivery to WebScrapingAPI.

What sites and content can I crawl?

Eligibility and coverage depend on the public source, your intended use, and the collection boundary. Share representative entry points, page families, exclusions, and success criteria.

How is crawl scope defined?

Start with approved entry points, the page families that count, explicit exclusions, stop conditions, and the criteria your team will use to judge coverage and usability.

Which output formats and delivery options are supported?

The supported result representation depends on the collection model. Request a crawl brief to align result structure, source context, delivery, and retention with your workflow.

Can I schedule, monitor, cancel, retry, or resume a crawl?

Job-lifecycle controls depend on the workload. We help define the operating behavior, error treatment, and responsibilities before your team builds against the workflow.

How are pricing, limits, and service commitments determined?

Pricing depends on source scope, page volume, cadence, result needs, and delivery model. Share a representative collection to receive a workload-specific plan.

你的爬行简报

Turn a source space into a bounded collection plan.

Bring one representative domain, the pages that matter, the boundaries that must hold, and the result your pipeline needs. We will help choose the right product path.