跳至内容

数据集和数据源

Data Marketplace for ready-to-use web datasets.

Browse ready-to-use structured datasets by domain and source. Inspect the schema and representative sample records, then select the collection that fits your analysis.

  • Find the source Domain, entities, markets
  • Inspect the records Fields, types, missing states
  • Choose the timeline Snapshot and available history
  • Plan the handoff Format, partitions, destination

Dataset catalog

Browse datasets by source and record type.

Start with the entities your team needs, then open a dataset page to review fields, applications, refresh options, and a representative record.

01 · record

Start with the entity

Separate products from offers, companies from locations, and properties from listings before comparing fields.

Object and use case

02 · sample

Test fields with real context

Review identifiers, source context, timestamps, optional values, and missing states in representative rows.

Schema and joins

03 · next step

Choose the right delivery path

Select a prepared snapshot, or continue to Data Feeds when the collection needs a recurring schedule and destination.

Snapshot or recurring feed

Schema and fields

Understand every record before it enters your pipeline.

Use the data dictionary, source context, identifiers, timestamps, and missing states to verify that each field supports the joins, filters, and calculations your team will run.

样本记录插图式JSON
{
  "product_id": "wsa_20491",
  "title": "Trail shoe · blue · 42",
  "offer": {
    "price": 119.00,
    "currency": "EUR",
    "availability": "in_stock"
  },
  "source_url": "https://example.test/item",
  "captured_at": "2026-08-05T09:30:00Z",
  "schema_version": "commerce.v1"
}
A representative structure connecting product, offer, source, and collection time.
数据词典审查检查的内容
Entity identity
Stable identifiers, source identifiers, matching keys, and parent-child relationships.
Observation context
Source, market, requested context, capture time, and schema version.
Value semantics
Data type, unit, currency, normalized value, and source value where relevant.
Missing states
The difference between absent, unavailable, not observed, and not applicable.
Field presence
Reported presence across a representative sample, separated by source or record type.
Change handling
How fields, enumerations, and schema versions are communicated across deliveries.
采样接受
  1. Can the record be joined?Identity and keys fit the destination model.
  2. Can absence be interpreted?Missing states do not become false evidence.
  3. Can change be managed?Schema versions and delivery dates remain explicit.

Freshness and history

Choose what should follow dataset discovery.

Use a current snapshot for a baseline, available history for change analysis, or move the selected collection into Data Feeds for recurring scheduled delivery.

快速拍摄

建立一个时间点的基准.

Use a dated collection for one-time analysis, model preparation, market mapping, or a controlled initial load.

Inspect
Observation window and source mix
Plan
Full initial delivery
常见的 数据馈送

保持所选范围的更新.

Move to 数据馈送 to define a recurring schedule, schema, and destination for structured delivery.

Inspect
Cadence and change pattern
Plan
Scheduled delivery
历史深度

让最新的记录与相关.

Where history is available, align the requested time window and grain with trend analysis, backtesting, or change detection.

Inspect
Available dates and continuity
Plan
Backfill plus future refreshes
01ObserveSource and capture window02ValidateSchema and quality checks03VersionDated delivery and manifest04DeliverFull or incremental handoff

Delivery design

Receive data in the shape your platform can operate.

Select JSON, CSV, or Parquet for your application, warehouse, or data lake. Clear file organization, manifests, schema versions, and refresh behavior keep each handoff predictable.

其他类型

保存嵌入式对象和记录应用程序或文件导向工作流的文本.

Nested records

公司的资产

提供对分析工具,仓库和受控进口的熟悉表格交付.

Flat tables

支持输入,用于更大的分析工作负载和数据湖摄入的列处理.

Columnar files

Samples and pricing

Price the records and delivery you need.

Dataset pricing reflects the selected scope, history, refresh schedule, format, and delivery route. Start with a relevant collection and sample, then size the production data package.

数据集样本

在你量度交付之前,检查记录.

运行代表行和数据词典,通过您的生产堆所使用的合并,过器,计算和质量检查.

  • Representative source and entity mix
  • Sample record and data dictionary
  • Field presence and missing states
  • Snapshot date and available historical scope
  • Format and receiving environment
Request a dataset sample

FAQ

Data Marketplace questions.

See how catalog selection and samples connect to Data Feeds, Data API, Web Archive, and Managed Data.

Match my data brief

What is the Data Marketplace?

The Data Marketplace is a catalog of ready-to-use structured web datasets. Browse by domain and source, then inspect fields and representative sample records before choosing a prepared collection.

How do I find a relevant dataset?

Start with the entity and source, then narrow by markets, fields, historical scope, and format. Use the category pages to compare adjacent collections or send us the brief for help choosing.

Can I inspect the data before choosing a dataset?

Yes. Request representative records with the data dictionary and source context, then run the joins, filters, calculations, and quality checks your production stack will use.

Can a dataset be filtered to my scope?

Start with the countries, categories, entities, dates, sources, and fields you need. We will map that scope to the relevant collection and delivery structure.

What if I need data on a recurring schedule?

Use Data Marketplace to find and sample the right prepared collection. Choose 数据馈送 when the workload needs scheduled structured delivery with a defined cadence and destination.

Which delivery formats are available?

Choose JSON, CSV, or Parquet for the selected collection. Delivery can also define partitions, compression, file naming, manifests, and the destination used by your data stack.

How is Data Marketplace different from Data API?

Choose Data Marketplace to discover and sample a prepared collection. Choose Data API when your application requests source-specific structured results on demand, or 数据馈送 for recurring scheduled delivery.

When should I choose Managed Data instead?

Choose Managed Data when the required sources, schema, matching, quality rules, cadence, or destination need an operated program beyond the available catalog. WebScrapingAPI then manages collection, extraction, maintenance, monitoring, and delivery to your defined requirements.

Your dataset brief

Turn your data brief into a usable collection.

Tell us the domain, entities, markets, fields, and dates. We will connect the brief to a relevant prepared collection—or to Data Feeds when you need recurring delivery.