Media asset
An image, video, audio item, or screenshot reference with a type and explicit file-delivery state.
- asset_id & type
- asset_url or file_ref
- format & dimensions
Images, video & audio data
Collect documented image and video search metadata, reverse-image results, Maps photo references, and browser screenshots—or scope a maintained multimodal dataset that keeps files, metadata, renditions, provenance, and derived labels distinct.
Multimodal object model
Separate the file or media binary from its observed metadata, source page, renditions, text, rights context, and derived labels. Each object can then change without corrupting the others.
An image, video, audio item, or screenshot reference with a type and explicit file-delivery state.
The public page or search result where the asset and its metadata were observed.
Alternate renditions, captions, transcripts, alt text, and derived labels remain linked records with their own provenance.
A discovered URL, caption, thumbnail, or delivered file does not establish ownership or permission to reuse it. Rights, licensing, attribution, and downstream use remain the customer's responsibility.
Coverage brief
Coverage describes the source, result or page family, request context, metadata fields, and whether file acquisition is included. It is never a blanket claim about media across the web.
Asset manifest
The schema tells consumers whether a record describes a media reference or a delivered file, where it came from, which rendition was observed, and which enrichment is derived.
{
"asset_id": "media_001",
"media_type": "image",
"file": {
"state": "not_requested",
"storage_ref": null
},
"metadata": {
"asset_url": "https://media.example/image.jpg",
"width": 1200,
"height": 800,
"caption": "Example caption"
},
"source": {
"page_url": "https://publisher.example/page",
"surface": "image_search"
},
"rights_state": "customer_review",
"observed_at": "YYYY-MM-DDThh:mm:ssZ",
"schema_version": "media.v1"
}Identity & relationships
Keep every source reference first. Exact file relationships and perceptual candidates are different claims and must remain distinguishable.
Source key
Exact evidence
Candidate evidence
Relationship state
Freshness & history
A source result, referenced media file, transcript, and derived label may be produced at different times. One generic “updated” field cannot describe them honestly.
The source, input, locale, and requested asset scope entered collection.
The page or result exposed the metadata or reference.
A media binary was delivered when file acquisition is in scope.
The source reference, rendition, or metadata entered tracked history.
Quality & missingness
Quality states distinguish source absence, reference-only scope, a failed acquisition, unsupported media, and an enrichment that has not run.
Required metadata, provenance, type, and state fields are present.
The page exposed a media reference; binary delivery was not requested.
A candidate relationship remains inspectable rather than silently merged.
Source absence, access failure, and out-of-scope media retain different reasons.
Required keys, types, URL shape, file-state consistency, and schema version.
Format, dimensions, duration, hash, decode state, or transcript presence only where contracted.
The customer approves source eligibility, rights workflow, label taxonomy, usefulness thresholds, and downstream use.
Operating model
Maximum control
Applications
Each application uses the same asset, source, rendition, time, and provenance foundation while the customer owns interpretation and permitted use.
Observe image and video result metadata for defined queries and markets.
query · result · asset reference · source · positionFind public media references for review without treating discovery as a rights conclusion.
asset · source page · caption · observed_atCompare expected media references, formats, and dimensions across public product surfaces.
entity cue · asset · rendition · quality stateTrack first-seen and last-seen observations for agreed public media surfaces.
source · asset · observation · historyScope linked files, metadata, text, and labels against an approved public-source brief.
file · metadata · transcript · label provenanceCapture rendered public pages with request context and observation time attached.
URL · viewport · screenshot · timestamp · stateRepresentative pilot
Sample ordinary and difficult assets, clarify file-versus-metadata delivery, test provenance and missingness, and confirm rights responsibilities before a recurring program.
Name public sources, media types, files, metadata, transformations, cadence, and exclusions.
Include missing files, alternate renditions, duplicate candidates, unsupported formats, and changed pages.
Review manifests, files where included, labels, provenance, timestamps, and collection states.
Agree schema, identity rules, quality checks, rights boundary, delivery, and maintenance response.
Evaluation FAQ
These answers distinguish documented access from pilot-first enrichment and keep file delivery separate from rights and interpretation.
Documented products can return metadata from Google image and video result surfaces, reverse-image results, Google Maps photo references, and browser screenshots. Broader media-file acquisition, audio extraction, transcripts, labels, deduplication, and training-ready datasets begin with a representative pilot.
The contract states this explicitly. A record can describe a source page, media URL, thumbnail, dimensions, format, caption, or rendition without delivering the underlying media binary. When file delivery is in scope, the file and its metadata remain separate, linked objects with their own states.
No. Metadata is not media rights. Copyright, licensing, attribution, permitted use, retention, and downstream distribution remain the customer's responsibility; WebScrapingAPI does not provide a copyright-clear guarantee.
The model keeps the media asset separate from the page or result where it appeared, its observed metadata, alternate renditions, captions or transcripts, source and rights context, and any derived label. Relationships and provenance are retained rather than flattened into one ambiguous row.
Private accounts, restricted libraries, direct messages, and other non-public media are not standard scope. Collection is limited to agreed eligible public web surfaces and the access context confirmed during scoping.
Labeling, media deduplication, transcript enrichment, and training-dataset assembly are pilot-first capabilities. A pilot defines the public sources, label taxonomy, file and metadata boundary, rights responsibilities, acceptance sample, and delivery format before production scope is agreed.
Source URLs and exact file hashes can support deterministic relationships when files are delivered. Perceptual matching across resized, cropped, re-encoded, or excerpted media is a pilot-first rule set, and ambiguous candidates stay visible instead of being silently merged.
The record distinguishes a value not exposed by the source, a rendition unavailable in the requested context, a file outside the contracted scope, a collection failure, and a field awaiting review. Missing values are not automatically treated as empty content.
Request-time documented APIs return an observation for the submitted context. Scheduled and managed programs use a source-specific cadence agreed after sampling; observed_at, first_seen, last_seen, and delivery timestamps remain distinct.
Proxy customers own their collectors, parsers, and monitoring. WebScrapingAPI maintains documented API behavior within the product boundary. In scheduled or managed delivery, WebScrapingAPI can own the contracted extraction, schema, quality checks, source-change maintenance, and delivery operation.
Images, video & audio data
Use this guide to define the object your system needs, then continue to Video Data for AI for scenario-focused clips, audio, transcripts, metadata, and recurring delivery.