How Do You Use the Web Scraper to Extract Text and Attributes from a URL?
Enter a public URL, type a CSS selector into the Selector field, then choose whether to scrape the text contents of matched nodes or an attribute from the last matched node. The tool returns JSON, and you can turn on Prettify to make that JSON easier to read. This works for any page the scraper can reach without a login, and it is best suited to quick, one-off extractions rather than large crawling jobs.
What the tool actually does
Web Scraper (web.scraper.workers.dev) is a simple scraper powered by Cloudflare Workers. Its interface is built around four controls described on the page:
- URL — the page to fetch.
- Selector — a CSS selector that matches the nodes you want.
- Scrape the text contents of matched nodes — returns the text inside every matched node.
- Scrape an attribute from the last matched node — returns one attribute value from the final match.
- Add a space between children of matched nodes — inserts spacing when a matched node contains multiple child elements.
- Attribute — the attribute name to pull when using attribute mode.
- Prettify the JSON output — formats the returned JSON with indentation.
There is no mention of pricing, accounts, or login on the page, so treat it as a lightweight utility and verify access conditions yourself before relying on it for anything recurring.
Step-by-step: extracting text
- Paste the target URL into the URL field. Use a full address, e.g.
https://example.com/blog. - Write a CSS selector in the Selector field. For example,
h2matches every second-level heading, andarticle pmatches paragraphs inside an article element. - Select "Scrape the text contents of matched nodes." The scraper will collect the visible text of each match.
- Enable "Add a space between children of matched nodes" if a match contains inline children (like
<span>or<a>tags) that would otherwise run together. - Turn on Prettify to get indented JSON instead of a single dense line.
- Run the scrape and read the output. Expect a JSON structure containing the extracted strings in document order.
Expected result: a JSON array or object of text values, one per matched node.
Step-by-step: extracting an attribute
- Enter the URL and selector as above, but pick a selector that targets the element carrying the attribute — for example
afor links orimgfor images. - Select "Scrape an attribute from the last matched node." Note that this mode returns only the last match, not all of them.
- Type the attribute name into the Attribute field, such as
href,src, ordata-id. - Enable Prettify and run the scrape.
Expected result: a JSON value containing the attribute string from the final matched node.
Choosing between the two modes
| Goal | Mode | What you get |
|---|---|---|
| Collect all visible text from many elements | Text contents of matched nodes | Text from every match |
| Pull one link, image source, or data attribute | Attribute from the last matched node | A single attribute value |
If you need an attribute from every match rather than just the last one, this tool's attribute mode will not do it — you would need a different approach.
Common snags and how to handle them
- Empty results. The selector probably matches nothing. Open the page's developer tools, inspect the element you want, and copy its class or tag into the Selector field.
- Text runs together. Turn on the space-between-children option; adjacent inline elements often concatenate without it.
- Only one value returned when you expected many. You are likely in attribute mode, which uses the last matched node only.
- Unreadable output. Enable Prettify; raw JSON is hard to scan.
- Page requires login or blocks automated requests. The scraper fetches the URL directly, so private or protected pages will not return the content you see in a browser.
When this tool fits — and when it doesn't
Use it for quick, public, single-page extractions where you already know the CSS selector and just want clean JSON back. Skip it when you need to crawl many pages, handle authentication, respect rate limits at scale, or extract attributes from every match. For those cases, a scripted scraper with a proper HTTP client and parsing library is the more reliable choice.