Web scraping is the automated collection of data from websites using a program instead of copying it by hand. A script or API sends a request to a page, reads the content, and returns it in a structured format like JSON or CSV, instead of a person manually browsing and writing down what they see.
At its core, web scraping is three steps: send a request to a URL, receive the page back, and extract the pieces of data you actually need from it. Doing this by hand for one page takes a minute. Doing it for ten thousand pages by hand is not realistic: that's the entire reason scraping exists as a category of tooling.
A basic scraper can be a short script using a library like Python's requests plus an HTML parser. That works fine for simple, static pages. It breaks down fast against modern websites, which is where a managed web scraper API comes in, handling the parts that get complicated at scale: rotating IP addresses so you don't get blocked, rendering JavaScript for pages that build their content client-side, and solving anti-bot challenges automatically.
If a website offers an official API, that's usually the better path: it's sanctioned, stable, and won't break when the site redesigns. Scraping is what you reach for when no official API exists, or when the official API doesn't expose the specific data you need. Most of the public web falls into that second category.
Depending on the tool, scraped output can be raw HTML, cleaned-up Markdown, or fully structured JSON with named fields (price, title, rating, and so on). Structured JSON is generally the most useful for downstream work, since it skips the step of parsing HTML yourself.