Web scraping 101 in one line: a program downloads web pages and copies the pieces you care about, such as prices, titles or reviews, into a spreadsheet or database. This guide is web scraping for beginners. It covers how scraping works, manual versus automated scraping, the HTML and CSS you need, a working Python scraper you can run today, pages built with JavaScript, the rules to follow, and when rotating residential proxies start to matter.
The short version
Web scraping is the automated collection of data from websites: fetch a page, parse its HTML, pick out the elements you need with selectors, and save them in a structured format. Beginners can start with Python, Requests and Beautiful Soup, move to Playwright for pages built by JavaScript, and add proxies once a project needs many requests, several countries or steady access to sites that limit traffic per IP address.
Web Scraping 101: What Is Web Scraping?
Turning web pages into dataA web page is written for people. The price of a product sits inside a styled box, next to a picture, under a menu. Web scraping is the process of pulling that price out and putting it somewhere a computer can use it, like a row in a CSV file with columns for the product name, price, rating and link.
You can do it by hand, by copying and pasting, and for ten products that is fine. For ten thousand products, or for the same ten products every hour, you write a program instead. That program is called a scraper. It requests the page like a browser would, reads the HTML that comes back, finds the elements that hold the data, and saves them.
Two related terms come up constantly:
- Crawling is discovering pages: start from a URL, follow links, build a list of pages to visit. Search engines crawl.
- Scraping is extracting data from the pages you visit. Most real projects do both: crawl the category pages to find product URLs, then scrape each product page.
Key takeaways
- Scraping has four steps: request, parse, extract, store.
- If the data is in the page’s HTML, Requests and Beautiful Soup are enough. If JavaScript builds the page, use a browser tool such as Playwright, or find the JSON the page loads.
- Check the site’s robots.txt and terms, go slowly, and never collect more personal data than you need.
- Blocks usually come from too many requests per IP address or from the wrong location. That is where proxies come in.
How Web Scraping Works
Request, parse, extract, storeEvery scraper, from a 20-line script to a crawler fetching millions of pages, runs the same loop:
list of URLs
|
v
+--------------+ HTTP GET +---------------+
| scraper | --------------> | web server |
| | <-------------- | |
+--------------+ HTML / JSON +---------------+
|
| 1. parse the HTML into a tree
| 2. select elements (CSS selectors / XPath)
| 3. clean the values (trim, convert prices)
v
+--------------+
| CSV / JSON / | next page? --> back to the top
| database |
+--------------+
- Request. The scraper sends an HTTP GET request for a URL, the same request a browser sends when you type an address.
- Receive. The server returns a status code (200 means OK, 404 not found, 429 too many requests) and the page body, usually HTML.
- Parse. A parser turns the HTML text into a tree of elements you can search.
- Extract. Selectors point at the elements that hold the data: “the text of every
pwith the classprice_color“. - Store. The values are cleaned and written to a file or database.
- Repeat. The scraper follows the “next page” link or takes the next URL from its list, with a pause between requests.
Types of Web Scraping: Manual vs Automated
Copy-paste, no-code tools and codeThere are two main categories, and the second splits into a few approaches.
Manual web scraping means a person collects the data: opening pages, copying values into a spreadsheet, maybe using the browser’s developer tools to copy a table. It needs no setup and it is the right choice for a one-off list of a few dozen items. It is slow, it doesn’t repeat itself, and people make copy errors.
Automated web scraping means software does the collecting. It is faster, it can run on a schedule, and it gives the same result every time. It comes in several forms:
| Manual | Automated | |
|---|---|---|
| Setup | None | Minutes to days |
| Speed | A few pages a minute | Many pages a second, if the site allows it |
| Repeatable | No | Yes, on a schedule |
| Errors | Copy mistakes | Breaks when the page layout changes |
| Best for | One-off, small lists | Ongoing monitoring, large datasets |
What Is Web Scraping Used For?
The jobs that pay for itMost commercial scraping is about public information that changes often and matters to a business decision:
What You Need Before You Start
A small, free toolkitFor web scraping 101 you don’t need much. Python is the most common language for scraping because its libraries are simple and well documented, though JavaScript (Node.js) works just as well. This guide uses Python.
| Tool | What it does | When you need it |
|---|---|---|
| Browser developer tools | Inspect the HTML of any element, watch network requests | Always. It is how you find selectors. |
| Python 3 | Runs your scraper | Always |
| Requests | Sends HTTP requests, handles sessions, headers and proxies | Pages whose data is in the HTML |
| Beautiful Soup | Parses HTML and finds elements with CSS selectors | With Requests |
| Playwright | Drives a real Chromium, Firefox or WebKit browser | Pages built by JavaScript, clicks, scrolling |
| Scrapy | Full crawling framework with queues, retries and throttling | Thousands of pages or more |
Install the two libraries for the first scraper with one command:
python -m pip install requests beautifulsoup4
Understanding HTML and CSS Selectors
How to point at the dataHTML is a tree of nested elements. Each element has a tag name (div, p, a), and most have attributes such as class, id or href. Here is one product from books.toscrape.com, a practice site built for learning to scrape:
<article class="product_pod">
<p class="star-rating Three">...</p>
<h3>
<a href="catalogue/a-light-in-the-attic_1000/index.html"
title="A Light in the Attic">A Light in the ...</a>
</h3>
<div class="product_price">
<p class="price_color">£51.77</p>
<p class="instock availability">In stock</p>
</div>
</article>
A CSS selector is a short pattern that matches elements. The same syntax styles web pages, so the browser’s developer tools can test it for you: right-click an element, choose Inspect, and use Ctrl+F in the Elements panel to try a selector.
| Selector | Matches | In the example |
|---|---|---|
article.product_pod | Every article with class product_pod | Each product card |
h3 a | An a anywhere inside an h3 | The title link |
p.price_color | A p with class price_color | £51.77 |
p.star-rating | The rating paragraph; the second class holds the value | Three |
li.next a | The link inside the “next” list item | The next page URL |
Two habits save a lot of trouble. First, read attributes when the visible text is cut short: the title text above ends in “…”, but the title attribute has the full name. Second, prefer stable hooks such as meaningful class names or data- attributes over position (“the third div”), because layouts change.
How to Scrape a Website: Your First Scraper
A working Python example, step by stepThis is the practical core of web scraping 101. The scraper below collects the title, price, rating and link of every book on the first three catalogue pages of books.toscrape.com and saves them to books.csv. It waits one second between pages, retries politely when the server says it is busy, and can send its traffic through a proxy if you set one. We ran it on 26 September 2026 and it saved 60 books.
import csv
import os
import time
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
START_URL = "https://books.toscrape.com/catalogue/page-1.html"
HEADERS = {"User-Agent": "books-tutorial/1.0 (+https://example.com/contact)"}
# Optional: route requests through a proxy, e.g. http://user:pass@host:port
proxy = os.environ.get("PROXY_URL")
PROXIES = {"http": proxy, "https": proxy} if proxy else None
def fetch(session, url, tries=4):
"""GET a page, backing off on 429 and 5xx responses."""
for attempt in range(tries):
r = session.get(url, headers=HEADERS, proxies=PROXIES, timeout=20)
if r.status_code == 200:
return r.content # bytes: BeautifulSoup reads the charset itself
if r.status_code == 429 or r.status_code >= 500:
wait = int(r.headers.get("Retry-After", 2 ** attempt))
time.sleep(wait)
continue
r.raise_for_status()
raise RuntimeError(f"gave up on {url}")
def parse(html, page_url):
soup = BeautifulSoup(html, "html.parser")
rows = []
for card in soup.select("article.product_pod"):
link = card.select_one("h3 a")
rows.append({
"title": link["title"],
"price": card.select_one("p.price_color").get_text(strip=True),
"rating": card.select_one("p.star-rating")["class"][1],
"url": urljoin(page_url, link["href"]),
})
nxt = soup.select_one("li.next a")
return rows, (urljoin(page_url, nxt["href"]) if nxt else None)
def main(max_pages=3):
url, books = START_URL, []
with requests.Session() as s:
for _ in range(max_pages):
rows, url = parse(fetch(s, url), url)
books.extend(rows)
if not url:
break
time.sleep(1) # be polite: one request per second
with open("books.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=["title", "price", "rating", "url"])
w.writeheader()
w.writerows(books)
print(f"saved {len(books)} books")
if __name__ == "__main__":
main()
What each part does:
HEADERSgives your scraper an honest name and a way to contact you. Many sites treat the default library user agent with suspicion, and a clear one helps a site owner reach you instead of just blocking you.fetch()handles the status codes that matter. A 429 means “too many requests”. RFC 6585 defines it and says the server may send aRetry-Afterheader saying how long to wait. The function honours it, and otherwise waits 1, 2, 4 and 8 seconds.r.contentpasses raw bytes to Beautiful Soup so it can read the page’s own charset. Passingr.texthere turned “£” into “£” in our test, because the server doesn’t declare a charset in its headers. It’s a classic beginner bug.parse()selects each product card, then the fields inside it, and turns relative links into full URLs withurljoin.time.sleep(1)is the politeness delay. On a real site, start slow and only speed up if the site clearly copes.
Run it with python scrape_books.py. The first lines of books.csv look like this:
title,price,rating,url
A Light in the Attic,£51.77,Three,https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html
Tipping the Velvet,£53.74,One,https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html
From here the natural next steps are to open each book’s own page for its description and stock count, to convert prices into numbers, and to store results in a database with a timestamp so you can track changes. For a deeper Python walkthrough, see web scraping with Python and residential proxies.
Scraping JavaScript-Heavy Pages
When the HTML arrives emptyMany modern sites send a nearly empty HTML page and build the content in the browser with JavaScript. If you run the scraper above on such a page, the selectors find nothing. The practice site has a JavaScript version of its quotes page: fetched with a plain HTTP request, it contains zero quote elements, even though a browser shows ten.
You have two options.
Option 1: find the data request
Open developer tools, go to the Network tab, filter by Fetch/XHR and reload. Pages often load their content as JSON from an internal endpoint. If you can request that JSON directly, you skip HTML parsing entirely and get clean, structured data. Check that you are allowed to use it, and keep the same polite pace.
Option 2: use a real browser
Playwright drives a real browser engine, waits for the page to render, and lets you select elements as the user sees them. After pip install playwright and playwright install chromium:
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto("https://quotes.toscrape.com/js/")
page.wait_for_selector("div.quote")
quotes = page.locator("div.quote span.text").all_inner_texts()
browser.close()
print(len(quotes), "quotes")
This returned all 10 quotes in our test. Browsers use far more memory and bandwidth than plain requests, so use them only where you have to. Our engineering posts on web automation with Chromium and on running browsers at scale go deeper.
Web Scraping Best Practices and Rules
Be the scraper a site owner doesn’t noticeScraping public pages is common and widely done, but “public” doesn’t mean “anything goes”. These web scraping best practices keep you on the right side of site owners and the law:
- Read robots.txt. It lives at
/robots.txton every site and tells automated clients which paths they are asked not to visit. The standard, RFC 9309, is explicit that these rules “are not a form of access authorization”. It is a request, and following it is basic good conduct. Python’s built-inurllib.robotparsercan check a URL for you. - Read the terms of service. Some sites forbid automated access. Scraping behind a login is riskier, because you have accepted their terms to get in.
- Go slowly. One request per second per site is a sensible start. Honour 429 responses and
Retry-After, and scrape outside the site’s peak hours where you can. - Identify yourself. Use a user agent that names your project and gives a contact.
- Take only what you need. Don’t download images, scripts and every page when you want one table.
- Be careful with personal data. Names, emails and profiles are personal data under laws such as the EU’s GDPR, even when they are public. Collect them only with a lawful basis, and keep them only as long as you need them.
- Respect copyright. Facts like prices are generally fine to collect. Republishing articles, photos or reviews wholesale is a different matter.
- Prefer an API when the site offers one for the data you need.
None of this is legal advice. Rules differ between countries. If a project involves personal data, logins or large-scale republishing, ask a lawyer before you build it.
Why Scrapers Get Blocked and Where Proxies Fit
IP limits, location and reputationYour first scraper will run fine against a practice site. Real projects hit three kinds of walls:
- Rate limits per IP address. Sites count requests per address. Send too many from one IP, even at a polite pace, and you get 429s, CAPTCHAs or silent blocks.
- Location. Prices, search results, stock and even whole catalogues change with the visitor’s country or city. A scraper in a Frankfurt data center sees the German site, not the Brazilian one.
- IP reputation. Cloud and datacenter address ranges are well known. Some sites treat them with more suspicion than home or mobile connections.
A proxy sends your request from a different IP address. A proxy network gives you many addresses and lets you choose where they are. Instead of one IP making 10,000 requests, many IPs each make a few, from the country you need to see.
| Proxy type | What the site sees | Good for | ProxyEmpire targeting |
|---|---|---|---|
| Rotating residential | A home broadband connection | Most scraping: e-commerce, travel, search results | Country, region, city, ZIP, ISP, ASN, OS fingerprint |
| Rotating mobile | A phone on a mobile carrier | Mobile-first sites and apps, the strictest targets | Country, region, city, ZIP, carrier, ASN, OS fingerprint |
| Datacenter | A server in a data center | Tolerant sites, high volume at the lowest cost | Country |
| Static residential | The same home IP every time | Logged-in sessions that must keep one address | Country |
In the scraper above, set the PROXY_URL environment variable to the proxy address from your dashboard and every request goes through it. Requests also accepts socks5:// URLs once you install its SOCKS extra. Two settings matter most:
- Rotating gives a new IP on every request. Use it for independent page fetches, like one product page after another.
- Sticky sessions keep the same IP for a series of requests. Use them when a site ties a cart, a search or a login to your address. On ProxyEmpire there is no fixed time limit: the IP stays until you rotate it or the device behind it goes offline.
Proxies don’t replace good manners. They spread your load and let you see the right country. Pair them with the pace and rules above. ProxyEmpire’s network has more than 30 million ethically sourced IPs and 99.9% uptime, with live figures on its public status page. Residential bandwidth costs $3.50/GB pay-as-you-go and drops to $1.50/GB on larger plans, datacenter starts at $0.35/GB, unused bandwidth rolls over, and the trial is $1.97. For a comparison of options, see the best proxies for web scraping.
A Web Scraping Checklist for Beginners
Before, during and afterBefore you run it
- Is there an official API or a downloadable dataset? Use that first.
- Read robots.txt and the terms. Note any paths you must skip.
- Does the data appear in the raw HTML, or only after JavaScript runs?
- Write down exactly which fields you need, and nothing more.
- Will you collect personal data? If yes, check the rules first.
While it runs
- A descriptive user agent with a contact.
- A delay between requests, and backoff on 429 and 5xx responses.
- Timeouts on every request, so one slow page can’t hang the job.
- Log the URL, status code and time of every request.
- Proxies in the right country once volume or location requires it.
After
- Check a sample of rows against the live page by eye.
- Watch for empty fields, a sign the layout changed and your selectors broke.
- Store a timestamp with every row so you can compare runs.
- Delete data you no longer need.
Frequently Asked Questions
The short answersIs web scraping legal?
Collecting publicly available, non-personal data such as prices is common and generally accepted, but the answer depends on the country, the site’s terms, the kind of data and what you do with it. Personal data, content behind a login and republishing copyrighted material carry the most risk. Follow robots.txt, go slowly, and get legal advice for anything sensitive.
Is web scraping hard to learn?
The web scraping 101 basics take a weekend if you know a little programming: HTTP requests, HTML structure and CSS selectors. The scraper in this guide is under 60 lines. The harder parts come later: pages built with JavaScript, sites that change layout, and running thousands of requests reliably.
What is the best language for web scraping?
Python is the most popular choice for beginners because Requests, Beautiful Soup, Scrapy and Playwright are mature and well documented. JavaScript with Node.js is equally capable, especially for browser automation. Pick the language you already know.
What is the difference between manual and automated web scraping?
Manual scraping is a person copying data from pages by hand. It suits small, one-off lists. Automated scraping uses software, such as a script, a framework or a no-code tool, to collect data quickly and repeatedly, which suits monitoring and large datasets.
What is the difference between web scraping and web crawling?
Crawling discovers pages by following links. Scraping extracts specific data from the pages. Search engines mostly crawl and index. A price tracker crawls category pages to find products and then scrapes each product page.
Do I need proxies for web scraping?
Not for learning or small jobs. You need them when a site limits requests per IP address, when you need to see results from a specific country or city, or when you run enough volume that one address would be blocked. Residential proxies suit most targets. Datacenter proxies are cheaper for tolerant sites.
How do I avoid getting blocked when scraping?
Slow down, respect 429 responses and Retry-After, use a clear user agent, avoid unnecessary requests, and spread traffic across IP addresses with a proxy network when volume grows. Use sticky sessions when a site expects the same visitor across several pages.
Can I scrape a website with Excel or Google Sheets?
For simple tables, yes. Excel’s Power Query can import web tables, and Google Sheets has IMPORTHTML and IMPORTXML functions. They are a good bridge between manual and automated scraping, but they struggle with pagination, JavaScript and anything above a few pages.
References
Primary documentation- IETF — RFC 9309, “Robots Exclusion Protocol”. rfc-editor.org/rfc/rfc9309
- IETF — RFC 6585, “Additional HTTP Status Codes”, section 4 (429 Too Many Requests). rfc-editor.org/rfc/rfc6585
- Requests — Advanced usage: sessions, proxies and SOCKS. requests.readthedocs.io/en/latest/user/advanced
- Beautiful Soup — documentation. crummy.com/software/BeautifulSoup/bs4/doc
- Playwright for Python — installation. playwright.dev/python/docs/intro
- Scrapy — AutoThrottle extension. docs.scrapy.org/en/latest/topics/autothrottle.html
- Python — urllib.robotparser. docs.python.org/3/library/urllib.robotparser.html
- Books to Scrape and Quotes to Scrape — practice sites used in the examples. books.toscrape.com
Scrape from the country you need to see
More than 30 million ethically sourced residential and mobile IPs with country, city, ZIP, ISP and ASN targeting at no extra charge, sticky or rotating sessions, HTTP(S) and SOCKS5, and bandwidth that rolls over. 24/7 support from real people. Try ProxyEmpire for $1.97.














