Web Scraping 101: A Complete Guide to Web Scraping for Beginners

Beginner’s guide · Last reviewed 26 September 2026 · 14 min read

Web scraping 101 in one line: a program downloads web pages and copies the pieces you care about, such as prices, titles or reviews, into a spreadsheet or database. This guide is web scraping for beginners. It covers how scraping works, manual versus automated scraping, the HTML and CSS you need, a working Python scraper you can run today, pages built with JavaScript, the rules to follow, and when rotating residential proxies start to matter.

The short version

Web scraping is the automated collection of data from websites: fetch a page, parse its HTML, pick out the elements you need with selectors, and save them in a structured format. Beginners can start with Python, Requests and Beautiful Soup, move to Playwright for pages built by JavaScript, and add proxies once a project needs many requests, several countries or steady access to sites that limit traffic per IP address.

Web Scraping 101: What Is Web Scraping?

Turning web pages into data

A web page is written for people. The price of a product sits inside a styled box, next to a picture, under a menu. Web scraping is the process of pulling that price out and putting it somewhere a computer can use it, like a row in a CSV file with columns for the product name, price, rating and link.

You can do it by hand, by copying and pasting, and for ten products that is fine. For ten thousand products, or for the same ten products every hour, you write a program instead. That program is called a scraper. It requests the page like a browser would, reads the HTML that comes back, finds the elements that hold the data, and saves them.

Two related terms come up constantly:

  • Crawling is discovering pages: start from a URL, follow links, build a list of pages to visit. Search engines crawl.
  • Scraping is extracting data from the pages you visit. Most real projects do both: crawl the category pages to find product URLs, then scrape each product page.

Key takeaways

  • Scraping has four steps: request, parse, extract, store.
  • If the data is in the page’s HTML, Requests and Beautiful Soup are enough. If JavaScript builds the page, use a browser tool such as Playwright, or find the JSON the page loads.
  • Check the site’s robots.txt and terms, go slowly, and never collect more personal data than you need.
  • Blocks usually come from too many requests per IP address or from the wrong location. That is where proxies come in.

How Web Scraping Works

Request, parse, extract, store

Every scraper, from a 20-line script to a crawler fetching millions of pages, runs the same loop:

Figure 1 — the basic scraping loop
  list of URLs
       |
       v
+--------------+    HTTP GET     +---------------+
|   scraper    | --------------> |   web server  |
|              | <-------------- |               |
+--------------+   HTML / JSON   +---------------+
       |
       |  1. parse the HTML into a tree
       |  2. select elements (CSS selectors / XPath)
       |  3. clean the values (trim, convert prices)
       v
+--------------+
| CSV / JSON / |      next page? --> back to the top
|   database   |
+--------------+
  1. Request. The scraper sends an HTTP GET request for a URL, the same request a browser sends when you type an address.
  2. Receive. The server returns a status code (200 means OK, 404 not found, 429 too many requests) and the page body, usually HTML.
  3. Parse. A parser turns the HTML text into a tree of elements you can search.
  4. Extract. Selectors point at the elements that hold the data: “the text of every p with the class price_color“.
  5. Store. The values are cleaned and written to a file or database.
  6. Repeat. The scraper follows the “next page” link or takes the next URL from its list, with a pause between requests.

Types of Web Scraping: Manual vs Automated

Copy-paste, no-code tools and code

There are two main categories, and the second splits into a few approaches.

Manual web scraping means a person collects the data: opening pages, copying values into a spreadsheet, maybe using the browser’s developer tools to copy a table. It needs no setup and it is the right choice for a one-off list of a few dozen items. It is slow, it doesn’t repeat itself, and people make copy errors.

Automated web scraping means software does the collecting. It is faster, it can run on a schedule, and it gives the same result every time. It comes in several forms:

No-code toolsBrowser extensions and desktop apps where you click on the data you want and the tool builds the scraper. Good for simple sites and non-programmers.
ScriptsA short program, usually Python or JavaScript, that requests pages and parses them. Full control, and the best way to learn.
FrameworksTools such as Scrapy handle queues, retries, throttling and exports for you. Worth it once you scrape more than a few thousand pages.
Headless browsersPlaywright or Puppeteer drive a real browser, so pages built by JavaScript render fully before you extract anything.
Official APIsIf a site offers an API for the data you need, use it. It is more stable than scraping and clearly permitted.
ManualAutomated
SetupNoneMinutes to days
SpeedA few pages a minuteMany pages a second, if the site allows it
RepeatableNoYes, on a schedule
ErrorsCopy mistakesBreaks when the page layout changes
Best forOne-off, small listsOngoing monitoring, large datasets

What Is Web Scraping Used For?

The jobs that pay for it

Most commercial scraping is about public information that changes often and matters to a business decision:

Price monitoringTrack competitor prices, stock and promotions across shops and countries. See price monitoring.
SEO and SERP trackingRecord where pages rank for a keyword in a given city or country. See SEO monitoring.
Travel faresCompare flight and hotel prices, which often differ by the visitor’s location. See travel fare aggregation.
Market researchCollect product listings, reviews and ratings to spot trends and gaps.
Brand protectionFind counterfeit listings and unauthorised sellers. See brand protection.
Ad verificationCheck that ads appear where, and to whom, they were bought. See ad verification.
Research and AIBuild datasets for analysis or model training. See LLM training data.
Lead lists and directoriesCollect business listings, but mind the personal-data rules below.

What You Need Before You Start

A small, free toolkit

For web scraping 101 you don’t need much. Python is the most common language for scraping because its libraries are simple and well documented, though JavaScript (Node.js) works just as well. This guide uses Python.

ToolWhat it doesWhen you need it
Browser developer toolsInspect the HTML of any element, watch network requestsAlways. It is how you find selectors.
Python 3Runs your scraperAlways
RequestsSends HTTP requests, handles sessions, headers and proxiesPages whose data is in the HTML
Beautiful SoupParses HTML and finds elements with CSS selectorsWith Requests
PlaywrightDrives a real Chromium, Firefox or WebKit browserPages built by JavaScript, clicks, scrolling
ScrapyFull crawling framework with queues, retries and throttlingThousands of pages or more

Install the two libraries for the first scraper with one command:

Figure 2 — install Requests and Beautiful Soup
python -m pip install requests beautifulsoup4

Understanding HTML and CSS Selectors

How to point at the data

HTML is a tree of nested elements. Each element has a tag name (div, p, a), and most have attributes such as class, id or href. Here is one product from books.toscrape.com, a practice site built for learning to scrape:

Figure 3 — the HTML of one product card (shortened)
<article class="product_pod">
  <p class="star-rating Three">...</p>
  <h3>
    <a href="catalogue/a-light-in-the-attic_1000/index.html"
       title="A Light in the Attic">A Light in the ...</a>
  </h3>
  <div class="product_price">
    <p class="price_color">£51.77</p>
    <p class="instock availability">In stock</p>
  </div>
</article>

A CSS selector is a short pattern that matches elements. The same syntax styles web pages, so the browser’s developer tools can test it for you: right-click an element, choose Inspect, and use Ctrl+F in the Elements panel to try a selector.

SelectorMatchesIn the example
article.product_podEvery article with class product_podEach product card
h3 aAn a anywhere inside an h3The title link
p.price_colorA p with class price_color£51.77
p.star-ratingThe rating paragraph; the second class holds the valueThree
li.next aThe link inside the “next” list itemThe next page URL

Two habits save a lot of trouble. First, read attributes when the visible text is cut short: the title text above ends in “…”, but the title attribute has the full name. Second, prefer stable hooks such as meaningful class names or data- attributes over position (“the third div”), because layouts change.

How to Scrape a Website: Your First Scraper

A working Python example, step by step

This is the practical core of web scraping 101. The scraper below collects the title, price, rating and link of every book on the first three catalogue pages of books.toscrape.com and saves them to books.csv. It waits one second between pages, retries politely when the server says it is busy, and can send its traffic through a proxy if you set one. We ran it on 26 September 2026 and it saved 60 books.

Figure 4 — scrape_books.py
import csv
import os
import time
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

START_URL = "https://books.toscrape.com/catalogue/page-1.html"
HEADERS = {"User-Agent": "books-tutorial/1.0 (+https://example.com/contact)"}

# Optional: route requests through a proxy, e.g. http://user:pass@host:port
proxy = os.environ.get("PROXY_URL")
PROXIES = {"http": proxy, "https": proxy} if proxy else None


def fetch(session, url, tries=4):
    """GET a page, backing off on 429 and 5xx responses."""
    for attempt in range(tries):
        r = session.get(url, headers=HEADERS, proxies=PROXIES, timeout=20)
        if r.status_code == 200:
            return r.content  # bytes: BeautifulSoup reads the charset itself
        if r.status_code == 429 or r.status_code >= 500:
            wait = int(r.headers.get("Retry-After", 2 ** attempt))
            time.sleep(wait)
            continue
        r.raise_for_status()
    raise RuntimeError(f"gave up on {url}")


def parse(html, page_url):
    soup = BeautifulSoup(html, "html.parser")
    rows = []
    for card in soup.select("article.product_pod"):
        link = card.select_one("h3 a")
        rows.append({
            "title": link["title"],
            "price": card.select_one("p.price_color").get_text(strip=True),
            "rating": card.select_one("p.star-rating")["class"][1],
            "url": urljoin(page_url, link["href"]),
        })
    nxt = soup.select_one("li.next a")
    return rows, (urljoin(page_url, nxt["href"]) if nxt else None)


def main(max_pages=3):
    url, books = START_URL, []
    with requests.Session() as s:
        for _ in range(max_pages):
            rows, url = parse(fetch(s, url), url)
            books.extend(rows)
            if not url:
                break
            time.sleep(1)  # be polite: one request per second
    with open("books.csv", "w", newline="", encoding="utf-8") as f:
        w = csv.DictWriter(f, fieldnames=["title", "price", "rating", "url"])
        w.writeheader()
        w.writerows(books)
    print(f"saved {len(books)} books")


if __name__ == "__main__":
    main()

What each part does:

  • HEADERS gives your scraper an honest name and a way to contact you. Many sites treat the default library user agent with suspicion, and a clear one helps a site owner reach you instead of just blocking you.
  • fetch() handles the status codes that matter. A 429 means “too many requests”. RFC 6585 defines it and says the server may send a Retry-After header saying how long to wait. The function honours it, and otherwise waits 1, 2, 4 and 8 seconds.
  • r.content passes raw bytes to Beautiful Soup so it can read the page’s own charset. Passing r.text here turned “£” into “£” in our test, because the server doesn’t declare a charset in its headers. It’s a classic beginner bug.
  • parse() selects each product card, then the fields inside it, and turns relative links into full URLs with urljoin.
  • time.sleep(1) is the politeness delay. On a real site, start slow and only speed up if the site clearly copes.

Run it with python scrape_books.py. The first lines of books.csv look like this:

Figure 5 — output
title,price,rating,url
A Light in the Attic,£51.77,Three,https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html
Tipping the Velvet,£53.74,One,https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html

From here the natural next steps are to open each book’s own page for its description and stock count, to convert prices into numbers, and to store results in a database with a timestamp so you can track changes. For a deeper Python walkthrough, see web scraping with Python and residential proxies.

Scraping JavaScript-Heavy Pages

When the HTML arrives empty

Many modern sites send a nearly empty HTML page and build the content in the browser with JavaScript. If you run the scraper above on such a page, the selectors find nothing. The practice site has a JavaScript version of its quotes page: fetched with a plain HTTP request, it contains zero quote elements, even though a browser shows ten.

You have two options.

Option 1: find the data request

Open developer tools, go to the Network tab, filter by Fetch/XHR and reload. Pages often load their content as JSON from an internal endpoint. If you can request that JSON directly, you skip HTML parsing entirely and get clean, structured data. Check that you are allowed to use it, and keep the same polite pace.

Option 2: use a real browser

Playwright drives a real browser engine, waits for the page to render, and lets you select elements as the user sees them. After pip install playwright and playwright install chromium:

Figure 6 — rendering a JavaScript page with Playwright
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto("https://quotes.toscrape.com/js/")
    page.wait_for_selector("div.quote")
    quotes = page.locator("div.quote span.text").all_inner_texts()
    browser.close()

print(len(quotes), "quotes")

This returned all 10 quotes in our test. Browsers use far more memory and bandwidth than plain requests, so use them only where you have to. Our engineering posts on web automation with Chromium and on running browsers at scale go deeper.

Web Scraping Best Practices and Rules

Be the scraper a site owner doesn’t notice

Scraping public pages is common and widely done, but “public” doesn’t mean “anything goes”. These web scraping best practices keep you on the right side of site owners and the law:

  • Read robots.txt. It lives at /robots.txt on every site and tells automated clients which paths they are asked not to visit. The standard, RFC 9309, is explicit that these rules “are not a form of access authorization”. It is a request, and following it is basic good conduct. Python’s built-in urllib.robotparser can check a URL for you.
  • Read the terms of service. Some sites forbid automated access. Scraping behind a login is riskier, because you have accepted their terms to get in.
  • Go slowly. One request per second per site is a sensible start. Honour 429 responses and Retry-After, and scrape outside the site’s peak hours where you can.
  • Identify yourself. Use a user agent that names your project and gives a contact.
  • Take only what you need. Don’t download images, scripts and every page when you want one table.
  • Be careful with personal data. Names, emails and profiles are personal data under laws such as the EU’s GDPR, even when they are public. Collect them only with a lawful basis, and keep them only as long as you need them.
  • Respect copyright. Facts like prices are generally fine to collect. Republishing articles, photos or reviews wholesale is a different matter.
  • Prefer an API when the site offers one for the data you need.

None of this is legal advice. Rules differ between countries. If a project involves personal data, logins or large-scale republishing, ask a lawyer before you build it.

Why Scrapers Get Blocked and Where Proxies Fit

IP limits, location and reputation

Your first scraper will run fine against a practice site. Real projects hit three kinds of walls:

  • Rate limits per IP address. Sites count requests per address. Send too many from one IP, even at a polite pace, and you get 429s, CAPTCHAs or silent blocks.
  • Location. Prices, search results, stock and even whole catalogues change with the visitor’s country or city. A scraper in a Frankfurt data center sees the German site, not the Brazilian one.
  • IP reputation. Cloud and datacenter address ranges are well known. Some sites treat them with more suspicion than home or mobile connections.

A proxy sends your request from a different IP address. A proxy network gives you many addresses and lets you choose where they are. Instead of one IP making 10,000 requests, many IPs each make a few, from the country you need to see.

Proxy typeWhat the site seesGood forProxyEmpire targeting
Rotating residentialA home broadband connectionMost scraping: e-commerce, travel, search resultsCountry, region, city, ZIP, ISP, ASN, OS fingerprint
Rotating mobileA phone on a mobile carrierMobile-first sites and apps, the strictest targetsCountry, region, city, ZIP, carrier, ASN, OS fingerprint
DatacenterA server in a data centerTolerant sites, high volume at the lowest costCountry
Static residentialThe same home IP every timeLogged-in sessions that must keep one addressCountry

In the scraper above, set the PROXY_URL environment variable to the proxy address from your dashboard and every request goes through it. Requests also accepts socks5:// URLs once you install its SOCKS extra. Two settings matter most:

  • Rotating gives a new IP on every request. Use it for independent page fetches, like one product page after another.
  • Sticky sessions keep the same IP for a series of requests. Use them when a site ties a cart, a search or a login to your address. On ProxyEmpire there is no fixed time limit: the IP stays until you rotate it or the device behind it goes offline.

Proxies don’t replace good manners. They spread your load and let you see the right country. Pair them with the pace and rules above. ProxyEmpire’s network has more than 30 million ethically sourced IPs and 99.9% uptime, with live figures on its public status page. Residential bandwidth costs $3.50/GB pay-as-you-go and drops to $1.50/GB on larger plans, datacenter starts at $0.35/GB, unused bandwidth rolls over, and the trial is $1.97. For a comparison of options, see the best proxies for web scraping.

A Web Scraping Checklist for Beginners

Before, during and after

Before you run it

  • Is there an official API or a downloadable dataset? Use that first.
  • Read robots.txt and the terms. Note any paths you must skip.
  • Does the data appear in the raw HTML, or only after JavaScript runs?
  • Write down exactly which fields you need, and nothing more.
  • Will you collect personal data? If yes, check the rules first.

While it runs

  • A descriptive user agent with a contact.
  • A delay between requests, and backoff on 429 and 5xx responses.
  • Timeouts on every request, so one slow page can’t hang the job.
  • Log the URL, status code and time of every request.
  • Proxies in the right country once volume or location requires it.

After

  • Check a sample of rows against the live page by eye.
  • Watch for empty fields, a sign the layout changed and your selectors broke.
  • Store a timestamp with every row so you can compare runs.
  • Delete data you no longer need.

Frequently Asked Questions

The short answers
Is web scraping legal?

Collecting publicly available, non-personal data such as prices is common and generally accepted, but the answer depends on the country, the site’s terms, the kind of data and what you do with it. Personal data, content behind a login and republishing copyrighted material carry the most risk. Follow robots.txt, go slowly, and get legal advice for anything sensitive.

Is web scraping hard to learn?

The web scraping 101 basics take a weekend if you know a little programming: HTTP requests, HTML structure and CSS selectors. The scraper in this guide is under 60 lines. The harder parts come later: pages built with JavaScript, sites that change layout, and running thousands of requests reliably.

What is the best language for web scraping?

Python is the most popular choice for beginners because Requests, Beautiful Soup, Scrapy and Playwright are mature and well documented. JavaScript with Node.js is equally capable, especially for browser automation. Pick the language you already know.

What is the difference between manual and automated web scraping?

Manual scraping is a person copying data from pages by hand. It suits small, one-off lists. Automated scraping uses software, such as a script, a framework or a no-code tool, to collect data quickly and repeatedly, which suits monitoring and large datasets.

What is the difference between web scraping and web crawling?

Crawling discovers pages by following links. Scraping extracts specific data from the pages. Search engines mostly crawl and index. A price tracker crawls category pages to find products and then scrapes each product page.

Do I need proxies for web scraping?

Not for learning or small jobs. You need them when a site limits requests per IP address, when you need to see results from a specific country or city, or when you run enough volume that one address would be blocked. Residential proxies suit most targets. Datacenter proxies are cheaper for tolerant sites.

How do I avoid getting blocked when scraping?

Slow down, respect 429 responses and Retry-After, use a clear user agent, avoid unnecessary requests, and spread traffic across IP addresses with a proxy network when volume grows. Use sticky sessions when a site expects the same visitor across several pages.

Can I scrape a website with Excel or Google Sheets?

For simple tables, yes. Excel’s Power Query can import web tables, and Google Sheets has IMPORTHTML and IMPORTXML functions. They are a good bridge between manual and automated scraping, but they struggle with pagination, JavaScript and anything above a few pages.

References

Primary documentation
  1. IETF — RFC 9309, “Robots Exclusion Protocol”. rfc-editor.org/rfc/rfc9309
  2. IETF — RFC 6585, “Additional HTTP Status Codes”, section 4 (429 Too Many Requests). rfc-editor.org/rfc/rfc6585
  3. Requests — Advanced usage: sessions, proxies and SOCKS. requests.readthedocs.io/en/latest/user/advanced
  4. Beautiful Soup — documentation. crummy.com/software/BeautifulSoup/bs4/doc
  5. Playwright for Python — installation. playwright.dev/python/docs/intro
  6. Scrapy — AutoThrottle extension. docs.scrapy.org/en/latest/topics/autothrottle.html
  7. Python — urllib.robotparser. docs.python.org/3/library/urllib.robotparser.html
  8. Books to Scrape and Quotes to Scrape — practice sites used in the examples. books.toscrape.com

Scrape from the country you need to see

More than 30 million ethically sourced residential and mobile IPs with country, city, ZIP, ISP and ASN targeting at no extra charge, sticky or rotating sessions, HTTP(S) and SOCKS5, and bandwidth that rolls over. 24/7 support from real people. Try ProxyEmpire for $1.97.

Flexible Pricing Plan

logo purple proxyempire

Our state-of-the-art proxies.

Experience online freedom with our unrivaled web proxy solutions. Pioneering in collecting location specific data at scale, our premium, ethically-sourced network boasts a vast pool of IPs, expansive location choices, high success rate, and versatile pricing. Advance your digital journey with us.

🏘️ Rotating Residential Proxies
  • 30M+ Premium Residential IPs
  •  170+ Countries
    Every residential IP in our network corresponds to an actual desktop device with a precise geographical location. Our residential proxies aare fast and reliable, with 99.9% uptime, and work for a wide range of use cases. You can use Country, Region, City and ISP targeting for our rotating residential proxies.

See our Rotating Residential Proxies

📍 Static Residential Proxies
  • 19 Countries
    Buy a dedicated static residential IP from one of the 19 countries that we offer proxies in. Keep the same IP for a month or longer, while benefiting from their fast speed and stability.

See our Static Residential Proxies

📳 Rotating Mobile Proxies
  • 4M+ Premium Mobile IPs
  •  170+ Countries
    Access millions of clean mobile IPs with precise targeting including Country, Region, City, and Mobile Carrier. Get far fewer IP blocks and CAPTCHAs with our 4G and 5G proxies.

See our Mobile Proxies

📱 Dedicated Mobile Proxies
  • 5+ Countries
  • 50+ Locations
    Get your own dedicated mobile proxy in one of our supported locations, with unlimited bandwidth and unlimited IP changes on demand. A great choice when you need a small number of mobile IPs and a lot of proxy bandwidth.

See our 4G & 5G Proxies

🌐 Rotating Datacenter Proxies
  • 197,000+ IPs Premium IPs
  •  62 Countries
    On a budget and need to do some simple scraping tasks? Our datacenter proxies are the perfect fit! Get started with as little as $2

See our Datacenter Proxies

proxy locations

30M+ rotating IPs

99% uptime - high speed

99.9% uptime.

dedicated support team

24/7 Dedicated Support.

fair price

Fair Pricing.

🏠 Residential Proxies Rotating / Static / Unlimited
📱 Mobile Proxies Rotating and Dedicated
🖥️ Datacenter Proxies Rotating
🌍 IP Pool 30M+ residential + 4M+ mobile IPs
📶 Uptime 99.9% · Live status
💳 Payment Card · PayPal · Crypto · Bank transfer
💬 Support 24/7 live chat · [email protected]