Python web scraping is the most common way to collect public data from websites: a few lines of requests and BeautifulSoup can turn a page into structured data. Scaling it is where most projects struggle, because sites limit how many requests one IP can make and show different content in different countries. This guide walks through web scraping with Python step by step, from a first scraper to Scrapy and Playwright, and shows exactly how to use residential proxies with requests, httpx, Scrapy and headless browsers, with working code for each.
The short version
Fetch pages with requests or httpx, parse them with BeautifulSoup, and pass a proxy URL in the format http://
Set Up Your Python Web Scraping Environment
Five libraries cover almost everything| Library | Job | Use it when |
|---|---|---|
| requests | HTTP client | Simple, synchronous scraping of HTML and APIs |
| BeautifulSoup (bs4) | HTML parser | Extracting elements with CSS selectors or tags |
| httpx | HTTP client with async support | Hundreds of pages concurrently |
| Scrapy | Crawling framework | Large crawls with scheduling, pipelines and retries built in |
| Playwright | Browser automation | Pages that only render with JavaScript |
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install requests beautifulsoup4 httpx scrapy playwright
playwright install chromium
Which library for which job?
If you are unsure where to start, pick by the page and the size of the job. A few dozen static pages a day: requests and BeautifulSoup, nothing more. A few thousand pages on a schedule: httpx with asyncio, or Scrapy if you also need crawling, retries and export pipelines without writing them yourself. Pages that render in the browser: find their JSON first, and use Playwright only if there is none. Many production scrapers mix all three, using a headless browser once to log the API calls and a lightweight client for the daily work. Whatever the stack, the proxy settings are the same host, port, username and password, so switching tools later costs minutes, not days.
Your First Python Scraper
Fetch, parse, saveEvery scraper does the same three things: download a page, pull the values you need out of the HTML, and save them. This example collects titles and prices from a product listing page. Replace the URL and selectors with those of the page you are working on; your browser’s “Inspect” tool shows the right selectors.
import csv, requests
from bs4 import BeautifulSoup
URL = "https://www.example.com/category/laptops"
HEADERS = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)",
"Accept-Language": "en-US,en;q=0.9"}
r = requests.get(URL, headers=HEADERS, timeout=30)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
rows = []
for card in soup.select("div.product-card"):
title = card.select_one("h2").get_text(strip=True)
price = card.select_one(".price").get_text(strip=True)
rows.append({"title": title, "price": price})
with open("products.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=["title", "price"])
w.writeheader(); w.writerows(rows)
print(len(rows), "products saved")
This works for a handful of pages. Run it across thousands of pages or every hour, and the site will eventually slow you down with 429 responses, serve a CAPTCHA page instead of products, or show you prices for the wrong country. That is where proxies come in. New to the topic? Start with our web scraping basics.
Parsing tips that save hours
- Prefer stable hooks. Choose selectors based on data attributes, IDs or semantic classes such as
.price, not on long chains of layout classes that change with every redesign. - Handle missing elements.
select_onereturnsNonewhen nothing matches; check before callingget_text, and log which page and field failed. - Clean values once. Strip currency symbols and thousands separators, convert prices to numbers and dates to ISO format in one function, so every scraper stores the same format.
- Follow pagination carefully. Read the “next” link from the page rather than guessing page numbers, and stop when it disappears.
- Save the raw HTML for a sample of pages. When a parser breaks, you can fix and re-run it without downloading everything again.
How to Use Proxies in Python Requests
One dictionary, every requestThe requests library takes a proxies dictionary that maps the URL scheme to a proxy URL. For an authenticated proxy, put the username and password in the proxy URL. With ProxyEmpire, the dashboard’s connection builder generates the proxy details for the country, city and rotation you choose, so the code stays the same.
import requests
PROXY = "http://USERNAME:[email protected]:5000"
proxies = {"http": PROXY, "https": PROXY}
r = requests.get("https://httpbin.org/ip", proxies=proxies, timeout=30)
print(r.json()) # shows the proxy's exit IP, not yours
Use a Session for many requests
A requests.Session reuses connections and keeps cookies, which is faster and behaves more like a real visitor. Set the proxies once on the session:
s = requests.Session()
s.proxies.update({"http": PROXY, "https": PROXY})
s.headers.update({"User-Agent": "Mozilla/5.0", "Accept-Language": "en-US"})
for url in urls:
r = s.get(url, timeout=30)
Environment variables
requests also reads the HTTP_PROXY, HTTPS_PROXY and NO_PROXY environment variables, which is handy for tools you cannot edit. Explicit proxies in code are clearer and easier to debug.
SOCKS5 proxies
SOCKS support is optional: install it with pip install "requests[socks]", then use a socks5:// URL. The requests documentation notes that socks5 resolves host names on your machine; use socks5h to have the proxy resolve them.
proxies = {"http": "socks5h://USERNAME:PASSWORD@PROXY_HOST:PROXY_PORT",
"https": "socks5h://USERNAME:PASSWORD@PROXY_HOST:PROXY_PORT"}
If a password contains special characters such as @ or :, URL-encode it with urllib.parse.quote before building the proxy URL, or the request fails with a 407 error.
Rotating vs Sticky Sessions
Choose per task, not per project- A new IP per request suits lists of independent pages: product URLs, search result pages, profile pages. Each request leaves from a different household IP, so no single IP sends much traffic.
- A sticky session keeps the same IP for a series of requests. Use it when a site ties cookies or a basket to the IP, or when you page through results that must come from one visitor.
- Location targeting matters for anything local: prices, availability, search results. ProxyEmpire’s rotating residential proxies target country, region, city, ZIP, ISP and ASN at no extra charge.
Our guide to sticky vs rotating proxies covers the trade-offs in more depth.
Retries, Timeouts and Error Handling
Robust scrapers fail gracefullyimport requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
retry = Retry(total=4, backoff_factor=1.5,
status_forcelist=[429, 500, 502, 503, 504],
respect_retry_after_header=True)
s = requests.Session()
s.mount("http://", HTTPAdapter(max_retries=retry))
s.mount("https://", HTTPAdapter(max_retries=retry))
s.proxies.update({"http": PROXY, "https": PROXY})
r = s.get("https://www.example.com/", timeout=(10, 30)) # connect, read
- Always set a timeout. A request without one can hang forever, and a stuck worker quietly slows down the whole job.
- Retry 429 and 5xx responses with backoff; with per-request rotation, each retry leaves from a new IP.
- Never retry a 407 automatically: it means the proxy credentials are wrong.
- Validate content, not only status codes: a 200 page can still be a CAPTCHA or an empty template.
Our proxy error codes guide explains every error you are likely to meet.
Async Python Web Scraping with httpx
Hundreds of pages at onceSynchronous requests handle one page at a time. httpx offers the same style of API with async support; it takes the proxy through the proxy parameter on the client. Limit concurrency with a semaphore so you stay polite to each site.
import asyncio, httpx
PROXY = "http://USERNAME:[email protected]:5000"
sem = asyncio.Semaphore(10)
async def fetch(client, url):
async with sem:
r = await client.get(url, timeout=30)
return url, r.status_code, len(r.text)
async def main(urls):
async with httpx.AsyncClient(proxy=PROXY, headers={"User-Agent": "Mozilla/5.0"}) as client:
for result in await asyncio.gather(*(fetch(client, u) for u in urls)):
print(result)
asyncio.run(main(["https://www.example.com/"] * 20))
Scrapy with Proxies
For crawls of thousands of pagesScrapy handles scheduling, concurrency, retries, throttling and export pipelines for you. Its built-in HttpProxyMiddleware reads the proxy from each request’s proxy meta key, and also obeys the standard proxy environment variables.
import scrapy
PROXY = "http://USERNAME:[email protected]:5000"
class ProductsSpider(scrapy.Spider):
name = "products"
start_urls = ["https://www.example.com/category/laptops"]
custom_settings = {"AUTOTHROTTLE_ENABLED": True, "CONCURRENT_REQUESTS_PER_DOMAIN": 4,
"ROBOTSTXT_OBEY": True}
def start_requests(self):
for url in self.start_urls:
yield scrapy.Request(url, meta={"proxy": PROXY})
def parse(self, response):
for card in response.css("div.product-card"):
yield {"title": card.css("h2::text").get(), "price": card.css(".price::text").get()}
nxt = response.css("a.next::attr(href)").get()
if nxt:
yield response.follow(nxt, meta={"proxy": PROXY})
From script to scheduled job
A scraper that runs once is a script; a scraper that runs every day is a data product, and it needs a little more care. Schedule it with cron, a task scheduler or your orchestration tool. Store results with a timestamp and the country the request came from, so you can compare runs. Deduplicate on a stable key such as the product URL or ID. Track three numbers on every run: pages requested, pages parsed successfully and records saved. A sudden drop in the second number usually means the site changed its layout; a drop in the first usually means a network or proxy problem. Alert on both, and keep the last good dataset available until the scraper is fixed.
Finally, budget bandwidth. With pay-per-GB proxies, fetching only the pages you need, skipping images and reusing sessions keeps costs predictable. ProxyEmpire’s unused bandwidth rolls over, so a quiet month is not wasted.
JavaScript Pages with Playwright
Only when the HTML is emptyIf the HTML you download has no data and the page fills itself in with JavaScript, drive a real browser. Playwright for Python takes the proxy at launch, with separate username and password fields. Block images and media to save bandwidth. Our guide to headless browsers covers this in detail.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(proxy={"server": "http://v2.proxyempire.io:5000",
"username": "USERNAME", "password": "PASSWORD"})
page = browser.new_page()
page.goto("https://www.example.com/app", wait_until="networkidle")
prices = page.locator(".price").all_inner_texts()
print(prices)
browser.close()
Find the JSON Behind the Page
Often the fastest scraper of allMany modern pages load their data from a JSON endpoint and then build the HTML in the browser. Open your browser’s developer tools, go to the Network tab, filter by Fetch/XHR and reload the page. If you find the request that returns the products or listings, call it directly with requests and parse the JSON: no HTML parsing and no browser. Many pages also embed structured data in application/ld+json script tags, which you can read with BeautifulSoup and json.loads. Our guide to JSON web scraping goes further.
Key takeaways
- Start with requests and BeautifulSoup; move to httpx or Scrapy for scale and Playwright only for JavaScript pages.
- Pass the proxy as http://
USERNAME: PASSWORD@ host:port and URL-encode special characters. - Use per-request rotation for independent pages and sticky sessions for multi-step flows.
- Set timeouts, retry with backoff, and validate the content of every response.
Scrape Responsibly
Public data, reasonable load- Collect public data only. Do not scrape content behind logins you are not entitled to, and avoid personal data unless you have a lawful basis.
- Check robots.txt with Python’s built-in
urllib.robotparser, and read the site’s terms. - Keep the load low. Limit concurrency per domain, add delays, and prefer off-peak hours for large jobs.
- Use official APIs where a site offers one.
ProxyEmpire proxies are for lawful use. Some site categories, such as banking, payment and government websites, are blocked by default on all products.
Python Web Scraping FAQ
Straight answersWhich Python library is best for web scraping?
requests with BeautifulSoup for simple jobs, httpx for async scraping, Scrapy for large crawls, and Playwright for JavaScript-rendered pages. Most projects combine two of them.
How do I use a proxy with Python requests?
Pass a proxies dictionary: {“http”: URL, “https”: URL}, where URL is http://
Why do I need proxies for web scraping with Python?
Websites limit how many requests one IP can make and show different content by location. Rotating residential proxies spread requests over many household IPs and let you choose the country or city.
Why does requests return a 407 error through my proxy?
The proxy rejected your credentials. Check the username and password, and URL-encode any special characters in them.
How do I use a SOCKS5 proxy with requests?
Install the optional dependency with pip install “requests[socks]” and use a socks5:// or socks5h:// URL. socks5h lets the proxy resolve host names.
How do I rotate proxies in Python?
With a rotating gateway you do not need to: every request through the same endpoint can leave from a new IP. Choose per-request rotation or a sticky session in the dashboard.
Is web scraping with Python legal?
Scraping publicly available data is generally lawful, but website terms, copyright, database rights and privacy laws still apply. Collect public, non-personal data, keep the load reasonable and use official APIs where they exist.
References
Primary documentation- Requests: Advanced usageProxies, sessions, SOCKS and socks5h
- Beautiful Soup documentationParsing HTML and CSS selectors
- HTTPX: ProxiesThe proxy parameter and mounts
- Scrapy: Downloader middlewareHttpProxyMiddleware and the proxy meta key
- Playwright: HTTP proxyBrowser and context proxies with credentials
- urllib3: RetryBackoff, status codes and Retry-After
- Python: urllib.robotparserReading robots.txt rules
Proxies for your Python scrapers
One endpoint for more than 30 million residential IPs, a new IP per request or sticky sessions, targeting by country, city, ZIP, ISP and ASN at no extra cost, no concurrent session limit and 24/7 support from real people. Try it for $1.97.














