scrapy-stealth
Stealthy Crawling. Maximum Results.
A pluggable anti-bot and stealth framework for Scrapy.
Overview
scrapy-stealth
extends Scrapy with browser impersonation, proxy rotation, fingerprint cycling, and intelligent retry strategies β
designed for large-scale, production-grade crawling.
π§ Why scrapy-stealth?
Scrapy is fast and powerful, but modern websites use advanced anti-bot protections such as:
- TLS fingerprinting
- Browser behavior detection
- Rate limiting and IP blocking
scrapy-stealth
helps by adding:
𧬠Browser-level impersonation
Deep emulation of modern TLS structures and HTTP/2 fingerprints.
π Smarter retry strategies
Heuristics engineered to intercept temporary challenges without loops.
π Proxy and fingerprint rotation
Dynamically cycles routing footprints cleanly matching agent properties.
π‘οΈ Anti-bot detection
Defeats strict platform defenses out of the box.
π Expected Result
- π Higher success rate
- π° Lower proxy cost
- π‘οΈ More stable crawls
π Comparison
| Feature | scrapy-stealth | scrapy-impersonate | scrapy-playwright | scrapy-splash | Scrapy (default) |
|---|---|---|---|---|---|
| TLS fingerprint spoofing | β | β | β | β | β |
| HTTP/2 support | β | β | β | β | β |
| Browser impersonation | β | β | β οΈ partial | β | β |
| Proxy rotation (built-in) | β | β | β | β | β |
| Fingerprint rotation | β | β | β | β | β |
| Anti-bot detection | β | β | β | β | β |
| Smart browser selection | β | β | β | β | β |
| Smart retry logic | β | β | β | β | β |
| Per-request engine switching | β | β | β | β | β |
| Headless browser required | β | β | β | β | β |
| JavaScript rendering | β | β | β | β | β |
| Screenshot / snapshot | β | β | β | β | β |
| Native Scrapy integration | β | β | β | β | β |
| Memory footprint | π’ Low | π’ Low | π΄ High | π΄ High | π’ Low |
β οΈ scrapy-playwright passes real browser TLS but does not spoof
fingerprint profiles like scrapy-stealth does.
βΉοΈ scrapy-impersonate provides TLS/HTTP2 impersonation via curl_cffi but lacks built-in rotation, detection, or per-request
engine switching.
π‘ JavaScript rendering is available via the optional browser driver
β use it selectively for pages that require a full browser.
β¨ Features
- π Pluggable engine system (
scrapy,stealth) - π§ Per-request engine selection via
request.meta - π Proxy support and rotation
- 𧬠Browser fingerprint rotation
- π Smart retry logic
- π‘οΈ Anti-bot detection (status + content-based, Cloudflare, Akamai)
- π§ Smart browser selection β start with fast
basic/turbo, auto-retry once with visible Chrome when a JS challenge or ban is detected - β‘ Thread-safe async integration
- π₯οΈ Real-browser engine (CDP) for JS-heavy pages
- π Intelligent session recycle β after consecutive bans, browser restarts Chrome; basic/turbo clear HTTP sessions
- π« Static asset blocking β skip images, fonts, CSS, and media for faster, lighter browser fetches
- π― Proxy bypass list β send chosen domains straight to
the origin instead of through the proxy (
--proxy-bypass-list) - π§ Custom DNS overrides β pin hosts to fixed IPs (connect via IP, keep hostname for TLS/SNI/Host) to dodge poisoned or geo-shifted public DNS
- πΈ Built-in snapshot decorator (
scrapy_stealth.decorators.snapshot)
π¦ Installation
pip install scrapy-stealth
π‘ Requires Python 3.11+ and Scrapy 2.12β2.x
βοΈ Setup
Option 1 β Global (settings.py)
# 1. Enable the middleware
DOWNLOADER_MIDDLEWARES = {
"scrapy_stealth.middlewares.StealthDownloaderMiddleware": 950,
}
# 2. (Optional) Route ALL requests through stealth automatically β no meta needed per request
STEALTH_ENABLED = True
STEALTH_DRIVER = "turbo" # "basic" (default), "turbo", "browser", or "auto"
STEALTH_AUTO_FALLBACK = True # retry basic/turbo JS challenges once with browser (headless=False)
# 3. (Optional) Proxy list β seeded as engine default; rotated on ban-streak session recycle
# Supported schemes: http, https, socks4, socks5
STEALTH_PROXIES = [
"http://proxy1:8080",
"http://proxy2:8080",
"http://user:pass@proxy3:8080", # with authentication
"socks5://proxy4:1080",
]
# 4. (Optional) Pin hosts to fixed origin IPs (bypass public DNS)
# Connects to the IP while keeping the hostname for TLS SNI / Host / certs
STEALTH_DNS_OVERRIDES = {
"example.com": "203.0.113.10",
"www.example.com": "203.0.113.10",
}
Option 2 β Per-spider (custom_settings)
Configure the middleware and all stealth settings directly on the spider β no changes to settings.py required.
class MySpider(scrapy.Spider):
name = "example"
custom_settings = {
"DOWNLOADER_MIDDLEWARES": {
"scrapy_stealth.middlewares.StealthDownloaderMiddleware": 950,
},
"STEALTH_ENABLED": True,
"STEALTH_DRIVER": "turbo",
"STEALTH_AUTO_FALLBACK": True,
"STEALTH_PROXIES": [
"http://proxy1:8080",
"http://user:pass@proxy2:8080",
"socks5://proxy3:1080",
],
"STEALTH_DNS_OVERRIDES": {
"example.com": "203.0.113.10",
},
}
ValueError immediately. DNS overrides are
validated the same way β invalid IPs raise ValueError
immediately.
π Quick Start
Option A β Per-request (stealth only on specific requests)
yield scrapy.Request(
url="https://example.com",
meta={"stealth": {}},
)
Option B β Global mode (stealth on every request automatically)
# settings.py or custom_settings
STEALTH_ENABLED = True
STEALTH_DRIVER = "turbo"
STEALTH_AUTO_FALLBACK = True # optional: basic/turbo -> browser on JS challenge
# No meta needed β all requests go through stealth
yield scrapy.Request(url="https://example.com")
# Opt out for a specific request
yield scrapy.Request(url="https://api.internal/health", meta={"stealth": False})
Option C β Smart browser selection (fast HTTP first, Chrome only when needed)
# settings.py or custom_settings β enable fallback for all stealth requests
STEALTH_ENABLED = True
STEALTH_DRIVER = "turbo" # primary: fast HTTP driver
STEALTH_AUTO_FALLBACK = True # retry once with browser (headless=False) on JS challenge / ban
# Or per-request β same behaviour without a global setting
yield scrapy.Request(
url="https://example.com",
meta={"stealth": {"driver": "auto"}},
)
See Smart browser selection for full details.
π§ Global Configuration
Customise package-wide defaults via the shared config
instance. All settings must be applied at module level, before the spider class β the engine
client is created at middleware initialisation, so changes inside start_requests or parse will have
no effect.
# myspider.py
import scrapy
from scrapy_stealth.config import config
config.DEFAULT_ENGINE = "stealth" # "scrapy" (native) or "stealth" (browser impersonation)
config.DEFAULT_PROFILE = "chrome_147" # browser profile when meta["stealth"]["profile"] is not set
config.DEFAULT_TIMEOUT = 30 # stealth request timeout in seconds
config.STEALTH_DRIVER = "turbo" # "basic" (default), "turbo", "browser", or "auto"
config.STEALTH_AUTO_FALLBACK = True # basic/turbo -> browser on JS challenge (headless=False)
config.HTTP2 = True # False for servers that only support HTTP/1.1
config.BLOCK_CODES |= {407} # extend blocked status codes (|= keeps defaults)
config.BLOCK_KEYWORDS.append("banned") # extend blocked body-text patterns
config.BROWSER_HEADLESS = True # browser driver: headless mode (False = visible window, more stealthy)
config.BROWSER_SETTLE_S = 4.0 # browser driver: seconds to wait after navigation for JS to finish
config.BROWSER_EXECUTABLE_PATH = "/usr/bin/brave-browser" # custom browser binary (default: auto-detect Chrome)
config.STEALTH_RECYCLE_AFTER_BANS = 5 # recycle Chrome / HTTP sessions after 5 consecutive bans
config.BROWSER_STATIC_ASSETS_BLOCK = True # block images/fonts/CSS/media (skipped when snapshot=True)
config.BROWSER_PROXY_BYPASS_LIST = ["example.com", "*.internal"] # these bypass the proxy
config.STEALTH_DNS_OVERRIDES = {"example.com": "203.0.113.10"} # pin host β origin IP
class MySpider(scrapy.Spider):
name = "example"
...
# β wrong β too late, the engine client is already created
class MySpider(scrapy.Spider):
def start_requests(self):
config.HTTP2 = False # has no effect
...
You can also read any value programmatically:
config.get("DEFAULT_ENGINE") # "scrapy"
config.get("MISSING_KEY", "default") # "default"
| Attribute | Type | Default | Description |
|---|---|---|---|
| DEFAULT_ENGINE | str | "scrapy" | Engine used when request.meta["stealth"] key is
absent
|
| DEFAULT_PROFILE | str | "chrome_147" | Browser profile used when none is specified |
| DEFAULT_TIMEOUT | int | 30 | Request timeout in seconds |
| STEALTH_DRIVER | str | "basic" | Default driver: "basic", "turbo", "browser", or
"auto". Also readable from Scrapy settings as STEALTH_DRIVER |
| STEALTH_AUTO_FALLBACK | bool | False | Smart browser selection: when True,
basic / turbo responses
that look like a JS challenge or session ban are retried once with browser in visible mode (headless=False).
Per-request equivalent: meta["stealth"]["driver"] = "auto" |
| HTTP2 | bool | True | HTTP/2 mode; overridable per-request via meta["stealth"]["http2"]
|
| BLOCK_CODES | frozenset[int] | {403, 429, 503} | HTTP status codes considered blocked |
| BLOCK_KEYWORDS | list[str] | ["captcha", "access denied", β¦] | Body-text patterns considered blocked |
| BROWSER_HEADLESS | bool | True | Browser driver: headless mode (False = visible
window, more stealthy)
|
| BROWSER_SETTLE_S | float | 4.0 | Browser driver: seconds to wait after navigation for JS to finish rendering |
| BROWSER_NO_SANDBOX | bool | None | None | Browser driver: disable Chrome sandbox. None =
auto-detect (enabled when running as root, e.g. Docker)
|
| BROWSER_EXECUTABLE_PATH | str | None | None | Browser driver: path to the browser binary. None =
auto-detect Chrome/Chromium. Set to use Brave or a custom install (e.g. "/usr/bin/brave-browser")
|
| BROWSER_MAX_TABS | int | 10 | Browser driver: max concurrent Chrome tabs across in-flight requests |
| STEALTH_RECYCLE_AFTER_BANS | int | 5 | After this many consecutive bans: browser
restarts Chrome; basic / turbo clear cached HTTP sessions/clients. Any clean response
resets the count
|
| BROWSER_STATIC_ASSETS_BLOCK | bool | False | Browser driver: block images, fonts, CSS, and media via CDP. Overridable per-request via
meta["stealth"]["static_assets_block"]; always off when snapshot=True |
| BROWSER_PROXY_BYPASS_LIST | list[str] | [] | Browser driver: domains/patterns that bypass the proxy and connect to the origin directly,
via Chrome's --proxy-bypass-list. Supports wildcards (*.example.com), IP/CIDR, ports, and <local>. Only applies when a proxy is in use; set at
browser launch (config/settings, not per-request)
|
| STEALTH_DNS_OVERRIDES | dict[str, str] | {} | HostβIP map used by basic / turbo (and Chrome --host-resolver-rules
for browser). Connects to the IP while keeping the hostname for
TLS SNI, Host header, and cert verification. Also readable from Scrapy settings as STEALTH_DNS_OVERRIDES. Per-request override via meta["stealth"]["dns"] |
For one-off overrides on a single request, set meta["stealth"]["driver"] or meta["stealth"]["http2"] (see Per-Request Configuration below).
βοΈ Per-Request Configuration
All options are passed via request.meta["stealth"].
The presence of meta["stealth"] (a dict)
activates the stealth engine. Omit the key to use the default Scrapy engine. When STEALTH_ENABLED
= True, all requests are stealth by default β pass meta={"stealth":
False} to opt out for a specific request.
yield scrapy.Request(
url,
meta={
"stealth": {
"driver": "turbo",
# optional overrides β otherwise profile/proxy come from defaults and
# rotate automatically when the session recycles after consecutive bans
"profile": "chrome_147",
"proxy": "http://user:pass@proxy:8080",
"stealth_timeout": 60,
"http2": True,
"dns": "203.0.113.10", # or {"example.com": "203.0.113.10"}
}
},
)
| Key | Type | Description |
|---|---|---|
| driver | str | "basic", "turbo", "browser", or
"auto" β "auto" enables
smart browser selection for this request (HTTP first, then browser on challenge/ban). Requires STEALTH_AUTO_FALLBACK = True globally, or use "auto" alone to opt in per-request
|
| fallback | bool | Set to False to opt out of auto-fallback for this
request (when STEALTH_AUTO_FALLBACK or driver="auto" is active)
|
| profile | str | Browser profile (e.g. "chrome_147", "safari_ios_18_1_1"). Omit to use engine default; default rotates
on ban-streak session recycle
|
| proxy | str | Explicit proxy URL. Omit to use STEALTH_PROXIES
default; default rotates on ban-streak session recycle
|
| dns | str or dict | Pin DNS: bare IP for this request's hostname, or {host:
ip} mapping. Merges over STEALTH_DNS_OVERRIDES. Works
with basic/turbo
per-request; browser uses global overrides at Chrome launch
only
|
| stealth_timeout | int | Per-request timeout in seconds (overrides default 30s) |
| http2 | bool | True = HTTP/2, False = HTTP/1.1 (overrides config.HTTP2
for this request)
|
| headless | bool | Browser driver only: True = headless, False = visible window (more stealthy)
|
| settle | float | Browser driver only: seconds to wait for JS after navigation (default 4.0)
|
| snapshot | bool | Browser driver only: capture a PNG snapshot β result available as response.meta["snapshot_content"] (bytes)
|
| static_assets_block | bool | Browser driver only: block images, fonts, CSS, and media for this request (overrides config.BROWSER_STATIC_ASSETS_BLOCK). Ignored β always unblocked β
when snapshot is True
|
π§ Custom DNS Overrides
Pin a hostname to a fixed origin IP so the package dials that address directly
instead of trusting public DNS. The request URL stays as https://example.com/...
β TLS SNI, the Host header, and certificate verification still use the
hostname.
Global (settings.py /
custom_settings / config):
STEALTH_DNS_OVERRIDES = {
"shop.example.com": "203.0.113.10",
"cdn.example.com": "203.0.113.11",
}
Per-request (overrides or extends the global map):
yield scrapy.Request(
"https://shop.example.com/item/1",
meta={"stealth": {"driver": "turbo", "dns": "203.0.113.10"}},
)
# Or a full mapping:
meta = {"stealth": {"dns": {"shop.example.com": "203.0.113.10"}}}
Supported on basic and turbo per-request. The browser driver applies the effective map (config + meta["stealth"]["dns"]) via a local CONNECT relay
that dials the pinned IP (Chrome's --host-resolver-rules is not
used β it is unreliable). Chrome is pointed at the relay with --proxy-server;
when the DNS map changes, the browser restarts so the relay is rebuilt.
β οΈ Do not put DNS-pinned hosts on BROWSER_PROXY_BYPASS_LIST or they
will skip the relay.
π‘ With an HTTP proxy on basic/turbo, DNS is often resolved by the proxy β prefer direct
connections
or SOCKS when using overrides.
π§ Smart Browser Selection
Pick the right driver automatically: stay on fast HTTP impersonation (basic / turbo) for normal pages, and
escalate to real Chrome only when the response looks like a JS challenge or session ban (403/429/503, Cloudflare
βJust a momentβ, Akamai, DataDome, and similar signals).
| Phase | Driver | When |
|---|---|---|
| 1 | basic or turbo
|
Default β low memory, high throughput |
| 2 | browser (headless=False)
|
One retry when phase 1 is blocked or challenged |
The fallback always opens a visible Chrome
window (headless=False) for better evasion β regardless of BROWSER_HEADLESS or any prior meta["stealth"]["headless"]
value.
Global (settings.py /
custom_settings):
STEALTH_ENABLED = True
STEALTH_DRIVER = "turbo" # or "basic" β primary HTTP driver
STEALTH_AUTO_FALLBACK = True # off by default; set True to enable smart selection globally
Per-request (no global setting β fallback for that URL only):
yield scrapy.Request(
url,
meta={"stealth": {"driver": "auto"}}, # primary = STEALTH_DRIVER, else basic
)
Always use browser (skip phase 1):
meta = {"stealth": {"driver": "browser"}}
Opt out for one request:
meta = {"stealth": {"driver": "auto", "fallback": False}}
# or globally: STEALTH_AUTO_FALLBACK = False
Each request is retried at most once. If the browser fetch fails, the original basic / turbo response
is returned.
Console output and stats (stealth/fallbacks, stealth/requests/browser) show when escalation happened.
π₯οΈ Browser Engine
For sites protected by Cloudflare JS challenges or heavy JavaScript rendering,
use the browser driver. It runs a real Chrome instance via the DevTools
Protocol (no WebDriver), keeping one persistent browser and opening a new tab per request.
basic/turbo (~5-15s per page vs <2s).
Use it selectively β route only JS-protected URLs to "browser" and keep
everything else on "turbo".
βΈ Basic Usage
Per-request (most common):
yield scrapy.Request(
url,
meta={
"stealth": {
"driver": "browser",
"headless": False, # visible window β harder to detect (default: True)
"settle": 4.0, # seconds to wait for JS after page load
}
},
)
Heavy Cloudflare sites β increase settle time:
meta={"stealth": {"driver": "browser", "headless": False, "settle": 12}}
Global default (all stealth requests use browser engine):
from scrapy_stealth.config import config
config.STEALTH_DRIVER = "browser"
config.BROWSER_HEADLESS = False # more stealthy
config.BROWSER_SETTLE_S = 6.0 # longer wait for JS
βΈ Custom Browser Binary
Use Brave, Chromium, or a non-default Chrome install by setting the executable path.
Via config:
from scrapy_stealth.config import config
config.BROWSER_EXECUTABLE_PATH = "/usr/bin/brave-browser" # Linux
# config.BROWSER_EXECUTABLE_PATH = r"C:\Program Files\BraveSoftware\Brave-Browser\Application\brave.exe" # Windows
Via settings.py / custom_settings:
BROWSER_EXECUTABLE_PATH = "/usr/bin/brave-browser"
π‘ When BROWSER_EXECUTABLE_PATH is None (default), scrapy-stealth
auto-detects Google Chrome or Chromium from standard system paths. Set it explicitly when using Brave or a
non-standard installation β a clear error is raised if the path does not exist.
βΈ Intelligent Restart / Session Recycle
After STEALTH_RECYCLE_AFTER_BANS
consecutive banned/challenged responses (as classified by Anti-Bot Detection), scrapy-stealth recycles the driver
session:
- browser β restarts Chrome (fresh fingerprint, cookies, CDP session)
- basic / turbo β clears cached HTTP clients/sessions and rotates default
fingerprint profile + proxy from
STEALTH_PROXIES(no per-requestrotate_*needed)
A single clean response resets the streak, so a healthy crawl is never recycled just because it has served a lot of requests.
from scrapy_stealth.config import config
config.STEALTH_RECYCLE_AFTER_BANS = 5 # recycle after 5 consecutive bans (default)
βΈ Docker / Running as Root
Chrome requires --no-sandbox when
the process runs as root. scrapy-stealth detects this automatically,
but you can also set it explicitly.
Via settings.py:
BROWSER_NO_SANDBOX = True # force no-sandbox (Docker, any root environment)
BROWSER_EXECUTABLE_PATH = "/usr/bin/chromium" # use Chromium instead of Chrome in Docker
Via config:
config.BROWSER_NO_SANDBOX = True
config.BROWSER_EXECUTABLE_PATH = "/usr/bin/chromium"
βΈ Static Asset Blocking
Block images, fonts, CSS, and media in the browser to speed up page loads and cut
bandwidth, via the CDP Fetch domain. Off by default. Blocking is always skipped when snapshot=True, since a snapshot needs the fully rendered page.
Global (via settings.py):
BROWSER_STATIC_ASSETS_BLOCK = True
Per-request override:
meta = {"stealth": {"driver": "browser", "static_assets_block": True}}
Snapshot always wins (assets never blocked):
# even with BROWSER_STATIC_ASSETS_BLOCK = True globally, this request loads everything
meta = {"stealth": {"driver": "browser", "snapshot": True}}
βΈ Proxy Bypass List
When a proxy is configured, send specific domains straight to the origin instead
of through the proxy. Passed to Chrome's --proxy-bypass-list launch flag β
supports bare hostnames, wildcards (*.example.com), IP/CIDR ranges, ports,
and <local>.
Via settings.py:
STEALTH_DRIVER = "browser"
STEALTH_PROXIES = ["http://user:pass@proxy:8080"]
BROWSER_PROXY_BYPASS_LIST = [
"example.com", # exact host
"*.internal.net", # wildcard subdomains
"127.0.0.1", # IP
"<local>", # any plain hostname without dots
]
Via config:
from scrapy_stealth.config import config
config.BROWSER_PROXY_BYPASS_LIST = ["example.com", "*.internal.net"]
π‘ The bypass list is a Chrome launch flag β read once at browser start and applies for the whole browser lifetime. Configured globally (config/settings), not per-request. Has no effect unless a proxy is in use.
πΈ Screenshots
Capture a PNG screenshot of any page rendered by the browser driver and save it to disk.
1. Enable on the request
yield scrapy.Request(
url,
meta={
"stealth": {
"driver": "browser",
"snapshot": True,
}
},
callback=self.parse,
)
The raw PNG bytes are available at response.meta["snapshot_content"]
inside your callback.
2. Auto-save with snapshot decorator
from scrapy_stealth.decorators import snapshot
class MySpider(scrapy.Spider):
@snapshot
def parse(self, response): ...
@snapshot(path="stealth_shots/page.png")
def parse(self, response): ...
@snapshot(path=lambda r: r.url.split("/")[-1] + ".png")
def parse(self, response): ...
β οΈ Note: Requires driver="browser"
and snapshot=True in the request meta. Logs an error if no snapshot
data is found in the response.
3. Custom handling (without the built-in helper)
The screenshot is just bytes in
response.meta["snapshot_content"] β do anything you like with it:
def parse(self, response):
shot: bytes | None = response.meta.get("snapshot_content")
if shot is None:
return # screenshot was not requested or capture failed
# Save manually
with open("page.png", "wb") as f:
f.write(shot)
# Pass to a pipeline via item
yield {"url": response.url, "screenshot": shot}
π Automatic Rotation
Profile + proxy stay stable for speed (session reuse). After STEALTH_RECYCLE_AFTER_BANS consecutive bans, the session recycles and a new
default profile + proxy (from STEALTH_PROXIES) are chosen automatically.
# settings.py
STEALTH_PROXIES = ["http://proxy1:8080", "http://proxy2:8080"]
STEALTH_RECYCLE_AFTER_BANS = 5 # default
# spider
yield scrapy.Request(url, meta={"stealth": {}})
Scrapy stats
After the crawl (or mid-run via crawler.stats),
inspect:
| Key | Meaning |
|---|---|
| stealth/requests Β· stealth/requests/{driver} | Stealth fetches |
| stealth/responses Β· stealth/responses/{driver} | Completed responses |
| stealth/successes Β· stealth/successes/{driver} | Non-banned responses below HTTP 400 |
| stealth/failures Β· stealth/failures/{driver} | Banned responses or HTTP 400+ |
| stealth/status/{code} | Response count by HTTP status |
| stealth/bans Β· stealth/bans/{driver} | Session-ban responses |
| stealth/recycles Β· stealth/recycles/{driver} | Session / Chrome recycles |
| stealth/ban_streak | Current consecutive ban streak |
| stealth/driver | Last stealth driver used |
| stealth/profile | Last fingerprint profile used |
| stealth/proxy | Last proxy as host:port (no credentials) |
| stealth/proxy/requests/{driver} | Requests sent through a proxy |
| stealth/dns/requests/{driver} | Requests using DNS overrides |
| stealth/dns/hosts | Total pinned hosts applied |
| stealth/dns/active_hosts | Pinned hosts on latest request |
| stealth/fallbacks | Smart browser selection escalations (basic/turbo β browser) |
# e.g. in spider_closed
stats = spider.crawler.stats.get_stats()
print(stats.get("stealth/bans"), stats.get("stealth/recycles"))
π§© Strategies
Proxy Rotation
from scrapy_stealth.strategies.proxy import ProxyRotator
proxy_rotator = ProxyRotator([
"http://proxy1:8080",
"http://proxy2:8080",
])
yield scrapy.Request(
url,
meta={
"stealth": {
"proxy": proxy_rotator.get(),
}
},
)
Fingerprint Rotation
from scrapy_stealth.strategies.fingerprint import ProfileRotator
fp = ProfileRotator()
yield scrapy.Request(
url,
meta={
"stealth": {
"profile": fp.get(),
}
},
)
Intelligent Retry
from scrapy_stealth.strategies.retry import RetryHandler
retry = RetryHandler()
def parse(self, response):
if retry.should_retry(response):
yield retry.build(response.request)
return
π‘οΈ Anti-Bot Detection
from scrapy_stealth.detectors.antibot import AntiBotDetector
detector = AntiBotDetector()
if detector.is_blocked(response):
print("Blocked!")
π Examples
βΈ Minimal Spider
Smallest working spider with middleware, global stealth mode, and turbo driver:
import scrapy
class ExampleSpider(scrapy.Spider):
name = "example"
custom_settings = {
"DOWNLOADER_MIDDLEWARES": {
"scrapy_stealth.middlewares.StealthDownloaderMiddleware": 950,
},
"STEALTH_ENABLED": True,
"STEALTH_DRIVER": "turbo",
}
def start_requests(self):
yield scrapy.Request("https://example.com")
def parse(self, response):
yield {"title": response.css("title::text").get(), "url": response.url}
βΈ Full Spider Example
Complete working spider from the package repo (examples/full_spider.py). It shows middleware + STEALTH_ENABLED, default turbo driver, per-request basic / browser overrides, optional
snapshot with @snapshot, ban detection, and stealth stats on close.
# from a Scrapy project
scrapy crawl stealth_demo
# or one-off
scrapy runspider examples/full_spider.py
"""Full scrapy-stealth spider example.
Run from a Scrapy project after installing scrapy-stealth::
scrapy crawl stealth_demo
Or as a one-off script::
scrapy runspider examples/full_spider.py
"""
from __future__ import annotations
import scrapy
from scrapy_stealth.decorators import snapshot
from scrapy_stealth.detectors import AntiBotDetector
class StealthDemoSpider(scrapy.Spider):
"""Demonstrates global stealth settings + per-request driver overrides."""
name = "stealth_demo"
custom_settings = {
"DOWNLOADER_MIDDLEWARES": {
"scrapy_stealth.middlewares.StealthDownloaderMiddleware": 950,
},
# Route every request through stealth (no meta needed by default).
"STEALTH_ENABLED": True,
"STEALTH_DRIVER": "turbo", # "basic" | "turbo" | "browser"
# Optional: seeded as engine default; rotated on ban-streak recycle.
# "STEALTH_PROXIES": [
# "http://user:pass@proxy1:8080",
# "socks5://proxy2:1080",
# ],
# Optional: pin hosts to fixed origin IPs.
# "STEALTH_DNS_OVERRIDES": {"example.com": "93.184.216.34"},
"STEALTH_RECYCLE_AFTER_BANS": 5,
"BROWSER_HEADLESS": True,
"BROWSER_SETTLE_S": 4.0,
"BROWSER_STATIC_ASSETS_BLOCK": True,
"LOG_LEVEL": "INFO",
"ROBOTSTXT_OBEY": False,
"TELNETCONSOLE_ENABLED": False,
}
start_urls = [
"https://example.com",
"https://httpbin.org/html",
]
def start_requests(self):
# 1) Default turbo path (STEALTH_ENABLED + STEALTH_DRIVER).
yield scrapy.Request(
self.start_urls[0],
callback=self.parse,
dont_filter=True,
)
# 2) Force basic for a lightweight page.
yield scrapy.Request(
self.start_urls[1],
callback=self.parse,
meta={"stealth": {"driver": "basic"}},
dont_filter=True,
)
# 3) Real browser + snapshot for JS / challenge-heavy pages.
yield scrapy.Request(
self.start_urls[0],
callback=self.parse_browser,
meta={
"stealth": {
"driver": "browser",
"settle": 3.0,
"snapshot": True,
"static_assets_block": False, # keep assets for a real screenshot
}
},
dont_filter=True,
)
# 4) Opt out of stealth for a specific request.
# yield scrapy.Request(
# "https://httpbin.org/get",
# callback=self.parse,
# meta={"stealth": False},
# )
def parse(self, response: scrapy.http.Response):
detector = AntiBotDetector()
if detector.is_blocked(response):
self.logger.warning("Blocked response from %s", response.url)
return
yield {
"url": response.url,
"status": response.status,
"title": response.css("title::text").get(),
"driver": response.flags, # e.g. ['scrapy-stealth', 'turbo']
"bytes": len(response.body),
}
@snapshot # saves PNG under stealth_snapshots/ when snapshot=True was set
def parse_browser(self, response: scrapy.http.Response):
yield {
"url": response.url,
"status": response.status,
"title": response.css("title::text").get(),
"driver": response.flags,
"has_snapshot": bool(response.meta.get("snapshot_content")),
}
def closed(self, reason: str):
stats = self.crawler.stats.get_stats()
self.logger.info(
"Done (%s) β requests=%s successes=%s bans=%s recycles=%s driver=%s",
reason,
stats.get("stealth/requests"),
stats.get("stealth/successes"),
stats.get("stealth/bans"),
stats.get("stealth/recycles"),
stats.get("stealth/driver"),
)
What this covers: middleware + global stealth settings, per-request
basic / turbo /
browser drivers, snapshots + @snapshot, anti-bot detection, and Scrapy stealth stats on
close.
π‘ Keep most traffic on turbo; reserve browser for JS-protected URLs only. Enable STEALTH_AUTO_FALLBACK = True or driver="auto" for automatic escalation.
β‘ Performance Insight
Using stealth selectively offers key structural improvements:
- β‘ Faster crawling (Scrapy for simple pages)
- π° Lower proxy cost
- π‘οΈ Better success rate on protected pages
π Dynamic Release Tracker
Recent Artifact Versions (GitHub API)
| Release Tag | Published Date | Documentation |
|---|---|---|
| Loading telemetry details... | ||
π€ Contributing & Project Links
See CHANGELOG.md for a full history of changes, browse GitHub Releases, or review guidelines in CONTRIBUTING.md to contribute.
π License
This project is licensed under the MIT License β free to use, modify, and distribute. See LICENSE for the full text.