快速答案
如何在Python中使用代理?
向你的HTTP客户端传递代理URL,如http://USERNAME-country-ae:[email protected]:7777。Nameless Proxy,一个注重隐私、无需 KYC 的代理提供商,支持requests, httpx, aiohttp 和 Scrapy,并在端口7778支持socks5h://。目标信息包含在用户名中。住宅流量起价为US$0.53/GB。无需 KYC — 无需身份证、自拍和文件。支持加密货币支付。
Python快速设置
| 代理URL | http://USER-country-xx:[email protected]:7777 |
|---|---|
| SOCKS5 | socks5h://…:7778搭配requests[socks] |
| requests | proxies={"http": url, "https": url} |
| httpx 0.26+ | httpx.Client(proxy=url)和httpx.AsyncClient(proxy=url) |
| aiohttp | session.get(url, proxy=proxy_url)(HTTP代理) |
| Scrapy | request.meta["proxy"],通过下载中间件设置 |
| 粘性IP | 在城市后添加 -session-{id}-ttl-{minutes},最长 60 分钟 |
如何用 requests 使用代理?
给 requests 的 http 和 https 键都设置相同的代理 URL。HTTPS 目标通过 CONNECT 隧道,所以 TLS 连接只在你的脚本和网站之间。务必设置超时。
# pip install requests
import requests
USERNAME = "USERNAME" # base proxy username from the client area
PASSWORD = "PASSWORD" # proxy password (not your account password)
HOST = "gw.namelessproxy.com"
HTTP_PORT = 7777
SOCKS_PORT = 7778
proxy = f"http://{USERNAME}-country-ae:{PASSWORD}@{HOST}:{HTTP_PORT}"
proxies = {"http": proxy, "https": proxy}
r = requests.get("https://api.ipify.org?format=json", proxies=proxies, timeout=30)
print(r.json()) # {'ip': '...'}如何构建定向且固定的代理 URL?
每个库用一个助手。用户名跟在 USERNAME-country-{cc}-city-{city}-session-{id}-ttl-{minutes} 后面,助手按顺序添加:国家、城市、会话和 TTL。没有会话则 IP 会轮换;有会话则 IP 在 TTL 内保持不变。
import uuid
def proxy_url(country, city=None, session=None, ttl=30, scheme="http"):
"""Build a gateway URL. Targeting lives in the username, in this order:
country -> city -> session -> ttl."""
user = f"{USERNAME}-country-{country.lower()}"
if city:
user += f"-city-{city.lower().replace(' ', '_')}"
if session:
user += f"-session-{session}-ttl-{ttl}"
port = SOCKS_PORT if scheme.startswith("socks") else HTTP_PORT
return f"{scheme}://{user}:{PASSWORD}@{HOST}:{port}"
def new_session():
"""A fresh session ID = a fresh sticky IP."""
return uuid.uuid4().hex[:10]
# Rotating IP in Egypt, and a sticky IP in Abu Dhabi for 30 minutes
rotating = proxy_url("eg")
sticky = proxy_url("ae", city="Abu Dhabi", session=new_session(), ttl=30)如何在 Python 中使用 SOCKS5?
安装 SOCKS 扩展并使用 socks5h:// 协议。h 让代理解析主机名,DNS 查询也会从目标国家出口。普通的 socks5:// 则在本机解析 DNS。
# pip install "requests[socks]"
# socks5h:// = the proxy resolves DNS, so lookups also exit in the target country
proxy = proxy_url("ke", scheme="socks5h")
r = requests.get("https://api.ipify.org", proxies={"http": proxy, "https": proxy}, timeout=30)
print(r.text)如何在新 IP 上重试失败请求?
连接错误、429 和 5xx 响应时重试,使用指数退避和随机延迟。用户名轮换时,每次新连接都可能使用新 IP,重试即是新 IP。固定会话时,IP 失败则启动新会话 ID。
import random
import time
RETRY_STATUS = {429, 500, 502, 503, 504}
def fetch(url, country="ae", attempts=4):
"""Retry with exponential backoff. Each attempt opens a new connection,
so a rotating username exits from a new IP."""
for attempt in range(attempts):
proxy = proxy_url(country)
try:
r = requests.get(url, proxies={"http": proxy, "https": proxy}, timeout=30)
if r.status_code not in RETRY_STATUS:
return r
except (requests.exceptions.ConnectionError, requests.exceptions.Timeout):
pass # ProxyError is a ConnectionError: retry on a new IP
time.sleep(min(2 ** attempt + random.random(), 30))
raise RuntimeError(f"Giving up on {url} after {attempts} attempts")如何使用 httpx 的同步和异步?
httpx 0.26 起,客户端只接受一个 proxy= 参数;旧的 proxies= 在 0.28 被移除。客户端保持连接,多个请求共享同一出口 IP。根据需求新建客户端或故意使用固定会话。
# pip install httpx (0.26 or later: the argument is "proxy", not "proxies")
import asyncio
import httpx
# Sync: one sticky IP in Nairobi for 10 minutes
with httpx.Client(proxy=proxy_url("ke", city="nairobi", session=new_session(), ttl=10), timeout=30) as client:
print(client.get("https://api.ipify.org").text)
# Async: 20 connections at most, rotating IPs in Nigeria
async def main(urls):
limits = httpx.Limits(max_connections=20)
async with httpx.AsyncClient(proxy=proxy_url("ng"), timeout=30, limits=limits) as client:
results = await asyncio.gather(*(client.get(u) for u in urls), return_exceptions=True)
for url, res in zip(urls, results):
print(url, res.status_code if isinstance(res, httpx.Response) else repr(res))
asyncio.run(main(["https://api.ipify.org"] * 5))如何用 aiohttp 发送大量请求?
aiohttp 每个请求设置代理,凭证放在 URL 中。只支持 HTTP 代理;SOCKS5 需额外包如 aiohttp-socks。用信号量限制并发:初始 10-20 个并行请求,错误率低时可提升。
# pip install aiohttp (HTTP proxies only; the credentials can sit in the proxy URL)
import asyncio
import aiohttp
async def fetch(session, url, sem):
async with sem:
async with session.get(url, proxy=proxy_url("eg")) as r:
return r.status, await r.text()
async def main(urls):
sem = asyncio.Semaphore(20) # concurrency cap
timeout = aiohttp.ClientTimeout(total=30)
async with aiohttp.ClientSession(timeout=timeout) as session:
results = await asyncio.gather(*(fetch(session, u, sem) for u in urls), return_exceptions=True)
print(results)
asyncio.run(main(["https://api.ipify.org"] * 5))如何给 Scrapy 添加代理?
写个小的下载中间件设置 request.meta["proxy"]。Scrapy 内置 HttpProxyMiddleware 会把凭证转成 Proxy-Authorization 头。中间件优先级设低于 750 以先运行。请求元数据可让同一爬虫混用国家和固定会话。
# myproject/middlewares.py
USERNAME = "USERNAME"
PASSWORD = "PASSWORD"
GATEWAY = "gw.namelessproxy.com:7777"
class NamelessProxyMiddleware:
"""Set the proxy per request from request.meta:
country (default 'ae'), optional city, optional session + ttl."""
def process_request(self, request, spider):
meta = request.meta
user = f"{USERNAME}-country-{meta.get('country', 'ae')}"
if meta.get("city"):
user += f"-city-{meta['city']}"
if meta.get("session"):
user += f"-session-{meta['session']}-ttl-{meta.get('ttl', 10)}"
# Scrapy's HttpProxyMiddleware turns the credentials into Proxy-Authorization
meta["proxy"] = f"http://{user}:{PASSWORD}@{GATEWAY}"# myproject/settings.py
DOWNLOADER_MIDDLEWARES = {
# Below 750 so it runs before Scrapy's built-in HttpProxyMiddleware
"myproject.middlewares.NamelessProxyMiddleware": 350,
}
CONCURRENT_REQUESTS = 16
DOWNLOAD_TIMEOUT = 30
RETRY_TIMES = 3
RETRY_HTTP_CODES = [429, 500, 502, 503, 504, 408]
AUTOTHROTTLE_ENABLED = True
ROBOTSTXT_OBEY = True
# In a spider:
# yield scrapy.Request(url, meta={"country": "sa", "city": "riyadh", "session": "cart01"})常见问题
为什么我在 Python 中会遇到 407 或 ProxyError?
407 表示网关拒绝了凭证。请检查客户区中的基础用户名和代理密码,目标部分的顺序,以及密码中的任何特殊字符,这些字符必须进行 URL 编码。没有 407 的 ProxyError 通常意味着端口错误或防火墙问题。使用 快速入门指南] 中的 curl 测试相同的 URL。
为什么我的 IP 在请求之间没有变化?
您的客户端可能在重复使用保持连接。requests.Session、httpx 客户端或 aiohttp 会话会保持连接打开,而一个连接对应一个出口 IP。若想每次请求使用新 IP,请调用requests.get且不使用会话,或打开新的客户端。如果您添加了会话 ID,IP 会被固定。
我应该使用 requests、httpx 还是 aiohttp?
简单脚本和小型任务使用 requests。需要同时支持同步和异步代码或 HTTP/2 时使用 httpx。大规模异步爬取且自行控制并发时用 aiohttp。需要完整爬虫功能(调度、重试、管道)时用 Scrapy。四者均支持相同的代理 URL。
一个脚本可以同时使用多个国家的代理吗?
可以。目标定位是每个代理 URL 的一部分,因此每个请求或客户端构建不同的 URL:例如 proxy_url("ae") 代表阿联酋,proxy_url("ng") 代表尼日利亚。无论国家,流量均从同一余额扣费,且每 GB 价格不随国家变化。
爬虫使用多少流量?
取决于页面内容:HTML 页面通常几十到几百 KB,图片和脚本会增加更多。流量以十进制 GB 计费(1 GB = 1,000,000,000 字节),客户区显示每个订阅和子用户的每日使用量。只抓取需要解析的内容,跳过图片以降低成本。
Python 代码能用专用移动端口吗?
可以,使用客户区提供的专用端口的主机、端口、用户名和密码,替代网关值。专用端口的国家和运营商在租用时选择,之后可更换。想从 Python 更换 IP,请参见移动 IP 轮换指南]。
通过住宅代理运行您的 Python 代码
来自 US$0.53/GB,支持加密货币支付。无需 KYC — 无需身份证、自拍或文件。