Build a Rotating Proxy Pool with Multiple VPS IPs
Cover Image

Introduction
One VPS running Squid carried you through the early days. Then volume grew: more targets, more requests, tighter rate limits. Suddenly one IP — or even three on the same box — is not enough, and a single ASN-wide block takes your entire operation dark at once.
To build a rotating proxy pool with multiple VPS IPs, stand up N cheap Squid nodes, point one front proxy at them as parents with round-robin distribution, add health checks that route around dead nodes, and size the pool with per-IP rate math. This guide outlines a four-parent fleet with a separate front proxy — the natural sequel to our single-box rotation guide.
💡 TL;DR: This example uses four parent VPS nodes plus a separate front proxy. Configure each parent’s listener, access controls, and peer authentication; use
cache_peer ... no-query round-robinat the front. Verify exits with fresh connections, monitor reachability and destination responses separately, and size the pool from permitted rates and measured capacity.
This configuration sketch targets Squid versions whose linked directives are supported (through v7; the references mark them unavailable in v8). It has not been deployment-tested. Back up per-node configurations and validate on a test fleet first.
Assumes Squid with authentication already working. Configure each parent to listen on its private or VPN address; a loopback-only listener cannot accept connections from the front VPS. Allow the front’s address through host and provider firewalls and the parent’s access policy. Adapt and back up each node’s configuration before copying it; do not copy the front’s parent-routing configuration to the parents. New here? Start with proxy setup, then rotation, then this post.
When One Box Stops Being Enough
Three reasons can justify a pool: network-level filtering, a shared failure domain, and measured capacity limits. Another IP in the same autonomous system may still be affected by an ASN-wide block; a different region does not necessarily mean a different ASN. One server also shares its operating system, account, and network risks across its addresses. Measure CPU, bandwidth, connections, and latency under your workload before deciding the proxy is the bottleneck. In an ordinary HTTPS CONNECT tunnel, the client and destination perform the TLS handshake; Squid forwards the encrypted traffic.
Think of it like delivery trucks. One truck with rotating plates (single-box rotation) fools the toll camera for a while. A fleet of trucks from different depots (multi-VPS pool) survives the camera learning any single plate — and keeps delivering when one truck breaks down. Different failure domains are the whole point.
Consider a pool when separate network failure domains or measured capacity needs justify its extra cost and maintenance. Adding a third IP or a second target alone does not establish that another server is necessary.
Architecture: Front Proxy Plus Parent Nodes
The design is hub-and-spoke. Scrapers talk to exactly one endpoint — the front Squid. The front forwards to N parent Squids, each on its own VPS with its own public IP, cycling via round-robin. Clients use the same front endpoint when a parent is added. The new parent still needs provisioning, access controls, peer configuration, and a routing check.
# fleet layout: scrapers see one endpoint, front spreads load
[scrapers] --> [front Squid :3128] --> [node1 203.0.113.10]
--> [node2 198.51.100.20]
--> [node3 192.0.2.30]
--> [node4 203.0.113.40]
Spread nodes across providers where separate network failure domains matter, and verify the actual ASNs. Different regions can share an ASN. Size each VPS from measurements of connection load, bandwidth, memory, and latency; a low-cost instance is a starting candidate, not a capacity guarantee.
💡 Tip: Keep the front proxy and one parent in different failure domains too. If the front dies, the pool is unreachable regardless of healthy parents — back up its config (
squid.conf+ passwd file) somewhere you can use for a tested recovery procedure.
Build: Nodes, Then Front Round-Robin
Step 1 — Provision and harden each node. Fresh Ubuntu LTS, SSH keys, UFW default-deny, automatic updates. Identical base images keep debugging sane — bake one snapshot and clone it.
Step 2 — Install Squid identically on every node. Same version, same auth file structure, same header-hygiene rules from our header leak fix. Check header handling at every parent: an HTTP-forwarding hop can append identifying headers. Ordinary HTTPS CONNECT traffic has encrypted inner headers, which need separate client-side checks. Use scp or a shell loop to push config changes to all nodes at once — manual per-node editing can introduce drift.
# prepare parent-configs/node1.conf through node4.conf for each parent first
# back up the remote files before copying; parse, then reload sequentially
for h in node1 node2 node3 node4; do
scp "./parent-configs/$h.conf" "root@$h:/etc/squid/squid.conf"
ssh root@$h 'squid -k parse && systemctl reload squid'
done
Step 3 — Wire the front with round-robin parents. On the front node, cache_peer with round-robin selects parents in rotation when ICP queries are absent; set no-query explicitly. weight=N biases toward fatter pipes:
# front squid: spread load across 4 parents, node1 has a larger selection weight
cache_peer 203.0.113.10 parent 3128 0 no-query round-robin weight=2 name=node1
cache_peer 198.51.100.20 parent 3128 0 no-query round-robin weight=1 name=node2
cache_peer 192.0.2.30 parent 3128 0 no-query round-robin weight=1 name=node3
cache_peer 203.0.113.40 parent 3128 0 no-query round-robin weight=1 name=node4
never_direct allow all
The example addresses are documentation placeholders. Remove conflicting always_direct allow rules when using never_direct allow all; if no parent is eligible, this design should fail rather than fall back to direct origin access. Replace them with reachable private or VPN endpoints and keep proxy ports restricted to trusted clients. For authenticated parents, add login=PEER_USER:PEER_PASSWORD to each cache_peer line using the actual peer credentials, URL-escaping special characters as documented. Matching password files alone does not make Squid forward credentials. Protect the configuration file and use an encrypted inter-server path: Basic authentication to a plain HTTP proxy is not encrypted by the destination’s HTTPS connection. Keep the front proxy’s client authentication and the parents’ access controls in place.
Size the Pool: Per-IP Rate Math
Pool sizing starts with a capacity estimate. Use each destination’s permitted rate and your measured per-node capacity; limits may also apply to accounts or total activity. Divide the planned rate by the applicable per-node allowance, round up, and then choose spare capacity for the failures you intend to tolerate. Honor retry guidance and reduce traffic when access is denied.
# pool sizing: total RPM divided by safe per-IP RPM, plus headroom
pool_size = ceil(total_requests_per_min / allowed_per_ip_per_min) + spare_nodes
# example: 200 RPM total, 50 RPM allowed per IP -> 4 + 1 spare = 5 nodes
One spare is an illustrative resilience choice, not a universal requirement. It covers one unavailable node only if the remaining nodes can carry the load within the relevant limits. Recompute when the workload, failure assumptions, or destination limits change.
⚠️ Warning: Never size by IP count alone without the rate math. Ten nodes at 100 RPM each against a 20-RPM-per-IP target is still 5x over the limit on every node ; that exceeds the stated rate allowance regardless of whether the destination immediately blocks the addresses. Rate per IP is the variable; node count is just how you afford it.
Health Checks and Failover
Round-robin selects among eligible parents, but parent reachability does not establish success at a particular destination. Existing HTTPS CONNECT tunnels keep their selected parent; a reachable parent can still receive a site-specific denial. Separate those outcomes in monitoring.
Liveness (is the node up?): check a canary endpoint through each parent. You might run this once a minute and require two consecutive failures before removing a parent; those are example thresholds. A controller needs state, safe configuration updates, a parse check before reload, and a recovery policy. A single curl command below only performs a probe; it does not implement that controller.
# one manual canary probe; failover requires a separate controller
curl --proxy http://PARENT_VPN_IP:3128 --proxy-user PEER_USER --fail --show-error --silent --max-time 10 --output /dev/null --write-out "%{http_code}" https://example.com
Quality (is the node working well?): compare destination-specific failures in application logs. A high 403 rate warrants investigation; it does not prove an IP block. With ordinary HTTPS CONNECT tunneling, Squid’s access log records the tunnel outcome rather than the encrypted origin response status. Collect those HTTP statuses in the client. Centralized proxy logs still help diagnose forwarding and connection failures.
Fleet-Wide Hygiene and Logging
Three things must stay identical across nodes or the fleet rots: header rules (one leaky node unmasks shared traffic), auth policy (manage peer credentials consistently, even when passwords differ) , and Squid versions (directive syntax drifts between majors — pin versions in apt). The scp loop above handles config; add version output (squid -v) to your monthly checklist.
Centralize logs from day one. Per-node access.log forwarding to the front (or a tiny Loki/Graylog instance) can bring forwarding records together. The available fields depend on the log format; use client records for encrypted origin response codes and compare both when investigating failures.
Crossover Math: When Paid Beats Self-Hosted
Compare self-hosting and vendor costs for the same workload and permitted access requirements: infrastructure, bandwidth, monitoring, maintenance time, and failure handling versus the vendor’s quoted charges. Neither low volume nor high volume guarantees that one option costs less. Use actual quotes and your recorded maintenance effort.
As a hypothetical calculation, four parents plus a separate front proxy at $5 each cost $25 per month for instances alone. Add bandwidth, taxes, monitoring, and maintenance at your chosen hourly rate. That is not an all-in quote or a throughput benchmark. Compare the complete figure with vendor quotes before deciding whether to switch.
