Every residential proxy vendor publishes a pool size. None of them publishes a pool composition — how much of that pool is actually a home broadband line, and how much is a server in a data centre wearing a residential label. That second number is the one that decides whether the residential premium you are paying buys anything at all, and it is measurable in an afternoon with free tools.
This page is both the method and the result: the exact rule, the exact commands, the sample size fixed in advance, and whatever the run produced — published as measured.
Key takeaways
- Pool size is advertised; pool composition is not. The share of a residential pool that exits through hosting networks is the figure that decides whether a residential IP is doing residential work, and essentially no vendor publishes it.
- It is measurable with free, redistributable sources only: Team Cymru's public IP-to-ASN service and PeeringDB (CC BY). No commercial database, no paid API, no licence that forbids commercial use.
- The sample size is fixed BEFORE the run — 1,087 unique exit IPs for a plus-or-minus two point margin at 95% confidence. Choosing n after seeing the data is how a benchmark becomes an advertisement.
- Every share is published with a 95% Wilson confidence interval, and the hosting share is a LOWER BOUND: an address counts as hosting only when its registered network name says so, so neutrally-named operators fall into 'unclassified' instead.
- We hold an affiliate relationship with one of the vendors measured here. The result is published whatever it says, the raw dataset is downloadable, and the method is reproducible against any pool you can authenticate to — including one we have never touched.
- First run, 17 September 2026, DataImpulse US residential: at least 5.1% of 1,087 unique exits came from self-described data centres (95% CI 3.9–6.5%), 43.0% from registered consumer ISPs, and 52.0% from networks public registries cannot classify — over a third of those from four operators that appear on no consumer broadband bill.
What a "residential" label actually promises
A residential proxy is sold on the premise that its exit address belongs to a consumer internet connection, so the traffic looks like an ordinary person browsing rather than a machine in a rack. That is the entire value proposition, and it is why residential bandwidth costs multiples of what datacenter bandwidth costs — the difference between the two is explained in residential vs datacenter proxies, and the underlying mechanics in what is a residential proxy.
What the label does not promise is that every address in the pool satisfies that description. A pool is assembled — from SDK partnerships, from reward apps, from resellers, from other proxy networks — and the further an address sits from the vendor's own supply chain, the less the vendor can say about it. Our reasoning on why that supply chain deserves scrutiny is in ethical proxy sourcing.
So there is a gap between the label and the addresses, the gap is a number, and the number is checkable. Nobody checks it, which is the only reason this page is interesting.
The method, fixed before the run
Four decisions define this measurement. All four were made before any data was collected, and all four are published so the run can be disputed rather than merely believed.
1. Only free, redistributable sources
Origin networks come from Team Cymru's public IP-to-ASN service and network types from PeeringDB, which is CC BY — free to use and republish with attribution, which this page gives.
The obvious easier option was rejected on inspection. ip-api.com returns a direct hosting boolean, which would have been richer signal than anything reconstructed from network names, but its terms restrict the free tier to non-commercial use in a non-commercial environment. This site monetises, so publishing figures derived from it would breach those terms — the same category of shortcut the piece criticises vendors for. Its free tier is also HTTP-only, and a measurement whose entire value is credibility cannot source its inputs over a channel any intermediary can rewrite.
That rejection is not free of cost, and it is the main reason the headline figure is a lower bound rather than an exact share.
2. A conservative classification rule
An address is counted as hosting when the registered name of its origin network contains one of a published list of terms — hosting, cloud, datacenter, colocation, hetzner, digitalocean, vps and similar. An address is counted as consumer when PeeringDB records its network as Cable/DSL/ISP. Everything else is unclassified and stays in the denominator.
Two properties of that rule matter more than the list itself. Hosting wins when both signals fire, because the conservative direction is to not under-report servers in a pool sold as residential. And addresses whose network could not be resolved at all are kept in the denominator rather than dropped — quietly shrinking the denominator would inflate whichever share the reader cares about, which is precisely the move this measurement exists to check somebody else for.
3. A sample size chosen in advance
1,087 unique exit IPs, which is the size that gives a plus-or-minus two percentage point margin at 95% confidence for a proportion near 13%. The measuring script refuses to start without an explicit sample size and prints the full method before the first request, so the number cannot be quietly tuned after the data arrives.
4. Every figure carries its interval
Shares are reported with a 95% Wilson score interval rather than a bare percentage. Wilson rather than the normal approximation because at proportions near zero the normal interval runs below zero and reports a lower bound that cannot exist. A pool statistic without an interval invites the reader to treat it as exact, which is the error the whole piece is arguing against.
Results
| Pool measured | Sample (unique exit IPs) | Hosting-likely (95% CI) | Consumer ISP | Unclassified | Distinct ASNs | Measured |
|---|---|---|---|---|---|---|
| DataImpulse | 1,087 | 5.1%3.9%–6.5% | 43.0% | 52.0% | 131 |
DataImpulse — hosting networks seen most often in the sample: AS394380 LEASEWEB-USA-DAL - Leaseweb USA, Inc., US (11); AS21769 AS-COLOAM - Colocation America Corporation, US (9); AS30633 LEASEWEB-USA-WDC - Leaseweb USA, Inc., US (9).
“Hosting-likely” is a LOWER BOUND: an address counts as hosting only when its registered network name says so, so operators with neutral names land in “Unclassified” instead. Origin networks come from Team Cymru's public IP-to-ASN service and PeeringDB (CC BY). Not comparable with figures derived from a commercial usage-type database.
Every figure above is downloadable as JSON at /data/pool-composition.json, licensed CC BY 4.0, including the sample size, the interval, the exact classification rule each run used, and a per-network tally that lets anyone re-classify the run without re-sampling the pool.
The 17 September 2026 run: DataImpulse, US residential
Setup. DataImpulse's standard residential pool, targeted to the United States in the dashboard, rotating session, default targeting, anonymity filter off, no ASN exclusions — the pool a customer gets without touching a setting. 1,120 proxied requests yielded the 1,087 unique exit addresses fixed in advance. They resolved to 131 origin networks, all classified against a single PeeringDB export of 35,393 networks fetched at 06:12 UTC the same day. The run consumed roughly 1 MB of evaluation traffic that DataImpulse provided free of charge, which is disclosed here because it is the kind of thing that should be.
Hosting: at least 5.1%. 55 of the 1,087 exits came from networks whose registered name says data centre — a 95% interval of 3.9% to 6.5%. The list is short and specific: 45 of the 55 were Leaseweb USA, across eight of its sites (Dallas, Washington, Los Angeles, Miami, Phoenix, Chicago, New York and Seattle); nine were Colocation America; one was Dedicated.com. Read as a lower bound, that is one exit in twenty from a hosting rack, in a pool sold and priced as residential.
Consumer ISPs: 43.0%. 467 exits sat on networks PeeringDB registers as cable, DSL or ISP access — Comcast (109), RCN (98), Charter (81 across six registrations), T-Mobile (40), Cox (28) and Windstream (20) lead. Those counts are scoped to this bucket: Charter holds four further registrations, 22 exits, that PeeringDB gives no type at all, and they sit in the unclassified section below. These are the addresses the residential label promises.
Unclassified: 52.0%, and this time it is a fact about the pool. An earlier attempt at this run had to be discarded when PeeringDB rate-limited the per-network lookups halfway through, which would have made "unclassified" a measurement of the lookup service. This run judged every network against one export, so the 565 unclassified exits decompose cleanly:
- 334 (30.7%) are on networks PeeringDB has no record or no type for. Four operators account for 215 of them — DynaNode LLC (83), Rocks Computer Services (77), eSited Solutions (28) and Flash Edge Information Technology (27). None is a name that appears on a consumer broadband bill; none has a name the published rule recognises as hosting either, so the rule does not count them. Between them they are one exit in five. What they are is the single largest open question in this result.
- 194 (17.8%) are registered as NSP — network service providers such as Verizon Business (61), AT&T (57), Frontier (25) and CenturyLink (9). These carriers also sell consumer lines, so this bucket most likely leans residential; the rule does not assume so.
- 32 (2.9%) are registered as Content — Enzu, Hype Enterprises, Latitude.sh and Psychz Networks among them, operators that are hosting businesses in practice but whose names do not say so.
- 5 (0.5%) fall in one-off registry types — Educational/Research (2), Non-Profit (1), Network Services (1) and Enterprise (1). Too few to move the result, listed so the four groups add up to 565.
Put those pieces together and the true hosting share of this sample sits somewhere between the 5.1% the rule can prove and roughly 8% if the Content bucket is counted, while the consumer share sits between 43% and roughly 61% if the NSP carriers are. The two operators that decide where in that range the truth lies are DynaNode and Rocks Computer Services, 160 exits between them. The classification rule was not amended after these names appeared: adding them now would be choosing the answer after seeing the data, which is the practice this page exists to refuse.
What this does and does not say about Proxyway's figure. Proxyway's April 2026 benchmark put roughly one US exit in eight on a server using IP2Location's usage-type database. This run cannot confirm or contradict that number — different input, different rule, different quantity — and it does not try to. What can be said is that a free, conservative, name-only rule already finds one exit in twenty from named data centres, and that the direction of the two findings is the same.
Vendor response. These figures were sent to DataImpulse before publication, as promised. The company had told us in writing that it does not measure or publish a composition figure of its own. Any reply will be appended here, dated.
Run the check yourself, before you pay
The useful version of this article is the one you run against the pool you are about to buy. Rotate the endpoint, collect unique exit addresses, then ask Team Cymru what network each one came from. The lookup is a plain DNS query, so it needs no key and no account:
# 1. Get an exit IP through the proxy you are evaluating.
ip=$(curl -s --proxy "http://$PROXY_HOST:$PROXY_PORT" \
--proxy-user "$PROXY_USER:$PROXY_PASS" https://api.ipify.org)
# 2. Ask Team Cymru which network announces it (reverse the octets).
reversed=$(echo "$ip" | awk -F. '{print $4"."$3"."$2"."$1}')
asn=$(dig +short TXT "$reversed.origin.asn.cymru.com" | awk -F'"' '{print $2}' | awk '{print $1}')
# 3. Resolve that ASN to the name its operator registered.
dig +short TXT "AS$asn.asn.cymru.com"
Repeat that a few hundred times and count how often the name that comes back reads like a data centre. If you would rather not build the loop, the full pre-registered run — sampling, classification, Wilson intervals and the JSON artifact — is one command in our open method:
pnpm run measure:pool -- --provider=dataimpulse --sample=1087
Two practical warnings. Send the proxy password over stdin or an environment variable, never as a command-line argument: arguments are readable in the process list of a shared machine for as long as the request runs, and this check issues thousands of them. And be polite to PeeringDB — it is a volunteer-run service that allows roughly one anonymous request an hour for its full export, so fetch the whole registry once, keep it on disk, and judge every address in a run against that one snapshot rather than looking networks up one at a time. That is also what makes the run reproducible: the artifact records which snapshot it used.
What a hosting share does and does not tell you
A high hosting share is not proof of dishonesty, and reading it that way is the fastest way to get the analysis wrong.
Some hosting-registered ranges are genuinely assigned to consumer connections. Some consumer ISPs operate their own hosting ranges under the same name. Some addresses are simply unclassifiable from public data, which is why unclassified is a published column and not a rounding error.
What a high hosting share does tell you is commercial. Datacenter bandwidth is cheap and residential bandwidth is not; a pool with a large server component is selling you cheap inventory at the expensive price. It is also operationally worse, because hosting ranges are the first thing a target site's anti-bot vendor buys a list of — which is the mechanism behind most of the blocks described in how to scrape without getting blocked. And if you are collecting data under a compliance obligation, "we could not account for one address in eight" is a sentence you do not want in an audit.
The inverse is equally worth stating: a low hosting share is evidence about composition, not about consent. A pool can be entirely consumer-ISP and still have been assembled from people who never meaningfully agreed to route strangers' traffic. Composition and sourcing are two different questions, and this measurement answers only the first.
Choosing a provider you can actually account for
Until composition figures are routine, the practical proxy for them is whether a vendor will document where its addresses come from and who it will sell to. That is observable today, from published pages rather than sales calls, and it is summarised across the market in the proxy market report. The short version:
Paid link (ad): we earn a commission if you buy.
Bright Data
The strictest published customer vetting we have found: company-only, human-reviewed KYC before residential access. The pick when you need to defend your supply chain to somebody else.
Paid link (ad): we earn a commission if you buy.
Oxylabs
Signup KYC with risk-based escalation, and a founding member of the industry's compliance initiative — externally checkable rather than self-asserted.
Paid link (ad): we earn a commission if you buy.
Decodo
Publishes both a vetting policy and an explicit list of target categories it refuses to serve; prices exclude VAT.
And the vendor whose pool this benchmark was pointed at first:
Paid link (ad): we earn a commission if you buy.
DataImpulse
Measured here because it is the pool we hold an evaluation account on, and because it told us in writing that it does not measure or publish its own composition figure. Disclosure: this is an affiliate link, and the result above is published exactly as measured.
Choosing between them on price rather than posture is a different exercise — that one is in best residential proxies and proxy pricing comparison.
Limitations
Stated plainly, because a method section that hides its weaknesses is marketing.
The hosting share is a lower bound, not an estimate of the true share. The rule only fires on self-describing network names, so every neutrally-named hosting operator is missed.
These figures are not comparable with any number derived from a commercial usage-type database, including the Proxyway benchmark cited above. Different inputs, different rule, different quantity. Presenting them as two readings of the same thing would be wrong even when they agree.
A pool is not static. A sample describes the addresses that pool handed out during one run, from one vantage point, at one moment. Composition can differ by target country, by time of day, and by how much you are spending.
And the sample says nothing about success rate, latency or consent. It is one property of one pool, measured carefully. Treat it as that and it is useful; treat it as a verdict on a vendor and it is not.
Cite this measurement
Hinata Tomoda (2026). Are Residential Proxies Really Residential? How to Measure Any Pool. ProxyFacts. https://proxyfacts.com/blog/residential-proxy-pool-composition
Machine-readable dataJSON
Free to cite and reuse with attribution (CC BY 4.0). Machine-readable data: /data/pool-composition.json. Method published 2026-08-27; each measurement carries its own run date and the classification rule in force at the time.
This is an independent measurement, not a vendor-supplied statistic, and not legal advice. If you re-run the method and get a different answer, we want the correction — that is what publishing the rule is for.