TUTORIAL / WEB DATA FUNDAMENTALS

What Is Concurrency Control? Bounding Parallel Scraping Requests

Nstdata WikiTutorial

Concurrency control governs how many requests a scraper sends in parallel at any given moment — too few and throughput suffers, too many and a target site's rate limits trigger or the scraper's own resources get exhausted. Semaphores, thread pools, and async event loops are the common mechanisms for enforcing a bound, and the right bound depends on the target's tolerance, not just the scraper's own raw capacity.

⚡ Key Takeaways

  • Concurrency control bounds how many requests run in parallel, independent of whether those requests reuse pooled connections or not.
  • A semaphore is the standard mechanism for capping in-flight requests in both threaded and async code, blocking new requests until a slot frees up.
  • Async I/O generally outperforms thread pools above roughly 50-100 concurrent requests for I/O-bound scraping workloads, though the exact crossover shifts with latency and payload size.
  • Concurrency limits interact with connection pooling — some clients silently open extra connections when the pool is busy rather than enforcing the intended cap through queueing.
  • HTTP/2 changes the concurrency math: with multiplexing, one connection can serve many concurrent requests, so connection-based limits don't map directly to request-based limits anymore.
  • Per-host bounded concurrency is the standard pattern for polite crawling — capping parallelism per target host, not just globally across all targets.

What Is Concurrency Control?

Concurrency control is the practice of bounding how many requests, tasks, or operations run simultaneously, rather than letting a program launch every available request all at once. In a scraping context specifically, it's the mechanism that keeps a fetcher from overwhelming a target server with a burst of hundreds or thousands of simultaneous connections, while still processing enough requests in parallel to make real progress on a large collection job.

Common Concurrency Control Mechanisms

MechanismHow it bounds parallelism
SemaphoreA counter that permits a fixed number of concurrent operations; a new request blocks until a permit frees up.
Thread poolA fixed pool of worker threads; the pool size itself becomes the concurrency ceiling.
Async event loop with bounded tasksA single-threaded event loop processing many tasks, with a semaphore or queue capping how many run concurrently at once.
Connection pool limitsCapping the pool size itself indirectly bounds concurrency, though behavior varies by client when the pool is exhausted.

A semaphore is the most explicit and portable version of this pattern: wrapping each request in an acquire/release around the semaphore guarantees no more than the configured number of requests are ever in flight simultaneously, regardless of how many total requests are queued up waiting.

Async I/O vs. Thread Pools for I/O-Bound Scraping

For I/O-bound work like web scraping — where a request spends most of its time waiting on network response rather than doing CPU-bound computation — async I/O (a single event loop processing many concurrent tasks) generally outperforms a thread-pool approach above roughly 50-100 concurrent requests, though the exact crossover point shifts based on latency and response payload size for the specific target being scraped. Below that range, the difference is often small enough that either approach performs adequately, and simplicity of implementation can reasonably be the deciding factor instead of raw throughput.

Per-Host vs. Global Concurrency Limits

A global concurrency cap alone — "never more than 50 requests in flight across everything" — doesn't protect any individual target site from an unfair share of that total if a crawl happens to concentrate on one host. Politeness-aware crawling instead applies concurrency limits per host specifically, ensuring no single target receives an overwhelming burst even when the crawler's aggregate concurrency across all targets is high. This per-host bounding is standard practice in production crawler architectures for exactly this reason — a global-only limit optimizes for the crawler's own resource use, while a per-host limit additionally protects each individual target from disproportionate load.

Per-host concurrency, tuned automatically

Nstdata Crawl applies bounded, per-host concurrency automatically across every request, balancing throughput against each target's actual tolerance without manual semaphore tuning.

Try Nstdata Crawl →

Concurrency Control vs. Adjacent Concepts

Connection pooling and concurrency control are related but distinct: pooling governs whether a connection gets reused, concurrency control governs how many requests run in parallel regardless of connection reuse. Request throttling is a third, related but separate concept — throttling paces requests over time (adding delays between them), while concurrency control bounds how many are simultaneously in flight at any one instant; a scraper can have low concurrency but no throttling (few requests at once, but back-to-back with no delay), or high concurrency with heavy throttling (many requests allowed in parallel, but each individually delayed), and production systems typically tune both together rather than relying on either alone.

Limits

Setting concurrency too high relative to a target's actual tolerance risks triggering rate limits or an outright block regardless of how well-paced any individual request is, since concurrent volume itself is a signal detection systems watch. Setting it too low leaves collection throughput far below what a target would actually tolerate, needlessly extending job completion time. There's no universal correct value — it depends on the target's specific infrastructure and defenses, which is why adaptive approaches that adjust concurrency based on real-time response signals (similar to adaptive proxy rotation) generally outperform a fixed concurrency setting chosen once in advance.

Conclusion

Concurrency control bounds parallel requests through mechanisms like semaphores, thread pools, or async task limits, and the right bound is target-specific rather than a universal constant — per-host limits protect individual targets even when overall crawler throughput is high. It works alongside, not instead of, connection pooling and request throttling; a well-tuned scraper coordinates all three rather than relying on any single lever.

For collection workloads that need concurrency automatically tuned per target, evaluate Nstdata Crawl against your own use case.

Try Nstdata Crawl for automatic concurrency tuning

Per-host limits balanced against target tolerance, without manual configuration.

Try Nstdata for Free →

FAQ

Q: What's the difference between concurrency control and connection pooling?

Connection pooling governs whether connections are reused or recreated. Concurrency control governs how many requests run in parallel at once — the two are independent and typically need separate tuning.

Q: Should I use async I/O or a thread pool for concurrent scraping?

For I/O-bound scraping, async I/O generally outperforms thread pools above roughly 50-100 concurrent requests, though the exact crossover depends on target latency and payload size.

Q: Why do I need per-host concurrency limits instead of just a global limit?

A global limit alone doesn't prevent one target site from receiving a disproportionate burst if a crawl happens to concentrate there. Per-host limits protect each individual target regardless of the crawler's overall aggregate throughput.

Q: What's the difference between concurrency control and request throttling?

Concurrency control bounds how many requests are simultaneously in flight. Throttling paces requests over time with delays between them. They're independent settings usually tuned together.

Q: Is there a universal correct concurrency setting for scraping?

No. The right value depends on the specific target's infrastructure and defenses — adaptive approaches that adjust based on real-time response signals generally outperform a fixed value chosen once in advance.

Was this guide helpful?

Your choice is saved on this device.