Crawlee Data Storage Types: Dataset, KVS, RequestQueue
TL;DR
Crawlee uses three core storage abstractions: Dataset for append-oriented result records, KeyValueStore for named values and artifacts, and RequestQueue for scheduled and deduplicated requests.
Named storage is the safer choice when data must survive across runs; unnamed default storage is convenient for temporary jobs and can be purged at startup.
Storage types should follow access patterns, not convenience. Mixing crawl state, binary artifacts, and result rows in one store makes retries and retention harder.
A production crawler should define stable record keys, idempotent writes, retention, and recovery before scaling concurrency.
Nstdata Crawl can provide the managed acquisition layer while Crawlee storage organizes application-side requests, artifacts, and accepted records.
What are the Crawlee data storage types?
Crawlee's primary storage types are Dataset, KeyValueStore, and RequestQueue. A Dataset stores result records that are typically appended during a crawl. A KeyValueStore stores values addressed by keys, including configuration, crawl state, or larger artifacts. A RequestQueue stores URLs or requests to process and prevents duplicate scheduling according to the queue's request identity rules.
When a team wants managed acquisition while retaining its own application-side storage model, Nstdata Crawl can supply the page or bounded-site collection layer.
The official Crawlee for Python storage guide distinguishes high-level storage interfaces from the storage clients that persist their data. The JavaScript implementation uses the same conceptual split, but teams should read the documentation for their chosen language and version rather than assuming APIs are identical.
When should you use Dataset?
Use Dataset for append-oriented crawler outputs that downstream code needs to iterate, export, or analyze. Examples include product records, article metadata, accepted extraction results, and page-quality measurements. Each record should contain enough provenance to stand alone: canonical URL, retrieval timestamp, status, content hash, schema version, and validation outcome.
Experience Nstproxy Crawl - Start Your Free Trial Today
Dataset is not ideal for mutable singleton state or for repeatedly updating a large artifact under one key. If a workflow needs to save the latest cursor, configuration, or screenshot, KeyValueStore is usually the clearer abstraction. If it needs to schedule URLs with deduplication, RequestQueue is the correct choice.
When should you use KeyValueStore?
Use KeyValueStore for values retrieved by a known name. Typical examples include input configuration, a checkpoint, a robots-policy snapshot, a screenshot, raw HTML, or a JSON summary of the run. The key should be deterministic and documented so a recovery process can locate it without scanning unrelated records.
KeyValueStore can hold structured values, but it should not become an unindexed substitute for a result dataset. Define retention separately for temporary debugging artifacts and durable compliance evidence. Large artifacts may need external object storage depending on the selected Crawlee storage client and deployment environment.
When should you use RequestQueue?
Use RequestQueue for requests that a crawler still needs to process. The queue supports discovery, prioritization, and deduplication, allowing a crawler to add links without processing the same request repeatedly. Store request-specific context such as discovery source, depth, label, retry count, and business identifier in the request's user data where the documented API supports it.
Request identity needs deliberate design. Query-string order, tracking parameters, fragments, and canonical redirects can create distinct URLs that represent the same content. Normalize only parameters that are known not to change the resource. Over-aggressive normalization can merge pages that should remain separate.
Turn Web Pages into Usable Data
Use Nstdata Crawl to convert a URL into clean outputs for AI, RAG, and data workflows.
Named storage is intended to be discoverable and reusable across runs, while unnamed or default storage is often scoped to a run and may be purged at startup according to configuration. Named stores are appropriate for scheduled jobs, handoffs, backfills, and debugging across process restarts. Unnamed stores are convenient for tests and disposable runs.
The exact lifecycle depends on the active storage client and environment. Confirm purge behavior before relying on default storage for recovery. The official Crawlee for JavaScript result-storage guide documents local persistence behavior for that implementation.
For implementation details and current releases, use the official Crawlee repository rather than copying storage behavior from an undated tutorial.
How should the three storage types work together?
A clean design uses RequestQueue for work, KeyValueStore for run state and artifacts, and Dataset for accepted result rows. This separation makes recovery understandable:
RequestQueue shows what is pending, handled, or eligible for retry.
KeyValueStore preserves configuration, checkpoints, and diagnostic artifacts.
Dataset contains records that passed the pipeline's acceptance rules.
Do not push a page into Dataset simply because the request completed. Validate content type, canonical URL, required fields, and business rules first. A separate rejected-record dataset or diagnostic key can preserve failures without contaminating accepted data.
How do you make Crawlee storage reliable in production?
Reliability begins with idempotency. Assign a stable business key or document ID so retries do not create uncontrolled duplicates. Store a content hash to detect changes. Version the schema and crawler configuration so a record can be interpreted after the code changes.
Define checkpoint boundaries. A crawler should know whether a request can be marked handled before result storage succeeds. If result persistence fails after a request is acknowledged, the pipeline can lose data. Use the failure and retry semantics documented by the active crawler and storage client.
Monitor queue depth, oldest pending request, handled rate, retry rate, dataset write failures, artifact size, and storage latency. Nstdata's guide to scaling web scraping provides broader operational context, while its web data pipeline guidance shows why accepted business records should be separated from raw retrieval.
The website URL discovery guide is also useful when deciding which discoveries belong in RequestQueue and which should be rejected before scheduling.
How can Nstdata Crawl work with Crawlee storage?
Nstdata Crawl can serve as a managed acquisition service while a Crawlee application manages its own queue, artifacts, and result records. This pattern is useful when the application needs custom scheduling or business logic but does not want to operate every browser and routing component.
Put authorized target requests and business context in RequestQueue.
Store crawl configuration and diagnostic artifacts in KeyValueStore.
Call the managed collection layer within bounded concurrency.
Validate the returned page or task result.
Push only accepted, attributable records to Dataset.
Use Nstdata Crawl pricing to confirm the current billing model and include failed or rejected pages when calculating cost per accepted record. For current product behavior, consult the Nstdata Crawl documentation.
What storage mistakes cause the most trouble?
Common mistakes include relying on default purge behavior without checking it, storing large binary artifacts as ordinary result rows, acknowledging requests before durable result writes, and using raw URLs as identity without canonicalization. Another frequent problem is keeping no schema version, which makes old records ambiguous after a deployment.
Security and privacy also apply to crawler storage. Do not place credentials, authentication cookies, or sensitive headers in RequestQueue user data or logs. Minimize personal data, define retention, and restrict access to artifacts that may contain sensitive page content.
Conclusion
Crawlee storage types are simple when each one has a clear responsibility: RequestQueue schedules work, KeyValueStore holds named state or artifacts, and Dataset stores accepted rows. Define lifecycle, identity, and failure behavior before increasing concurrency. If browser and network acquisition are the main operational burden, pair application-side Crawlee storage with a managed collection layer such as Nstdata Crawl and keep the acceptance contract in your own system.
The three main types are Dataset, KeyValueStore, and RequestQueue.
Q: Which Crawlee storage type should hold scraped records?
Dataset should usually hold append-oriented scraped records that passed validation.
Q: Which Crawlee storage type deduplicates URLs?
RequestQueue manages scheduled requests and deduplicates them according to the request identity used by the implementation.
Q: Are default Crawlee storages persistent?
Default-storage lifecycle depends on the storage client and purge configuration, so teams must verify startup behavior before relying on it for recovery.
Q: Can KeyValueStore hold screenshots or raw HTML?
Yes, KeyValueStore is suitable for named artifacts, subject to the capabilities and size limits of the selected storage backend.
Q: How do you prevent duplicate Dataset records?
Use stable record identifiers, canonical URLs, content hashes, and idempotent application logic rather than relying on Dataset append behavior alone.
Lena Zhou
Aug. 11th 2026
Crawl entire websites with a single API request
99.8% success rate with JavaScript rendering
Get clean, LLM-ready data in multiple formats
Turn any website into Markdown, HTML, JSON, links, PDFs and more — without managing crawling infrastructure.