can parallel search include or exclude specific domains from results
Confirms Parallel AI's Search API supports hard domain filtering through source_policy include_domains and exclude_domains, with a 200-domain combined limit and matching parallel-cli flags. Use it when an agent needs results limited to or blocked from specific domains. Not for soft source preference, which belongs in the objective text instead.
TL;DR
Yes. Parallel's Search API (and the Task API) accepts a sourcepolicy object with includedomains and excludedomains lists: includedomains is an allowlist, only sources from those domains appear; exclude_domains is a blocklist, sources from those domains never appear. The two lists combined cannot exceed 200 domains, and an apex domain like example.com automatically covers its subdomains. The parallel-cli exposes the same controls as repeatable --include-domains and --exclude-domains flags.
can parallel search include or exclude specific domains from resultsSteps
- Decide whether you need a hard allowlist or a blocklist. Use includedomains when you know exactly which sources count; use excludedomains when you only want to cut noise like forums or social sites.
Expected: you can write down the domain list and say which of the two it is.
- Send the policy inside the search request body as source_policy. Get an API key from the Parallel platform console and send it in the x-api-key request header.
Expected: the API returns a 200 with a results array, not a validation error.
- Check that the returned URLs honor the policy. Scan the url fields of the results.
Expected: with includedomains, every result URL belongs to a listed domain; with excludedomains, no result URL belongs to a blocked one.
- If results look thin with an allowlist, widen it instead of dropping the policy. Small allowlists can return zero results on niche queries.
Expected: adding one or two more domains brings results back without losing focus.
- For quick checks use the CLI flags instead of the raw API: parallel-cli search with --include-domains or --exclude-domains, repeatable per domain.
Expected: the CLI returns JSON results scoped to the domains you passed.
curl -s -X POST "https://api.parallel.ai/v1beta/search" \
-H "Content-Type: application/json" \
-H "parallel-beta: search-extract-2025-10-10" \
-d '{"objective": "React hooks best practices", "max_results": 5, "source_policy": {"include_domains": ["react.dev", "github.com"]}}'To block domains instead:
curl -s -X POST "https://api.parallel.ai/v1beta/search" \
-H "Content-Type: application/json" \
-H "parallel-beta: search-extract-2025-10-10" \
-d '{"objective": "startup funding news", "max_results": 5, "source_policy": {"exclude_domains": ["reddit.com", "x.com"]}}'Use this when
- An agent needs search results limited to trusted sources like docs sites or publishers
- You want to block noisy or low-quality domains from agent research results
- You are building a domain-scoped research flow on the Parallel Search or Task API
- A query keeps returning forums and you want them gone without rewriting the query
Not for this skill when
- You only want to prefer certain sources, not hard-filter them (steer that in the objective text, e.g. prefer official documentation)
- You are using a different search provider (this policy object is Parallel-specific)
- You need more than 200 domains combined (the API rejects it; split into multiple searches)
- You need per-URL filtering rather than per-domain filtering (filter client-side after the call)
Variant phrasings
parallel search api include domains
Pass "source_policy": {"include_domains": ["arxiv.org", "nature.com"]} in the request body. Only those domains appear.
parallel search exclude domains
Pass "source_policy": {"exclude_domains": ["reddit.com"]} in the request body. Those domains never appear.
parallel-cli domain flags
parallel-cli search "your query" --include-domains react.dev --include-domains github.com --json or the --exclude-domains mirror. Flags are repeatable, up to 10 per search on the CLI.
Why it happens
Agent search calls often drown in forum threads and SEO filler when the agent actually wants docs or primary sources. Parallel answers that with a hard domain policy evaluated server-side: the allowlist version narrows the corpus before ranking, the blocklist version removes domains after retrieval. The 200-domain combined cap keeps the policy cheap to evaluate, and apex-domain matching means you list example.com once instead of every subdomain.
Edge cases
- includedomains plus excludedomains in one request is allowed; the 200 cap counts both lists together and exceeding it raises a validation error.
- Listing an apex domain covers www and all subdomains automatically, so do not list each subdomain separately.
- Overly narrow allowlists return empty results on niche topics. That is the policy working, not a bug; widen the list.
- Domain filtering is exact on the registered domain. Internationalized or lookalike domains are treated as different domains.
- afterdate and other filters combine with sourcepolicy; they narrow independently, so stacking many filters can also empty the results.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstl-nlnF4R_eV4nKHkjRYAg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.