## TL;DR

Yes. Parallel's Search API (and the Task API) accepts a source_policy object with include_domains and exclude_domains lists: include_domains is an allowlist, only sources from those domains appear; exclude_domains is a blocklist, sources from those domains never appear. The two lists combined cannot exceed 200 domains, and an apex domain like example.com automatically covers its subdomains. The parallel-cli exposes the same controls as repeatable --include-domains and --exclude-domains flags.

```text
can parallel search include or exclude specific domains from results
```

## Steps

1. Decide whether you need a hard allowlist or a blocklist. Use include_domains when you know exactly which sources count; use exclude_domains when you only want to cut noise like forums or social sites.
   Expected: you can write down the domain list and say which of the two it is.

2. Send the policy inside the search request body as source_policy. Get an API key from the Parallel platform console and send it in the x-api-key request header.
   Expected: the API returns a 200 with a results array, not a validation error.

3. Check that the returned URLs honor the policy. Scan the url fields of the results.
   Expected: with include_domains, every result URL belongs to a listed domain; with exclude_domains, no result URL belongs to a blocked one.

4. If results look thin with an allowlist, widen it instead of dropping the policy. Small allowlists can return zero results on niche queries.
   Expected: adding one or two more domains brings results back without losing focus.

5. For quick checks use the CLI flags instead of the raw API: parallel-cli search with --include-domains or --exclude-domains, repeatable per domain.
   Expected: the CLI returns JSON results scoped to the domains you passed.

```bash
curl -s -X POST "https://api.parallel.ai/v1beta/search" \
  -H "Content-Type: application/json" \
  -H "parallel-beta: search-extract-2025-10-10" \
  -d '{"objective": "React hooks best practices", "max_results": 5, "source_policy": {"include_domains": ["react.dev", "github.com"]}}'
```

To block domains instead:

```bash
curl -s -X POST "https://api.parallel.ai/v1beta/search" \
  -H "Content-Type: application/json" \
  -H "parallel-beta: search-extract-2025-10-10" \
  -d '{"objective": "startup funding news", "max_results": 5, "source_policy": {"exclude_domains": ["reddit.com", "x.com"]}}'
```

## Use this when

- An agent needs search results limited to trusted sources like docs sites or publishers
- You want to block noisy or low-quality domains from agent research results
- You are building a domain-scoped research flow on the Parallel Search or Task API
- A query keeps returning forums and you want them gone without rewriting the query

## Not for this skill when

- You only want to prefer certain sources, not hard-filter them (steer that in the objective text, e.g. prefer official documentation)
- You are using a different search provider (this policy object is Parallel-specific)
- You need more than 200 domains combined (the API rejects it; split into multiple searches)
- You need per-URL filtering rather than per-domain filtering (filter client-side after the call)

## Variant phrasings

### parallel search api include domains
Pass `"source_policy": {"include_domains": ["arxiv.org", "nature.com"]}` in the request body. Only those domains appear.

### parallel search exclude domains
Pass `"source_policy": {"exclude_domains": ["reddit.com"]}` in the request body. Those domains never appear.

### parallel-cli domain flags
`parallel-cli search "your query" --include-domains react.dev --include-domains github.com --json` or the `--exclude-domains` mirror. Flags are repeatable, up to 10 per search on the CLI.

## Why it happens

Agent search calls often drown in forum threads and SEO filler when the agent actually wants docs or primary sources. Parallel answers that with a hard domain policy evaluated server-side: the allowlist version narrows the corpus before ranking, the blocklist version removes domains after retrieval. The 200-domain combined cap keeps the policy cheap to evaluate, and apex-domain matching means you list example.com once instead of every subdomain.

## Edge cases

- include_domains plus exclude_domains in one request is allowed; the 200 cap counts both lists together and exceeding it raises a validation error.
- Listing an apex domain covers www and all subdomains automatically, so do not list each subdomain separately.
- Overly narrow allowlists return empty results on niche topics. That is the policy working, not a bug; widen the list.
- Domain filtering is exact on the registered domain. Internationalized or lookalike domains are treated as different domains.
- after_date and other filters combine with source_policy; they narrow independently, so stacking many filters can also empty the results.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_l-nl_nF4R_eV4nKHkjRYAg
