theHarvester Tutorial 2026: OSINT Email, Subdomain & Domain Recon

·

theHarvester is the tool almost every red-team engagement and OSINT investigation opens with. Point it at a domain and it queries dozens of third-party datasets — certificate transparency logs, search engines, threat-intel feeds, breach corpora — and hands you back emails, subdomains, hostnames, IPs, ASNs and employee names, without sending a single packet at the target itself. That "passive by default" posture is exactly why it is the safest first pass, and exactly why defenders should run it against their own estate before an attacker does.

This guide installs theHarvester three ways (the modern uv workflow from source, Docker Compose, and the Kali/pipx route), walks every core recon workflow with copy-pasteable commands, wires up the API keys that unlock the high-value sources, and covers the newer HarvestView web UI and REST API for automation. Then it does the part most tutorials skip: shows you how to detect this reconnaissance against your own environment with KQL and SPL.

Everything here is written for authorized use — your own domains, a client with a signed scope, or a lab. Passive does not mean permissionless.

◈ Table of Contents

01 What theHarvester Is 02 Prerequisites & Requirements 03 Method 1 — uv Install (Recommended) 04 Method 2 — Docker Compose 05 Method 3 — Kali Linux & pipx 06 First Run & Reading Output 07 Data Sources & API Keys 08 Core Recon Workflows 09 HarvestView, REST API & Automation 10 Blue-Team Use & Troubleshooting 11 Sources & References
🔎

01 — What theHarvester Is & Why Defenders Care

Phase 1 / 11

theHarvester, maintained by Christian Martorella (laramies) since 2011, is an open-source reconnaissance collector written in Python. It has one job: take a target organisation and enumerate its externally visible footprint by asking other people's databases what they already know. It is a staple of the OSINT phase in the OSSTMM and PTES methodologies and ships pre-installed in Kali and Parrot.

A Passive by default — the source tier model Concept

Recent releases classify every data source into a capability tier so you always know whether a run touches the target. Understanding these three tiers is the single most important concept for using the tool safely on an engagement.

TierBehaviourExamplesTouches target?
P0 — PassiveQueries existing third-party datasets onlycrtsh, certspotter, bing, duckduckgo, otx, urlscan, hunter, shodanNo
P1 — DNSResolves / brute-forces DNS records-r resolution, -c brute force, -n reverse DNSDNS only
P2 — ActiveMakes direct HTTP(S) contact with target hosts-t takeover checks, --screenshotYes

Default runs (a bare -d with P0 sources) send zero traffic to the target's own infrastructure — the result is stitched together entirely from Google, Bing, Certificate Transparency logs and threat-intel APIs. Keep to P0 until your scope explicitly authorises active checks.

B What it actually collects Output

A single run against a mid-size organisation typically returns hundreds of artifacts. Each is deduplicated and tagged with the source that produced it, so you can weight confidence.

  • 1Emails — harvested from search engines, PGP key servers and breach/OSINT feeds; the seed list for password spraying tests and phishing-resistance reviews.
  • 2Subdomains & hosts — the backbone of external attack-surface mapping, largely from Certificate Transparency (crt.sh, certspotter).
  • 3IP addresses & ASNs — netblocks that anchor further scanning scope.
  • 4Employee names — feed username-convention guessing (e.g. first.last@).
  • 5Open ports & banners — when Shodan/Censys keys are configured, without you scanning anything.
C Where a blue team uses it Defensive

theHarvester is not just an attacker tool. The same output is a defender's external-exposure inventory — the shadow subdomains, forgotten dev hosts and leaked mailboxes an attacker would build their target package from.

🗺️

Attack Surface Discovery

Enumerate every subdomain in Certificate Transparency to find the staging box nobody decommissioned.

📧

Phishing Exposure

See which employee mailboxes are already public before an attacker builds a target list from them.

🔗

Recon Pipelines

Feed JSON output straight into amass, httpx and nmap for a full external-assessment chain.

🛡️

Continuous Monitoring

Schedule runs and diff results to catch new exposed hosts the moment they appear.

📋

02 — Prerequisites & Requirements

Phase 2 / 11

theHarvester is cross-platform and light on resources — it is an I/O-bound HTTP client, not a scanner. The one modern requirement worth flagging up front: current releases target a recent Python and are built around the uv package manager.

RequirementMinimumRecommended
Python3.12 min3.14 rec (matches repo .python-version)
Package managerpipuv (used by the project for lockfile installs)
OSWindows 10 / macOS 12 / any modern LinuxKali or a Linux VM for the full toolchain
RAM1 GB2 GB+ (many concurrent source workers)
NetworkOutbound HTTPS (443)Unfiltered egress — proxies break some sources
AccountsNone (P0 free sources)Free API keys: Hunter, Shodan, Censys, VirusTotal

AUTHORIZATION: theHarvester is passive, but running recon against an organisation you have no permission to assess can still breach computer-misuse law and terms of service on the queried APIs. Only target domains you own or that fall inside a signed engagement scope. "It's just Google results" is not a legal defence.

A Install uv (all platforms) Prep

uv is a fast Rust-based Python package/venv manager. It creates an isolated environment from the project lockfile so you get the exact dependency set the maintainers tested — the most reliable way to avoid the dependency-hell issues older pip installs were famous for.

Linux / macOS
$ curl -LsSf https://astral.sh/uv/install.sh | sh $ uv --version
Windows (PowerShell)
PS> powershell -c "irm https://astral.sh/uv/install.ps1 | iex"

uv also manages Python itself — uv python install 3.14 pulls the exact interpreter the project pins, so you don't have to touch your system Python at all.

⚙️

03 — Method 1: uv Install from Source (Recommended)

Phase 3 / 11

This is the maintainers' recommended path and always gives you the newest sources and fixes. It clones the repository, resolves the locked dependency tree, and runs the tool inside its own environment.

1 Clone and sync Install

Three commands take you from nothing to a working install. uv sync reads pyproject.toml / uv.lock, provisions the pinned Python, creates a .venv, and installs every dependency.

$ git clone https://github.com/laramies/theHarvester.git $ cd theHarvester $ uv sync

Every subsequent invocation is prefixed with uv run, which activates the project environment for that command only. No source .venv/bin/activate needed.

2 Smoke-test the install Verify

Run a tiny passive query against a domain you control. If Certificate Transparency subdomains come back, the toolchain is healthy.

$ uv run theHarvester -d example.com -b crtsh,certspotter
Install is healthy when
  • The banner prints the version and "Starting harvesting process"
  • A list of hosts/subdomains prints under [*] Hosts found
  • No ModuleNotFoundError / SSL errors appear
3 Keeping it updated Maintenance

Source modules break constantly as upstream sites change their markup or rate limits. Update often — a subdomain source that worked last month may silently return nothing today.

$ git pull $ uv sync # re-resolve after any dependency bump
🐳

04 — Method 2: Docker Compose

Phase 4 / 11

Docker is the cleanest option for CI pipelines, shared team boxes, or when you want the REST API running as a service. The project ships a Compose definition that runs the tool as an unprivileged user and persists runs in a named volume.

1 Build the image Install

Clone the repo and build from the bundled Dockerfile. This bakes the pinned dependency set into an image so every teammate runs identical code.

$ git clone https://github.com/laramies/theHarvester.git && cd theHarvester $ docker build -t theharvester:latest .
2 One-off CLI runs in a container Usage

Mount a host directory for output and pass CLI arguments straight through. The container is disposable; only the mounted results survive.

$ docker run --rm -v $(pwd)/results:/results theharvester:latest \ -d example.com -b crtsh,otx,duckduckgo -f /results/example

Mount your API-key file into the container or the paid sources return nothing: add -v $HOME/.theHarvester:/root/.theHarvester to the docker run line, or bake keys in via an env/secret in Compose.

3 docker-compose.yml for the API + web UI Service

To run theHarvester as a persistent service — the REST API and HarvestView UI — use a Compose file. This runs as a non-root user, persists the run database in a named volume, and reads provider keys from a mounted secret file so credentials never live in the image.

docker-compose.yml
services: theharvester: build: . image: theharvester:latest container_name: theharvester user: "1000:1000" # unprivileged ports: - "127.0.0.1:5000:5000" # bind to localhost only command: ["uv","run","harvestview","--host","0.0.0.0","--port","5000"] environment: - THEHARVESTER_API_KEY_FILE=/run/secrets/op_key volumes: - th_data:/home/harvester/.local/share/theHarvester # stash.sqlite - ./api-keys.yaml:/home/harvester/.theHarvester/api-keys.yaml:ro secrets: - op_key restart: unless-stopped volumes: th_data: secrets: op_key: file: ./operator.key
$ docker compose up -d $ docker compose logs -f theharvester

Binding the port to 127.0.0.1 is deliberate — the REST API and UI have no business being reachable from the network. Front them with an SSH tunnel or reverse proxy with auth if remote access is genuinely required.

🐉

05 — Method 3: Kali Linux & pipx

Phase 5 / 11

On Kali and Parrot, theHarvester is already installed. That is the fastest way to a working tool, but the packaged version can lag the GitHub master by weeks — worth knowing when a source suddenly breaks.

1 Use the pre-installed Kali package Fastest

Confirm it's present and current. If missing, the metapackage pulls it in.

$ theHarvester -h # already on Kali $ sudo apt update && sudo apt install theharvester
2 Isolated install with pipx (any distro) Portable

If you don't want the uv workflow but still want an isolated, PATH-linked binary, install straight from the Git repository with pipx. This is ideal on a non-Kali analyst workstation.

$ sudo apt install pipx -y && pipx ensurepath $ pipx install git+https://github.com/laramies/theHarvester.git $ theHarvester -d example.com -b crtsh

Do NOT sudo pip install into system Python. It clobbers OS-managed packages and is the number-one cause of a broken theHarvester and a broken Python at the same time. Use uv, pipx or a venv — never global pip as root.

MethodBest forFreshnessEffort
uv from sourceAnalysts who want latest sourcesNewestLow
Docker ComposeAPI/UI service, CI, teamsNewest (you build)Medium
Kali aptQuick one-off on KaliCan lag weeksNone
pipx from gitIsolated binary on any distroNewestLow
🚀

06 — First Run & Reading the Output

Phase 6 / 11

The command grammar is always the same: a target with -d, one or more sources with -b, and optional flags for resolution, output and limits. Learn the flags once and every workflow below is a variation on them.

FlagPurpose
-dTarget domain or company name
-bSource(s): names (crtsh,bing) or capabilities (subdomains, emails, all)
-lResult limit per source (-l 0 removes the cap)
-fOutput filename stem — writes .json / .xml / .jsonl
-rResolve discovered hostnames to IPs (P1 / DNS)
-cDNS brute force against the domain (P1)
-nReverse DNS on resolved ranges (P1)
-tSubdomain takeover checks (P2 / active)
--screenshotScreenshot discovered hosts to a directory (P2)
1 Your first real query Run

Run three free P0 sources against a domain and cap results generously. This is the safe "hello world" of external recon.

$ theHarvester -d yourcompany.com -b crtsh,duckduckgo,otx -l 500

New to the tool? Run theHarvester -b with no value, or read the help, to print the current list of source names — they change between releases and a stale source name just errors out.

2 Understanding what prints Read

Results stream to the terminal grouped by artifact type. A typical tail looks like this (values illustrative):

[*] Target: yourcompany.com [*] Searching crtsh, duckduckgo, otx. [*] Emails found: 7 --------------------- [email protected] [email protected] [*] Hosts found: 42 --------------------- vpn.yourcompany.com dev-staging.yourcompany.com:203.0.113.24 mail.yourcompany.com:198.51.100.10
Read every result as a lead
  • Hosts with dev, staging, test, uat = forgotten pre-prod exposure
  • A hostname resolving to an unexpected IP = shadow / third-party hosting
  • Role mailboxes = phishing and spray target seeds
3 Save machine-readable output Persist

The -f flag writes structured output for later parsing. Modern builds emit JSON (grouped by type) and a JSONL stream where the first line is run metadata and each following line is a deduplicated finding with provenance. Runs are also retained in a local SQLite stash.

$ theHarvester -d yourcompany.com -b crtsh,certspotter,otx -f recon_yourcompany $ ls recon_yourcompany.* # recon_yourcompany.json recon_yourcompany.xml $ jq '.hosts[]' recon_yourcompany.json | head
Stash location (persisted history)
$ ls ~/.local/share/theHarvester/stash.sqlite
🔑

07 — Data Sources & API Keys

Phase 7 / 11

theHarvester supports roughly 60 sources. Many of the best ones — Shodan, Hunter, Censys — require a free or paid API key. Configuring keys is what separates a thin free run from a genuinely deep one.

1 Where keys live: api-keys.yaml Config

Keys are read from a YAML file in your home directory. Copy the template from the repo, then fill in the providers you have. System-wide paths (/etc/theHarvester/) also work for shared installs.

$ mkdir -p ~/.theHarvester $ cp theHarvester/data/api-keys.yaml ~/.theHarvester/api-keys.yaml $ chmod 600 ~/.theHarvester/api-keys.yaml # keys are secrets $ nano ~/.theHarvester/api-keys.yaml
api-keys.yaml (excerpt)
apikeys: shodan: key: YOUR_SHODAN_KEY hunter: key: YOUR_HUNTER_KEY censys: id: YOUR_CENSYS_ID secret: YOUR_CENSYS_SECRET virustotal: key: YOUR_VT_KEY

Never commit api-keys.yaml to a repo or bake it into a Docker image layer. Add it to .gitignore, mount it read-only at runtime, and rotate any key that has ever touched a shared box.

2 The sources worth configuring Reference

You do not need all 60. These are the high-value sources per artifact type; the free tiers are enough for most assessments.

SourceBest forKey?Free tier
crtshSubdomains (CT logs)NoUnlimited
certspotterSubdomains (CT logs)NoYes
otx (AlienVault)Hosts, passive DNSOptionalYes
hunterCorporate emailsYes~25/mo
shodanOpen ports, bannersYesLimited
censysHosts, certs, servicesYesLimited
virustotalSubdomains, passive DNSYes4 req/min
urlscanHosts, URLsNoYes
intelxEmails, leaks, docsYesTrial
3 Run a keyed, multi-source sweep Run

Once keys are set, chain the sources that complement each other — CT logs for subdomains, Hunter for emails, Shodan for exposed services — in a single command.

$ theHarvester -d yourcompany.com -b crtsh,certspotter,hunter,shodan,otx -l 1000 -f deep_recon

Use the capability keyword -b subdomains or -b all to let theHarvester pick every configured source in that category, instead of naming them one by one. all is noisy but exhaustive for a first pass on your own estate.

🧭

08 — Core Recon Workflows

Phase 8 / 11

These are the recipes you will actually run on an engagement, ordered from safest (pure passive) to most active (direct target contact). Escalate down this list only as your scope allows.

1 Email harvesting Passive

Combine search-engine and email-specialist sources. The output seeds username-convention analysis and phishing-resistance reviews.

$ theHarvester -d yourcompany.com -b bing,duckduckgo,hunter,intelx -l 500
2 Subdomain enumeration via Certificate Transparency Passive

CT logs are the single richest passive subdomain source — every TLS certificate ever issued for the domain is public record. crt.sh and certspotter mine exactly that.

$ theHarvester -d yourcompany.com -b crtsh,certspotter,virustotal,urlscan -f subs

CT never forgets. A cert issued for old-intranet.yourcompany.com in 2019 still shows up today — which is precisely how attackers find hosts you decommissioned but forgot to remove from DNS.

3 Resolve discovered hosts to IPs DNS · P1

Add -r to resolve every discovered hostname. Now you have host-to-IP mappings — the input to netblock scoping and Shodan/Censys pivots.

$ theHarvester -d yourcompany.com -b crtsh,certspotter -r -f resolved
4 DNS brute force & reverse DNS DNS · P1

-c brute-forces subdomains from a wordlist; -n runs reverse DNS across the resolved ranges to catch hosts that share a netblock. These touch DNS (not the target's web tier), so confirm scope covers active DNS queries.

$ theHarvester -d yourcompany.com -b crtsh -c -n -f bruteforce

DNS brute force generates a burst of queries against the authoritative name servers and your resolver. On a monitored network this is one of the first things a DNS-anomaly rule flags — expect it to show up in the blue-team queries in Phase 10.

5 Subdomain takeover checks Active · P2

-t checks discovered subdomains for dangling CNAMEs pointing at deprovisioned cloud services — the classic subdomain-takeover condition. This makes direct HTTP contact, so it is a P2 active check.

$ theHarvester -d yourcompany.com -b crtsh,certspotter -t -f takeover_check
6 Screenshot the attack surface Active · P2

The --screenshot flag captures every reachable host to an output directory — a fast visual triage of what's exposed, from forgotten admin panels to default install pages.

$ theHarvester -d yourcompany.com -b crtsh -r --screenshot ./shots
7 Tune workers and result limits for large targets Scale

Big estates need higher limits and more concurrency. Raise the per-source cap and the worker count, but back off if sources start rate-limiting you (HTTP 429).

$ theHarvester -d bigcorp.com -b all -l 0 --source-workers 6 -f bigcorp_full

-b all with -l 0 (uncapped) can run for a long time and will hammer rate-limited APIs. Reserve it for your own estate or an authorised long-run; for time-boxed engagements name the three or four sources that matter.

🔁

09 — HarvestView, REST API & Automation

Phase 9 / 11

Recent versions ship far more than a CLI: a local web UI (HarvestView), a REST API for programmatic runs, scheduling, and hostname-change tracking. This is what turns theHarvester from a point-in-time command into a continuous exposure monitor.

Automation flow: schedule or API-trigger runs → theHarvester queries sources → export JSONL or pipe into the wider recon chain
1 Launch HarvestView (web UI) UI

HarvestView is a local browser interface for running, scheduling and reviewing harvests, and for tracking when a target's hostnames change between runs. It binds to localhost.

$ uv run harvestview # open http://127.0.0.1:5000 — Swagger API docs at /docs
2 Drive it from the REST API API

The API authenticates with an operator key sent in the X-API-Key header. Submit a run, poll for completion, then export JSONL — the pattern for wiring theHarvester into any orchestrator.

List available sources
$ curl -s -H "X-API-Key: $OP_KEY" http://127.0.0.1:5000/api/v1/sources
Submit an enumeration run
$ curl -s -X POST http://127.0.0.1:5000/api/v1/runs \ -H "X-API-Key: $OP_KEY" -H "Content-Type: application/json" \ -d '{"target":"yourcompany.com","sources":["crtsh","certspotter","otx"]}'
Poll and export results
$ curl -s -H "X-API-Key: $OP_KEY" http://127.0.0.1:5000/api/v1/runs/$RUN_ID $ curl -s -H "X-API-Key: $OP_KEY" http://127.0.0.1:5000/api/v1/runs/$RUN_ID/export -o run.jsonl
3 Pipe into the wider recon chain Integrate

theHarvester is the front of the funnel. Extract its hostnames with jq, then feed live-host probing (httpx) and port scanning (nmap) — the standard external-assessment pipeline.

$ theHarvester -d yourcompany.com -b crtsh,certspotter -f out $ jq -r '.hosts[]' out.json | cut -d: -f1 | sort -u > hosts.txt $ cat hosts.txt | httpx -silent | tee live.txt $ nmap -iL hosts.txt -T4 --top-ports 100 -oA scan

For continuous monitoring, schedule the JSONL export nightly and diff against yesterday's host list. A brand-new subdomain appearing at 2am is exactly the signal a defender wants — often a shadow-IT deployment nobody told security about.

⏰

Scheduled Harvests

HarvestView schedules recurring runs so exposure data is never stale.

📈

Hostname Diffing

Change-tracking flags new or removed hosts between runs automatically.

🧩

Pipeline Front-End

JSONL feeds amass, httpx, nuclei and nmap for a full assessment chain.

🛡️

10 — Blue-Team Use, Maintenance & Troubleshooting

Phase 10 / 11

Passive P0 recon is invisible to you by design — you cannot detect someone querying crt.sh. What you can do is (1) monitor your own Certificate Transparency footprint to shrink what recon returns, and (2) detect the active P1/P2 stages (DNS brute force, host probing) when an attacker escalates against your estate.

1 Shrink your passive footprint Hardening

Run theHarvester against yourself on a schedule and treat the output as a to-do list. You cannot stop CT logging, but you can control what it reveals.

  • 1Decommission dangling DNS records for hosts that no longer exist — kills both stale CT hits and takeover risk.
  • 2Use wildcard certs sparingly on internal names so pre-prod hostnames don't leak into CT.
  • 3Replace public role mailboxes with contact forms where practical to cut harvested emails.
  • 4Enforce phishing-resistant MFA so a harvested-email spray has nowhere to land.
2 Detect DNS brute force (KQL + SPL) Detection

The -c brute-force stage produces a spike of failed DNS lookups (NXDOMAIN) for non-existent subdomains from a single source in a short window. Both queries below alert on that burst.

DETECTS: a single client generating an abnormal volume of NXDOMAIN responses for one domain — the signature of subdomain brute forcing.
Microsoft Sentinel — KQL (DNS logs)
DnsEvents | where TimeGenerated > ago(1h) | where ResponseCode == "NXDOMAIN" | summarize nxdomains = count(), sample = make_set(Name, 10) by ClientIP, bin(TimeGenerated, 5m) | where nxdomains > 50 | order by nxdomains desc
Splunk — SPL
index=dns (reply_code="NXDOMAIN" OR response_code="NXDOMAIN") earliest=-1h | bin _time span=5m | stats dc(query) as unique_names count as nxdomains values(query) as sample by src_ip _time | where nxdomains > 50 | sort - nxdomains

Tune the 50 threshold to your baseline — busy resolvers on large networks generate NXDOMAIN legitimately. Group by an authoritative-zone field if you have one, so you alert on brute force against your domains specifically.

3 Detect host probing & screenshotting Detection

The P2 stages (-t, --screenshot, or a downstream httpx sweep) touch many of your hosts from one source in a short window. Web/WAF logs show one client hitting a large number of distinct hostnames.

DETECTS: a single source IP requesting an unusually large set of distinct hostnames/vhosts — fan-out characteristic of automated external probing.
Microsoft Sentinel — KQL (WAF / web logs)
AzureDiagnostics | where TimeGenerated > ago(1h) | where Category == "ApplicationGatewayAccessLog" | summarize hosts = dcount(host_s), hits = count() by clientIp_s, bin(TimeGenerated, 10m) | where hosts > 20 | order by hosts desc
Splunk — SPL
index=web sourcetype=access_combined earliest=-1h | bin _time span=10m | stats dc(host) as distinct_hosts count as hits by src_ip _time | where distinct_hosts > 20 | sort - distinct_hosts
4 Troubleshooting common errors Fix

Most theHarvester problems are source-side (rate limits, markup changes, missing keys), not bugs in the tool. This table covers the ones you will hit.

SymptomCauseFix
A source returns 0 resultsSource markup/API changed or is rate-limitedgit pull && uv sync; try another source; wait out the limit
HTTP 429 / "read timed out"API rate limit hitLower --source-workers, add a paid key, or reduce -l
"No API key found for..."Key missing or wrong YAML pathCheck ~/.theHarvester/api-keys.yaml and indentation
ModuleNotFoundErrorRan outside the project envPrefix with uv run, or reinstall via pipx
SSL: CERTIFICATE_VERIFY_FAILEDCorporate TLS-inspection proxyTrust the proxy CA, or run from an unfiltered network
Unknown source name errorSource renamed/removed in this releasePrint the current source list; update the -b value

If one specific source consistently fails while others work, it is almost always an upstream change — check the project's GitHub issues before assuming your install is broken.

📚

11 — Sources & References

Phase 11 / 11

Primary documentation and reference material used in this guide. Always verify source names and flags against the current release, which changes frequently.

laramies/theHarvester — Official GitHub repository (install, sources, usage) theHarvester Wiki — REST API, source configuration, API keys Kali Linux Tools — theHarvester package documentation Astral uv — Python package & environment manager documentation crt.sh — Certificate Transparency search (subdomain source) Shodan Developer — API key and query reference

Know your external footprint before an attacker maps it for you.

Run theHarvester against your own domains this week, then paste the resolved host list into CyberHawk's IOC Scanner and browse our blog for the rest of the external-assessment toolchain — Nmap, Shodan, BloodHound and more. Turn recon on yourself before someone else does.

◈ Stay Connected

Follow CyberHawk Threat Intel for threat intelligence, deployment guides and hands-on SOC tooling content.

🌐 Website ▶️ YouTube ▶️ YouTube (2) 𝕏 Twitter / X ♪ TikTok ✈️ Telegram
🔍 IOC Scanner 🛠️ Live Tools 📚 Courses 🚨 Threat Intel 📝 Blog 📋 SOPs

"They can't exploit you if you are the Exploit."