What this covers and who it is for
OSINT tools for penetration testing are the collection and correlation utilities that turn public, lawfully available information — DNS and WHOIS records, certificate transparency logs, search engine indexes, code repositories, document metadata, breach corpora, and social media — into a target profile you can scope and prioritize before sending a single packet. They exploit nothing. What they produce is the map: which domains are real, which services are reachable, which email addresses are valid, which credentials have already leaked, and which of those findings deserve a test window.
This hub is for three audiences: penetration testers and red teamers who want a repeatable reconnaissance pipeline rather than a folder of half-remembered one-off tools; SOC analysts and defenders who want to see their own organization the way an outsider sees it; and IT owners at small and mid-sized businesses who need to check their external footprint without an enterprise threat-intelligence platform.
Scope and authorization come before collection. Nothing here should be aimed at infrastructure you do not own or have written authorization to test — passive tools included. Reading public data is not automatically harmless: scraping in breach of a provider’s terms of service, profiling employees, or holding credential dumps can create contractual, civil, and even criminal exposure. Before the first query, get a signed authorization document that names in-scope domains and IP ranges, excludes affiliates and third-party hosting providers, sets rate limits, states whether staff personal data may be gathered, and names the person to call if you find live credentials. In the US, unauthorized access is addressed by the Computer Fraud and Abuse Act; if a target or data subject is in the EU or UK, data-protection law follows your notes too.
The short version
Take one idea from this page: OSINT is a pipeline, not a product. Collection, enrichment, correlation, verification, reporting. Tooling handles the first three imperfectly and the last two not at all — those are your job, and they are what the client pays for.
- Subdomain and DNS discovery: Amass in passive mode, Subfinder, dnsx, and certificate transparency searches against crt.sh.
- Infrastructure and service search: Shodan, Censys, Netlas, FOFA, and urlscan.io for rendered page and request history.
- Search engine and historical recon: Google and Bing dorks, the Wayback Machine CDX API, archived pages.
- People and identity: theHarvester, email-pattern discovery services, LinkedIn and job-posting analysis, username enumeration with Sherlock.
- Code and secret leakage: GitHub and GitLab code search, trufflehog, gitleaks.
- Documents and metadata: ExifTool for indexed downloadable files, Metagoofil for legacy share-heavy environments.
- Breach context: Have I Been Pwned Domain Search (after domain verification) and licensed breach-intelligence platforms.
- Orchestration and correlation: Recon-ng, SpiderFoot, Maltego.
Two of those orchestrators have deep-dive guides on this site and form the natural backbone of a repeatable workflow: Recon-ng: The Ultimate Guide to Footprinting and Web Reconnaissance and SpiderFoot: An Open-Source OSINT Automation Tool for Cybersecurity.
Key concepts you need first
Passive, active, and the legal line between them
Passive collection asks third parties what they already know: certificate transparency logs, WHOIS registries, search indexes, historical archives, and passive DNS feeds all sit on that side of the line because your traffic never reaches the target. Active collection makes the target, or something it controls, respond. DNS brute-forcing, HTTP probing, screenshotting live pages, port scanning, and scanners such as Nuclei are all active, even when they only read a banner. Practical test: if the target’s logs can show your IP address, treat the step as active and confirm it is inside the rules of engagement.
The line matters twice over. Legally, passive work usually falls under a generic OSINT clause in the engagement letter, while active work needs explicit authorization and often a rate limit. Operationally, active recon is visible to the client’s defenders — fine when the engagement should be seen, unhelpful when it should not.
Evidence quality beats volume. Tool output is a lead, not a finding. A hostname found in a certificate log may be expired, parked, shared with another tenant, or owned by a marketing agency. Record the source URL, query, UTC timestamp, and tool version, and keep raw output in the case folder; when someone asks how you know a host is theirs, you want a citation, not terminal scrollback. Time matters too: certificate logs and web archives show history, while service-search engines show a point in time. Confirm leads against live evidence before you write them up.
How to apply this
Work in order. Each stage feeds the next, and each stage should produce a file you can hand to somebody else.
- Freeze scope and build the asset list. Turn the authorization document into a working list of root domains, IP ranges, product names, and brands. Anything not on it stays out of the collection plan.
- Discover the domain surface passively. Certificate transparency first, then passive subdomain enumeration, then resolve results so you track live hosts rather than strings.
- Enrich the infrastructure and documents. Search hostnames and netblocks in service-search engines, review historical URLs, and note which hosts look forgotten. Old PDFs, slide decks, and spreadsheets leak internal hostnames, software versions, and usernames in metadata; ExifTool covers most of it.
- Map people carefully. Establish the corporate email pattern, identify technical contacts, and note public job postings that describe the technology stack. Collect the minimum needed to test, not a staff dossier.
- Check for already-leaked secrets. Search public repositories for keys, tokens, and configuration files, and check whether the client’s domains appear in known breach corpora. Never authenticate with a leaked credential — report it and let the client rotate.
- Correlate, verify, rank. Merge everything into one dataset, de-duplicate, and confirm the high-value items. Rank by what an attacker would reach first, not by what is easiest to find.
- Report each finding as a fix. Exposure, evidence, business impact, and the fix. A hostname without a remediation is trivia.
Orchestration keeps that sequence from becoming eight disconnected spreadsheets. Recon-ng gives you a database, a marketplace of modules, and a repeatable run:
recon-ng
marketplace install recon/domains-hosts/hackertarget
modules load recon/domains-hosts/hackertarget
options set SOURCE northwind-example.com
run
SpiderFoot is faster when you want broad automatic coverage and a correlation graph:
spiderfoot -s northwind-example.com -m sfp_dnsresolve,sfp_crt,sfp_hackertarget
A worked example: one day on a mid-sized manufacturer
You have written authorization to test Northwind Fabrication, a manufacturer with one primary domain and two regional sales domains, with passive collection permitted for the first 24 hours. Start with certificate transparency, usually the quickest honest view of a company’s real subdomains:
curl -s "https://crt.sh/?q=%25.northwind-example.com&output=json" | jq -r '.[].name_value' | sort -u
Then enumerate passively and resolve the results so the list shows live hosts:
subfinder -d northwind-example.com -all -silent | dnsx -silent -a -resp
Next, pull historical URLs and sample them rather than dumping thousands of lines into a report. Look for patterns: an old content-management path, a staging hostname, a password-reset endpoint, a document with a sequential identifier parameter.
curl -s "http://web.archive.org/cdx/search/cdx?url=northwind-example.com*&output=json&collapse=urlkey&fl=original" | jq -r '.[1:][] | .[0]' | sort -u | head -200
Finish with document metadata (ExifTool over anything already indexed), repository secrets (trufflehog against the organization’s public repos), and email authentication records for each domain.
A realistic first day yields a short list: four subdomains missing from the client’s asset list, one of them a reachable staging application returning verbose error output; a former contractor’s public repository holding a configuration file with a long-lived API key that still works; and a mail domain publishing an authentication policy in monitor-only mode, so spoofed mail is unlikely to be blocked. None of that required touching the client’s network, and all of it becomes a fix: decommission or authenticate the staging host, revoke the key and add secret scanning to the pipeline, and move the email policy from monitoring to enforcement.
The defender’s view of the same pipeline
Attackers run this pipeline too, which is why defenders should run it first. Certificate transparency monitoring flags new certificates for names resembling yours, catching forgotten internal hosts and lookalike domains. A recurring passive pass against your own domains finds dangling DNS records pointing at deprovisioned cloud resources — the setup for a subdomain takeover. Secret scanning plus a rotation policy closes the leak most likely to turn a public repo into initial access. Publishing and enforcing email authentication makes the email addresses you cannot hide less useful for phishing.
Be honest about detection limits. Passive collection may leave nothing in your logs, so you cannot alert your way out of being mapped. The only reliable control is removing what is findable: retire the host, revoke the key, delete the metadata, narrow the exposure. Pair that with log review for the active phase — spikes in DNS failures, 404s for common paths, scans across your address space — and treat those as a scope conversation, not a confirmed intrusion.
Where people get this wrong
- Treating “public” as “unrestricted.” A search index or social network is public, but bulk collection may violate its terms, and employee personal data is still regulated. Check the platform’s rules, not just the client’s local law.
- Forgetting the third parties. Regional sales domains, marketing microsites, SaaS portals, franchise sites, and employee-owned work accounts are frequently the easiest way in and the last thing on the asset list.
- Mistaking collection for conclusion. Volume feels like progress. Ten thousand lines of subdomain output with four verified findings is worse than forty lines with the same four, because the client will read the second one.
- Using leaked credentials to log in. Even with stated authorization, authenticating with a leaked password turns a documented exposure into unauthorized access, and it is out of scope in most engagements. Report it and let the client rotate.
- Ignoring freshness. Presenting a two-year-old service listing as current exposure destroys trust in the report. Timestamp everything and re-check before writing it up.
- Collecting people data you do not need. If the engagement is about an external attack surface, home addresses, personal phone numbers, and private social media add legal risk and no test value.
- Assuming passive means invisible. Service-search platforms log your queries, API keys identify you, and some sources rate-limit accounts. Plan so one blocked source does not stall the engagement.
- Handing over a tool dump. Clients cannot action a list of hostnames. They can action a prioritized list where every item carries evidence and a named fix.
If you would not be comfortable explaining a collection step to the client’s legal team and to the person whose data you gathered, do not run that step.
Where to go next
Build the pipeline in this order. Start with one orchestrator so findings land in a database rather than a text file, then add the collection habits: certificate transparency and passive subdomain enumeration first, infrastructure search second, documents and code third, people last and only as far as the engagement requires.
The two cluster guides cover the orchestration layer in depth. Read the Recon-ng footprinting and web reconnaissance guide for a scriptable, database-backed workflow; read the SpiderFoot open-source OSINT automation walkthrough for broad coverage with a visual correlation graph. Once recon produces a list you intend to test, move to scripting your own checks — the site’s guide to port scanning with Python, Bash, and PowerShell is the sensible next stage, where the work becomes active and must be firmly in scope.
Then do the part most testers skip: run the passive pipeline against your own primary domain on a schedule, and log each finding with its source and date. Add certificate transparency alerts and DNS change monitoring. Move every email authentication policy you own from monitoring to enforcement. Add secret scanning to the build pipeline and a rotation procedure for anything that was ever committed. Reconnaissance earns its keep only when it changes a configuration, and the fastest way to prove these tools’ value to a client is to have already fixed what they find.
Image credit: kewl, licensed under CC BY 2.0.