Pick almost any online service and you will find the same pattern: the username is the identity handle. One person tends to reuse the same handle on forums, gaming platforms, developer sites, and social networks, which means a single string like user123 can act as a thread that ties dozens of accounts together. Sherlock, an MIT-licensed project created by Siddharth Dushantha and maintained by the Sherlock Project, turns that observation into a command line tool. Its own documentation describes the purpose in one line: hunt down social media accounts by username across social networks. The bundled release this article examines is version 0.16.0, and the project’s Python package lives in a compact, well-factored module named sherlock_project.
What makes Sherlock interesting from an engineering standpoint is not the idea but the execution. The tool ships a manifest of 481 site entries in sherlock_project/resources/data.json, and every one of those entries behaves differently. Some sites return a clean 404 for a missing profile, some redirect to a search page, some answer with a text message that only a regex can interpret, and some sit behind a web application firewall that loves to hand out 403 pages to scripts. Sherlock models all of that variation as data rather than as 481 special cases in code, then drives the whole sweep through a concurrent HTTP session with twenty workers.
A note before we open the hood: Sherlock is published for research, security auditing, and personal account discovery, and this article reads it purely as a software architecture study. Checking whether a username is taken on public profile pages is a legitimate operation, but the same capability can be abused for stalking or harassment, which is why the project itself frames the tool around responsible use. If you experiment with it, respect the terms of service of the sites involved, follow the privacy laws that apply to you, and keep your probing to accounts you own or engagements you are authorized to run.
Sherlock's end-to-end shape: a manifest-driven command line layer feeds a concurrent query engine whose per-site verdicts flow into a notifier and three export writers.
Reading the overview from left to right:
- CLI entry (
sherlock_project/sherlock.py) —main()parses the argument flags, validates the timeout, checks for a newer release, and expands special usernames before anything is requested. See sherlock.py. - Manifest loader (
sherlock_project/sites.py) —SitesInformationturns raw JSON entries intoSiteInformationobjects, one per site, and honors an exclusions feed. See sites.py. - Site data (
sherlock_project/resources/data.json) — 481 entries describing URL formats, detection rules, and probes, validated againstdata.schema.json. See data.json. - Query engine (
sherlock_project/sherlock.py) — thesherlock()function builds aSherlockFuturesSessionwith a maximum of twenty workers and fans out one interpolated request per site. - Verdict model (
sherlock_project/result.py) — a five-stateQueryStatusenum and aQueryResultobject carry the outcome of every probe. See result.py. - Notifier (
sherlock_project/notify.py) —QueryNotifyPrintrenders live, colorized results as they arrive. See notify.py.
Why You Need This
The obvious use case is self-audit. If you have reused one handle for a decade, running it through Sherlock shows you exactly where that handle resolves to a live profile today. That is valuable when you are tightening your own privacy posture, pruning dormant accounts, or deciding whether a new project name is safe to adopt everywhere. Security teams use the same capability to check whether an organization’s brand handles have been squatted on lesser-known platforms.
The second reason is reliability. Hand-rolling a username checker sounds trivial until you meet the reality of 481 sites with 481 personalities. Sherlock’s data file records that, as of this release, 327 entries are judged by HTTP status code, 127 by a message in the response body, and 27 by where the response URL ends up. Fifty-two entries define a dedicated urlProbe that differs from the public profile URL, and 95 entries carry a regexCheck that constrains which usernames are even legal on that platform. Encoding that knowledge once, in one reviewed file, beats scattering guesswork across ad-hoc scripts.
The third reason is engineering quality. Sherlock is not a weekend scraper. It ships a JSON schema for the manifest, a dedicated exclusions feed for false positives, a regression test suite under tests/, and GitHub workflows named validate_modified_targets.yml, exclusions.yml, regression.yml, and update-site-list.yml that keep the community’s site contributions honest. Every site entry carries a username_claimed value used to verify that a rule actually detects a known account, and the code generates an unclaimed control name with secrets.token_urlsafe(32) for the opposite direction. That discipline is why the tool keeps working as platforms change.
How It Works
The full anatomy: CLI parsing and validation, manifest loading with NSFW and exclusions filtering, the twenty-worker query engine with interpolation and redirect policy, three verdict paths guarded by a WAF fingerprint filter, and three export formats.
Understanding the Architecture
Manifest-driven targeting. Everything starts with data. By default SitesInformation pulls the live manifest from the project’s feed at https://data.sherlockproject.xyz, so a fresh install sees curated updates without a code change; passing --local falls back to the data.json bundled with the package. Each entry becomes a SiteInformation object holding the profile URL format, an optional probe URL, the detection method, error codes or messages, redirect policy, and regex constraints. Unless --ignore-exclusions is passed, entries from the project’s false-positive exclusions feed are dropped, and remove_nsfw_sites() filters the 19 entries flagged isNSFW unless --nsfw is given.
CLI parsing and safety rails. main() wires up a long flag list: --version, --verbose, --folderoutput, --output, --csv, --xlsx, --site for pruning to a subset, --proxy, --json for loading results or pulling a manifest straight from a GitHub pull request, --timeout with a default of 60 seconds and a dedicated timeout_check() validator, --print-all, --print-found (on by default), --browse, --local, --nsfw, --txt, and --ignore-exclusions. A SIGINT handler lets you interrupt a long sweep cleanly. Before any network traffic, the tool also checks the project’s GitHub release API and tells you when a newer version exists.
Username expansion. A neat trick lives in check_for_parameter(): if the query contains the marker {?}, Sherlock builds variants by substituting the separator characters underscore, hyphen, and dot into that slot, collecting the results in multiple_usernames, and then runs the full sweep once per variant. interpolate_string() then performs the actual {} substitution of the username into each site’s URL template, so a site whose profile path is https://example.com/user/{} gets exactly the request the site expects.
A twenty-worker concurrent session. The engine class SherlockFuturesSession extends FuturesSession with max_workers=20, so the sweep of 481 targets runs as a bounded thread pool instead of a slow serial loop. A response hook stamps each response with a monotonic response_time, which later lands in the CSV and XLSX exports. A forged browser User-Agent header, Mozilla/5.0 (X11; Linux x86_64; rv:129.0) Gecko/20100101 Firefox/129.0, keeps naive bot blocking from skewing results, and --proxy routes everything through a proxy when needed.
Three detection methods. The heart of the tool is get_response() plus the verdict logic. For status_code sites, the request defaults to HEAD, an errorCode list marks responses that count as “not found”, and anything at or above 300 or below 200 means the username is available. For message sites, the code searches the body for one or more errorMsg strings. For response_url sites, the request is sent with redirects disabled and a 2xx means claimed. Before any request is made, a regexCheck guard can short-circuit the probe with an ILLEGAL verdict if the username cannot exist on that platform at all.
WAF awareness. A standout detail is the WAFHitMsgs fingerprint table. When a response body matches known challenge pages from Cloudflare, Cloudfront, or PerimeterX, Sherlock does not guess; it reports QueryStatus.WAF instead of a false “available”. That distinction between “no account” and “we cannot see” is what separates an honest OSINT tool from a false-positive generator.
Verdicts, notifier, and exports. Every probe resolves into a QueryResult carrying the username, site name, profile URL, one of the five QueryStatus values (CLAIMED, AVAILABLE, UNKNOWN, ILLEGAL, WAF), the query time, and an optional context string. QueryNotifyPrint prints results live with colorama-based coloring, and --browse opens each found profile in the system browser via webbrowser. At the end it prints the result count. Exports are handled in sherlock.py: --txt writes one claimed URL per line plus a “Total Websites Username Detected On” count; --csv writes rows with the columns username, name, url_main, url_user, exists, http_status, and response_time_s; --xlsx builds a pandas DataFrame with clickable =HYPERLINK() formulas and writes an Excel workbook.
End to end. A run flows like this: flags are parsed and validated, the manifest is fetched or loaded locally and filtered by exclusions, NSFW flag, and any --site subset, the username is expanded if it contains {?}, then sherlock() interpolates the name into 481 URL templates, fires them through the twenty-worker session with the right verb and redirect policy, classifies each response through the WAF filter and the matching detection method, and streams QueryResult objects into the live notifier and whichever export files you asked for.
Advantages
- Breadth without code sprawl. 481 sites are described as data entries, so adding or fixing a target is a JSON edit reviewed by CI, not a code change.
- Honest verdicts. Five states, including
UNKNOWNandWAF, refuse to disguise uncertainty as a clean negative, and the WAF fingerprint table catches challenge pages from major CDN providers. - Real concurrency. The
FuturesSessionwith a twenty-worker ceiling makes a full sweep practical while keeping load on targets reasonable, and per-response latency is measured by a hook. - False-positive hygiene. A curated exclusions feed, claimed-versus-unclaimed control usernames, and 95
regexCheckguards show a systematic attack on the classic OSINT failure mode of wrong answers. - Operational flexibility. Proxy support, configurable timeout,
--localoffline manifests,--sitesubsetting, and JSON manifests pulled from GitHub pull requests make the tool scriptable in many environments. - Multiple output formats. Plain text, CSV, and Excel with working hyperlinks cover humans, spreadsheets, and downstream automation alike.
Benefits
- Self-discovery and privacy auditing. Map where your own handle lives across the web in minutes, then close the accounts you forgot you had.
- Brand and impersonation checks. Verify that your project or company name has not been claimed by impostors on niche platforms you never monitor.
- A teaching codebase. The separation between manifest schema, engine, verdict model, and notifier is a compact masterclass in turning messy real-world variance into clean data-driven design.
- Trustworthy foundations. MIT license, an active test suite, validation workflows for community site contributions, and packaging via pipx, Fedora’s dnf, Docker, and major distro repositories.
- Extensible by non-programmers. Because targeting logic lives in
data.jsonunder a published schema, contributors can add sites without reading the engine code. - Predictable behavior. Deterministic verdict rules, measured response times, and explicit handling of redirects, HEAD versus GET, and status thresholds make runs reproducible and diffable.
Usage
Install the released package with pipx:
pipx install sherlock-project
On Fedora-based systems it is available from the distribution repositories:
dnf install sherlock-project
Run a search for a single username; results are written to user123.txt:
sherlock user123
Several usernames can be swept in one invocation:
sherlock user1 user2 user3
The project also ships a Dockerfile, so no local Python setup is required:
docker run -it --rm sherlock/sherlock
Useful flags for controlled runs include --csv and --xlsx for tabular exports, --site to restrict the sweep to chosen targets, --timeout to bound each request, --proxy for routing, --browse to open found profiles, --local to use the bundled manifest instead of the live feed, and --nsfw to include sites excluded by default.
Conclusion
Sherlock earns its reputation the same way good infrastructure software always does: by refusing to lie. It treats “not found”, “unknown”, “illegal name”, and “blocked by a firewall” as different answers, it encodes the quirks of 481 sites in a reviewed, schema-validated data file, and it does the waiting with a bounded thread pool instead of a naive loop. For anyone curious how a large-scale HTTP reconnaissance tool is engineered responsibly, the repository is a rewarding read.
Links:
- Project site: https://sherlockproject.xyz/
- Repository: https://github.com/sherlock-project/sherlock
- Manifest data file: data.json
- Engine source: sherlock.py
Enjoyed this post? Never miss out on future posts by following us