Give a detective one name and they come back with a life story. Maigret, the MIT-licensed project by soxoj, aims for the same trick in software: its documentation opens with the line that it “collects a dossier on a person by username only, checking for accounts on a huge number of sites and gathering all the available information from web pages”, and it does so with no API keys required. The release studied here is version 0.6.6, a Python 3.10+ package whose CLI is a thin wrapper around a fully asynchronous engine built on asyncio and aiohttp.

The scale is the first thing that stands out. The bundled manifest maigret/resources/data.json carries 7,922 site entries together with 28 reusable “engines” and 75 tags, and its companion db_meta.json stamps the database with a version, an update timestamp, a minimum-compatible Maigret version, and a SHA-256 of the data. A default run does not hammer all of them: it checks the 500 highest-ranked sites by Alexa traffic, while -a widens the sweep to everything and --tags narrows it to categories or countries. Where Sherlock, covered earlier in this series, treats each site as a hand-written rule, Maigret layers a second idea on top: thousands of forums and self-hosted platforms run the same underlying software, so one XenForo or phpBB engine can supply check logic to hundreds of sites at once.

As always in this series, this is an architecture study, not an instruction manual. The project’s own README carries a blunt disclaimer, “For educational and lawful purposes only”, naming GDPR and CCPA compliance as the user’s responsibility, and it classifies its techniques in the SOWEL OSINT taxonomy for transparency. Username checks against public profile pages are a legitimate research technique, but the same power enables stalking, which is why responsible use, respect for site terms of service, and applicable privacy law frame every section below.

Maigret overview architecture diagram

Maigret's end-to-end shape: a CLI and self-updating site database feed an asyncio queue executor whose per-site checks run through a stack of transport checkers into four-state verdicts, live notifications, eight report formats, and an optional AI summary.

Reading the overview from left to right:

  • CLI entry (maigret/maigret.py) — main() and a large ArgumentParser organize flags into site filtering, operating modes, output, and report-format groups. See maigret.py.
  • Auto-update (maigret/db_updater.py) — checks a lightweight meta file once per 24 hours, downloads and SHA-256-verifies a newer database, and caches it under ~/.maigret. See db_updater.py.
  • Site database (maigret/sites.py, maigret/resources/data.json) — MaigretDatabase loads 7,922 entries and 28 engines into MaigretSite objects.
  • Async engine (maigret/checking.py, maigret/executors.py) — the maigret() coroutine fans out per-site tasks through an AsyncioQueueGeneratorExecutor with a default of 100 parallel workers.
  • Verdicts (maigret/result.py) — MaigretCheckStatus keeps four states: Claimed, Available, Unknown, and Illegal.
  • Reports (maigret/report.py, maigret/ai.py) — eight output formats plus an optional OpenAI-compatible investigation summary streamed to the terminal.

Why You Need This

The core value is the dossier. A single search does not stop at “found” or “not found”: when parsing is enabled, Maigret runs the sister project socid_extractor over every claimed profile page and pulls out whatever identifiers the page exposes, including links to other accounts. Those identifiers join a supported set of twelve types in checking.py — username, yandex_public_id, gaia_id, vk_id, ok_id, wikimapia_uid, steam_id, uidme_uguid, yelp_userid, orcid, qq_id, and bilibili_id — and each new handle can trigger another search wave. This recursive mode, disabled with --no-recursion, is what turns a list of hits into an actual profile graph.

The second reason is the engine inheritance model. Platforms like XenForo, phpBB, vBulletin, Discourse, Mastodon, MediaWiki, GitLab, Gitea, Flarum, Lemmy, bbPress, BuddyPress, NodeBB, and uCoz all expose predictable profile URL shapes and “user not found” behaviors. Maigret encodes each family once as an engine, and the 5,280 site entries that carry no explicit checkType inherit their rules from one. Of the remaining entries, 1,075 use message checks, 1,476 use status codes, and 91 use response-URL checks. When a site deliberately deviates from its engine, the code records the override and strips it back out on save, so the database stays clean.

The third reason is operational maturity. The database updates itself daily without a release: db_updater.py fetches db_meta.json from GitHub, compares timestamps, validates the minimum version, verifies the SHA-256 of the payload, and atomically swaps it into ~/.maigret/data.json with a fallback to the bundled copy when offline. A --self-check mode audits usernameClaimed and usernameUnclaimed pairs against live sites for maintainers, with --auto-disable to retire failing entries. Settings load through a layered chain — package defaults, /etc/maigret/settings.json, ~/.maigret/settings.json, then the working directory — so administrators, distributions, and users can all customize without conflicts.

How It Works

Maigret detailed architecture diagram

The full anatomy: CLI and layered settings, the database layer with engines and auto-update, six transport checkers, the asyncio queue engine with activation and retries, the verdict and extraction path, and the reporting surface including the Flask web UI.

Understanding the Architecture

A queue-based asyncio engine. The heart is AsyncioQueueGeneratorExecutor in executors.py: site checks are pushed onto an asyncio.Queue, a bounded pool of workers (default 100, tunable with -n) consumes them, and results stream out of a results queue as a generator. The implementation is unusually careful — it uses asyncio.wait() instead of wait_for() so that a site holding its connection open during bot-protection cleanup cannot stall the whole scan past the timeout, and a done-callback logs late task exceptions that nobody retrieves. Per-site work lives in check_site_for_username(), coordinated by the maigret() coroutine with a timeout set to the user’s value plus half a second.

Six transport checkers behind one interface. Every request flows through a CheckerBase subclass chosen per site. SimpleAiohttpChecker handles the clearnet with a choice of async aiodns/c-ares resolution or a threaded resolver that respects system DNS (a workaround for documented Windows and VPN failures). ProxiedAiohttpChecker routes Tor and I2P traffic through user-supplied gateways. AiodnsDomainResolver supports the experimental --with-domains mode that checks whether domains registered to a username exist. CurlCffiChecker impersonates browser TLS fingerprints for sites that block Python clients, CloudflareWebgateChecker offloads JavaScript-challenge sites to a local FlareSolverr instance when --cloudflare-bypass is enabled, and CheckerMock silently skips a site when its protocol’s gateway was not configured.

Honest verdicts in four states. The classifier applies the site’s check type: for message sites the body must contain a presence string and none of the absence strings; for status_code sites a 2xx means claimed; for response_url sites a 2xx plus presence wins with redirects disabled. Anything that smells like a block page — the COMMON_ERRORS table in errors.py fingerprints Cloudflare’s “Attention Required!” and “Just a moment” pages — becomes Unknown with an explanatory error, never a false “Available”. There is even a dedicated guard for non-ASCII usernames: if a page does not contain the searched Chinese or Arabic handle at all, the result is downgraded rather than celebrated, closing a false-positive hole documented in the code.

Token activation and retries. Some sites gate profile checks behind session credentials. activation.py implements per-site activators: fetching a Twitter guest token, minting a Vimeo JWT, solving Wikimapia’s per-IP cookie challenge, signing OnlyFans requests with a computed checksum plus an x-bc header from secrets.token_hex(20), and starting a Proton anonymous session. When a response matches a site’s activation marks, the activator runs once and the request is retried with merged headers; minted tokens are cached per user across runs. On flaky networks, --retries restarts temporarily failed requests, and a mirror fallback swaps in an alternate site URL.

Extraction, recursion, and permutation. On every Claimed result with parsing enabled, the response body goes to socid_extractor, discovered usernames and IDs feed parse_usernames(), and the CLI queues the next search wave. --parse URL starts the whole process from a single profile page, --enrich additionally fetches secondary API endpoints derived from profile URLs via URL mutations, and --permute generates variants from two or more inputs using the separators empty string, underscore, hyphen, and dot — turning “john doe” into johndoe, j.doe, doe_john, and their neighbors.

Reports, web UI, and AI summary. Results land in report.py as eight formats: TXT, CSV, HTML, PDF (an optional maigret[pdf] extra), Markdown, an XMind XML mindmap with a manifest for modern readers, an interactive D3 graph, and a Neo4j Cypher script whose re-imports are idempotent. JSON output comes in simple and ndjson flavors. The built-in Flask app, launched with --web 5000, renders results as a graph and offers every report as a download, and the published soxoj/maigret:web Docker image deploys to Render’s free tier in one click. Finally, --ai compiles the findings into an internal Markdown report and streams a short, neutral investigation summary from any OpenAI-compatible chat endpoint, with --ai-model defaulting to gpt-5.4 per the CLI help.

End to end. A run flows like this: settings resolve through the layered chain, the database auto-updater guarantees a fresh manifest, site filters pick the ranked top 500 or a tag-scoped subset, MaigretDatabase merges engine rules into per-site objects, the executor queues one check per site, the matching transport checker issues the request with a random User-Agent and Connection: close, activation and retries smooth over hostile sites, the verdict classifier and error fingerprints produce a four-state result, extraction and recursion expand the search surface, and the notifier, report writers, web UI, and optional AI summary present the dossier.

Advantages

  • Engine leverage. 28 engines supply check logic to 5,280 entries, so one well-maintained rule covers entire software families instead of single sites.
  • Self-maintaining database. Daily meta-checked updates with SHA-256 verification, version gating, atomic cache swaps, and a bundled offline fallback.
  • Transport diversity. Clearnet, Tor, I2P, DNS-only domain checks, TLS-impersonating curl_cffi, and an opt-in Cloudflare webgate cover almost every hosting reality.
  • Honest uncertainty. Four verdict states, a fingerprint table for challenge pages, and the non-ASCII presence guard actively suppress false positives.
  • Recursive intelligence. socid_extractor integration over twelve identifier types turns one handle into a growing graph of linked accounts.
  • Consumer-grade outputs. Eight report formats, a Flask results UI, Docker images for CLI and web, and an embeddable async library interface.

Benefits

  • Self-discovery. Map your own username footprint across thousands of platforms, including niche forums no commercial service bothers to index.
  • Investigation-ready reporting. XMind mindmaps, Neo4j graph imports, and PDF dossiers feed directly into professional casework and incident reviews.
  • A masterclass in async engineering. The queue executor’s cancellation handling and the checker abstraction are instructive far beyond OSINT.
  • Community quality control. --self-check with --auto-disable, database statistics via --stats, and a documented contribution path keep 7,922 entries honest.
  • Privacy-conscious operation. Proxy, Tor, and I2P support let researchers match the traffic posture to the sensitivity of the task.
  • Extensible by design. Adding a site is a JSON edit with engine inheritance; adding a platform family benefits hundreds of entries at once.

Usage

Install from PyPI and run a first search:

pip install maigret
maigret YOUR_USERNAME

Produce a full report set for one user:

maigret user --html
maigret user --pdf
maigret user --xmind
maigret user --json ndjson
maigret user --graph
maigret user --neo4j

Narrow or widen the sweep:

maigret user --tags photo,dating
maigret user --keywords python rust
maigret user1 user2 user3 -a

Route through anonymity networks:

maigret user --proxy socks5://127.0.0.1:1080
maigret user --tor-proxy socks5://127.0.0.1:9050
maigret user --i2p-proxy http://127.0.0.1:4444

Kick off a recursive search from a profile URL, or start the web interface:

maigret --parse https://example.com/profile/username
maigret --web 5000

Snap and Windows standalone builds are also published: sudo snap install maigret on Linux, and maigret_standalone.exe from the GitHub releases on Windows, with the Docker images soxoj/maigret:latest for CLI and soxoj/maigret:web for the interface.

Conclusion

Maigret shows what happens when a username checker grows up: rule inheritance through engines, a database that maintains and verifies itself, six transports behind one interface, verdicts that refuse to guess, and a reporting surface that treats a sweep as an investigation rather than a list. It is also candid about its limits and its ethics, pairing a permissive MIT license with an explicit lawful-use disclaimer and a published technique taxonomy. For anyone studying how large-scale async HTTP tooling is engineered with care, the repository repays a long read.

Links:

Watch PyShine on YouTube

Contents