fraudurl › Guide
How to check a list of URLs for phishing
This guide shows how to check a whole CSV or text file of links for phishing on your own computer with fraudurl, a free, open-source phishing URL checker for Python. It takes one command, needs no account or API key, and never opens the links.
What you need
- Python 3.9 or later, on Windows, macOS or Linux. Nothing else to install.
- The tool: fraudurl_standalone.py
(one 450 KB file, models included). Or install the package with
pip install fraudurl(PyPI) and use thefraudurlcommand. - Your links, in a CSV file (any column layout) or a plain text file with one link per line.
Step 1: put the links in a file
A spreadsheet exported as CSV works as it is. fraudurl finds the column that holds the URLs by itself; if it
picks the wrong one, name it with --url-column "Website". Every other column is kept in the output.
Links without http:// or https://, such as github.com/python/cpython, are fine.
url
http://paypal.com.secure-login.test/webscr/login.php?cmd=verify
https://www.wikipedia.org/
github.com/python/cpython
Step 2: run one command
python fraudurl_standalone.py links.csv
This writes links.fraudurl.csv next to your file (choose another name with -o) and
prints a summary such as Verdicts: FRAUD=1, LEGITIMATE=2. Nothing is sent over the network.
Step 3: read the results
This is the real output for the three links above:
| url | fraud_verdict | fraud_probability | top_reasons |
|---|---|---|---|
| http://paypal.com.secure-login.test/webscr/login.php?cmd=verify | FRAUD | 0.998 | well-known brand name in the subdomain (impersonation pattern) (raises risk); uses unencrypted http:// (raises risk); 4 login/payment words in the path (raises risk) |
| https://www.wikipedia.org/ | LEGITIMATE | 0.036 | domain ending '.org' (8% of training URLs with it were phishing) (lowers risk); has a 'www.' prefix (lowers risk); subdomain is 3 characters long (lowers risk) |
| github.com/python/cpython | LEGITIMATE | 0.079 | domain ending '.com' (23% of training URLs with it were phishing) (lowers risk); uses https:// (or no scheme given) (lowers risk); 4 slashes (lowers risk) |
The columns fraudurl adds:
- fraud_verdict:
FRAUD,LEGITIMATE,REVIEW(the address alone is not conclusive) orERROR(not a usable URL; the error column says why). - fraud_probability: the calibrated probability that the link is phishing.
- verdict_confidence: how sure the verdict is, from 0.5 to 1.
- top_reasons: up to three plain-English reasons, taken from the model's own per-clue contributions.
- registrable_domain: the domain someone actually registered, such as
secure-login.testin the first row, notpaypal.com. - analysis_mode and lookup_status: which checks produced the verdict.
What to do with each verdict
- FRAUD: the address looks like phishing. If phishing is rare in your data, have a person confirm before blocking anything (see the next step).
- REVIEW: the address alone is not enough. Look at these by hand, or let the optional lookups below decide more of them.
- LEGITIMATE: nothing in the address looks like phishing. It is not a malware or scam check.
Sort the output by fraud_probability in Excel or any spreadsheet to work from the most suspicious
link down.
If phishing is rare in your list, say so
The probabilities are calibrated on data where about 47% of URLs were phishing. In everyday traffic phishing is usually far rarer, and then many FRAUD verdicts are false alarms. Tell fraudurl the share you expect:
python fraudurl_standalone.py links.csv --base-rate 0.01
The probabilities are re-weighted to that rate, and a FRAUD row whose re-weighted probability drops below 0.5
becomes REVIEW. With --base-rate 0.01, the PayPal look-alike above still comes out FRAUD, at 0.877.
Look closer at the unsure links (optional)
python fraudurl_standalone.py links.csv --enrich-review
For links the offline check leaves in REVIEW, this looks up the domain's DNS records (through Cloudflare's DNS
service) and its registration date (from the domain registry), and a second model uses them to settle the
verdict. It never visits the website. The lookups are cached per domain and politely rate-limited, so this mode
is much slower than the offline check. --enrich does the same for every link.
Add your own allow and block lists (optional)
Lists always win over the model. Put one domain, host, IP address or URL prefix per line; lines starting with
# are comments, and hosts-file lines such as 0.0.0.0 bad.example are understood. A domain
also covers its subdomains.
python fraudurl_standalone.py links.csv --allow-list partners.txt --block-list known_bad.txt
A listed row says so in top_reasons, for example on your block list (entry
'example-bad.test'), followed by what the model alone would have said. The block list wins if a link is on both.
Check a single link
python fraudurl_standalone.py --url "http://paypal.com.secure-login.test/webscr/login.php" --quiet
The result is one JSON object: the verdict, probability, reasons and the stages that produced it. Repeat
--url to check several links; add --format json to get JSON Lines for a whole file.
Use it from a Python script
Version 1.0 is a command-line tool with no import-level API, so call it with subprocess:
import json, subprocess, sys
url = "http://paypal.com.secure-login.test/webscr/login.php"
out = subprocess.run([sys.executable, "fraudurl_standalone.py", "--url", url, "--quiet"],
capture_output=True, text=True, check=True).stdout
result = json.loads(out)
print(result["fraud_verdict"], result["fraud_probability"], result["top_reasons"][0])
# FRAUD 0.998 well-known brand name in the subdomain (impersonation pattern) (raises risk)
Large lists
fraudurl streams the file, so memory stays flat. It checks about 1,000 URLs a second on one CPU core and
2,600–3,200 a second on four; a million URLs took 6 minutes 20 seconds on a 4-core desktop, using about 230 MB.
It uses up to 4 worker processes by default; change that with --workers.
fraudurl is built and tested for phishing URLs. How accurate it is, and where it is weaker, is on the home page and in the model card.