fraudurl

fraudurl › Guide

How to check a list of URLs for phishing

This guide shows how to check a whole CSV or text file of links for phishing on your own computer with fraudurl, a free, open-source phishing URL checker for Python. It takes one command, needs no account or API key, and never opens the links.

What you need

Step 1: put the links in a file

A spreadsheet exported as CSV works as it is. fraudurl finds the column that holds the URLs by itself; if it picks the wrong one, name it with --url-column "Website". Every other column is kept in the output. Links without http:// or https://, such as github.com/python/cpython, are fine.

url
http://paypal.com.secure-login.test/webscr/login.php?cmd=verify
https://www.wikipedia.org/
github.com/python/cpython

Step 2: run one command

python fraudurl_standalone.py links.csv

This writes links.fraudurl.csv next to your file (choose another name with -o) and prints a summary such as Verdicts: FRAUD=1, LEGITIMATE=2. Nothing is sent over the network.

Step 3: read the results

This is the real output for the three links above:

urlfraud_verdictfraud_probabilitytop_reasons
http://paypal.com.secure-login.test/webscr/login.php?cmd=verifyFRAUD0.998 well-known brand name in the subdomain (impersonation pattern) (raises risk); uses unencrypted http:// (raises risk); 4 login/payment words in the path (raises risk)
https://www.wikipedia.org/LEGITIMATE0.036 domain ending '.org' (8% of training URLs with it were phishing) (lowers risk); has a 'www.' prefix (lowers risk); subdomain is 3 characters long (lowers risk)
github.com/python/cpythonLEGITIMATE0.079 domain ending '.com' (23% of training URLs with it were phishing) (lowers risk); uses https:// (or no scheme given) (lowers risk); 4 slashes (lowers risk)

The columns fraudurl adds:

What to do with each verdict

Sort the output by fraud_probability in Excel or any spreadsheet to work from the most suspicious link down.

If phishing is rare in your list, say so

The probabilities are calibrated on data where about 47% of URLs were phishing. In everyday traffic phishing is usually far rarer, and then many FRAUD verdicts are false alarms. Tell fraudurl the share you expect:

python fraudurl_standalone.py links.csv --base-rate 0.01

The probabilities are re-weighted to that rate, and a FRAUD row whose re-weighted probability drops below 0.5 becomes REVIEW. With --base-rate 0.01, the PayPal look-alike above still comes out FRAUD, at 0.877.

Look closer at the unsure links (optional)

python fraudurl_standalone.py links.csv --enrich-review

For links the offline check leaves in REVIEW, this looks up the domain's DNS records (through Cloudflare's DNS service) and its registration date (from the domain registry), and a second model uses them to settle the verdict. It never visits the website. The lookups are cached per domain and politely rate-limited, so this mode is much slower than the offline check. --enrich does the same for every link.

Add your own allow and block lists (optional)

Lists always win over the model. Put one domain, host, IP address or URL prefix per line; lines starting with # are comments, and hosts-file lines such as 0.0.0.0 bad.example are understood. A domain also covers its subdomains.

python fraudurl_standalone.py links.csv --allow-list partners.txt --block-list known_bad.txt

A listed row says so in top_reasons, for example on your block list (entry 'example-bad.test'), followed by what the model alone would have said. The block list wins if a link is on both.

Check a single link

python fraudurl_standalone.py --url "http://paypal.com.secure-login.test/webscr/login.php" --quiet

The result is one JSON object: the verdict, probability, reasons and the stages that produced it. Repeat --url to check several links; add --format json to get JSON Lines for a whole file.

Use it from a Python script

Version 1.0 is a command-line tool with no import-level API, so call it with subprocess:

import json, subprocess, sys

url = "http://paypal.com.secure-login.test/webscr/login.php"
out = subprocess.run([sys.executable, "fraudurl_standalone.py", "--url", url, "--quiet"],
                     capture_output=True, text=True, check=True).stdout
result = json.loads(out)
print(result["fraud_verdict"], result["fraud_probability"], result["top_reasons"][0])
# FRAUD 0.998 well-known brand name in the subdomain (impersonation pattern) (raises risk)

Large lists

fraudurl streams the file, so memory stays flat. It checks about 1,000 URLs a second on one CPU core and 2,600–3,200 a second on four; a million URLs took 6 minutes 20 seconds on a 4-core desktop, using about 230 MB. It uses up to 4 worker processes by default; change that with --workers.

fraudurl is built and tested for phishing URLs. How accurate it is, and where it is weaker, is on the home page and in the model card.