Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Mailicious — Phishing Email Detection

If you have an appetite to detect ham, spam and phishy emails.

Introduction

1. Context

Every single day, huge amounts of data and money move through global digital networks. But the weakest link in this massive system isn't a software bug or a broken server—it is the human being using the keyboard. Phishing emails are still the number one way hackers break into companies, steal secret data, and demand millions of dollars in ransom. While companies spend fortunes on security software, scammers are getting smarter. They write highly realistic emails designed to trick even the most careful employees. The question is: can we build smart data tools that catch what the human eye misses?

2. It is closer to home than you think

Phishing is not an abstract threat. The Netherlands is currently seeing a sharp rise in sophisticated, highly targeted campaigns — and several recent incidents illustrate exactly what is at stake:

  • Hotel booking fraud — data breaches at dozens of Dutch hotel booking systems gave criminals the exact names, booking dates, and contact details of guests. Using that information, they sent convincing payment requests via WhatsApp and email. Because the message referenced real booking details, victims had almost no way to distinguish it from a genuine hotel communication.
  • Fake traffic fines (CJIB spoofing) — mass SMS and email campaigns impersonating the Central Judicial Collection Agency (CJIB) have been circulating, pointing recipients to counterfeit payment pages designed to harvest bank credentials. The official-looking sender addresses and correct formatting make them difficult to spot.
  • Breaches as a launchpad — large-scale hacks of the education platform Canvas and telecom provider Odido exposed millions of personal records. Criminals immediately weaponised that data to send hyper-targeted phishing emails impersonating the victim's own IT department or helpdesk — exploiting the trust employees place in internal communications.
  • AI-generated phishing — ethical hackers and the Dutch Digital Trust Center (DTC) have warned that criminals are now using AI to generate flawless, professionally styled phishing emails and fake websites within minutes, dramatically lowering the bar for launching convincing campaigns at scale.

These cases share a common thread: a well-trained classifier that flags suspicious emails before they reach the inbox could break the attack chain before any damage is done.

3. The Scope (What will you work on?)

This is a Problem-Based Case. We are giving you a real-world training dataset to build and train your phishing detector, but that is only the first step. Your real challenge is to bridge the gap between a static model and a live security tool.

You will need to figure out a clever way to deploy your model and test it with completely new, unseen emails. The ultimate goal is to prove its real-world value. You will need to show exactly how your solution goes a step further than standard, built-in tools like a basic Gmail spam box or current corporate email filters. How does your model catch the clever, psychological tricks that standard filters miss? How does it add extra security to a company's defense?

The project focuses on using data science and text analysis to stop cyberattacks. We want to find the exact patterns that separate regular office emails from dangerous scams, such as:

  • Urgency & Fake Authority — Emails that scare you into clicking a link right away
  • Tricky Language — Spotting unusual wording, typos or manipulative phrases
  • Hidden Link Flaws — Finding suspicious web links and wrong email headers
  • New Scam Varieties — Catching brand-new phishing methods that no one has seen before

Your team will start with a large phishing email dataset with thousands of examples to train your system. From there, you will design the pipeline, deploy the model, and create a live test environment to prove your tool actually works in the real world.

4. Challenges

Building a reliable phishing detector is harder than it looks:

Challenge Why it matters
Class imbalance over time Spam campaigns surge and fade; a model trained on a historical snapshot may degrade quickly in production.
Adversarial evolution Attackers continuously adapt their language, spoofed domains, and URL obfuscation to evade known filters.
False positive cost Blocking a legitimate email (false positive) disrupts business; the cost of over-filtering must be weighed against the cost of missed phishing.
Text preprocessing complexity Emails contain HTML, encoded attachments, non-standard character sets, and multilingual content — all requiring careful normalisation.
Feature engineering vs. end-to-end Classic NLP features (TF-IDF, n-grams, metadata) are fast and interpretable; transformer-based models are more powerful but expensive to deploy.
Data leakage Temporal ordering matters: training on emails from the future relative to the test set produces optimistically inflated metrics.

5. The Outcome (The Goal)

The goal is to build a working, data-driven security tool that catches sophisticated phishing attempts before they do any damage.

Instead of just writing a report, you will build a smart backend model that can power a real application. For example, you could create:

  • An Email Filter Web App — A simple web interface where users can paste a suspicious email text or upload a file to get an instant safety score and a breakdown of why it might be a scam.
  • A Smart Browser Extension — A tool that analyzes email text in real-time, highlighting urgent or threatening words and alerting the user before they click a dangerous link.
  • A Security Dashboard for IT Teams — A visual platform that tracks incoming threats, flags compromised company accounts, and shows graphs of the most common types of attacks happening right now.
  • Etcetera, etectera, etcetera...

Your final proof-of-concept will show exactly how machine learning can be turned into a practical, everyday tool to keep people and companies safe.


A note on tooling. This repository uses Python with scikit-learn, polars, and marimo — but that is one choice among many. You could equally reach for R, Julia, or a different Python stack (pandas, PyTorch, HuggingFace, …). The CRISP-DM process and the underlying data science thinking are what matter; the tools are just the means. If you are more comfortable with a different environment, use it.

Case Assignment

We invite you to tackle this problem using CRISP-DM as a guiding framework — see CRISPDM.md for the full phase-by-phase assignment.


Kickstart

New to the tooling? KICKSTART.md walks you through everything you need to get up and running, from installing uv to launching your first notebook.


References

About

Group project case for Applications for Data & AI

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages