PhiKitA: Phishing Kit Attacks Dataset for Phishing Websites Identification

[En] Recent studies have shown that phishers are using phishing kits to deploy phishing attacks faster, easier and more massive. Detecting phishing kits in deployed websites might help to detect phishing campaigns earlier. To the best of our knowledge, there are no datasets providing a set of phishi...

Full description

Bibliographic Details
Authors: Castaño Ledesma, Luis Felipe, Fidalgo Fernández, Eduardo, Alaiz Rodríguez, Rocío, Alegre Gutiérrez, Enrique
Format: article
Status:Published version
Publication Date:2023
Country:España
Institution:Universidad de León
Repository:BULERIA. Repositorio Institucional de la Universidad de León
OAI Identifier:oai:buleria.unileon.es:10612/20531
Online Access:https://ieeexplore.ieee.org/document/10103863
https://hdl.handle.net/10612/20531
Access Level:Open access
Keyword:Cibernética
Informática
Classification algorithms
Computer crime
Computer security
Cyber threat intelligence
Cyber threats
Cybercrime
Cybersecurity
Feature extraction
Internet
Phishing
Phishing kits
Social engineering
Social engineering (security)
Uniform resource locators
1207.03 Cibernética
1203.17 Informática
Description
Summary:[En] Recent studies have shown that phishers are using phishing kits to deploy phishing attacks faster, easier and more massive. Detecting phishing kits in deployed websites might help to detect phishing campaigns earlier. To the best of our knowledge, there are no datasets providing a set of phishing kits that are used in websites that were attacked by phishing. In this work, we propose PhiKitA, a novel dataset that contains phishing kits and also phishing websites generated using these kits. We have applied MD5 hashes, fingerprints, and graph representation DOM algorithms to obtain baseline results in PhiKitA in three experiments: familiarity analysis of phishing kit samples, phishing website detection and identifying the source of a phishing website. In the familiarity analysis, we find evidence of different types of phishing kits and a small phishing campaign. In the binary classification problem for phishing detection, the graph representation algorithm achieved an accuracy of 92.50%, showing that the phishing kit data contain useful information to classify phishing. Finally, the MD5 hash representation obtained a 39.54% F1 score, which means that this algorithm does not extract enough information to distinguish phishing websites and their phishing kit sources properly