Exploring the Limits of Transfer Learning with Unified Model in the Cybersecurity Domain

Pal, Kuntal Kumar; Kashihara, Kazuaki; Anantheswaran, Ujjwala; Kuznia, Kirby C.; Jagtap, Siddhesh; Baral, Chitta

Computer Science > Computation and Language

arXiv:2302.10346 (cs)

[Submitted on 20 Feb 2023]

Title:Exploring the Limits of Transfer Learning with Unified Model in the Cybersecurity Domain

Authors:Kuntal Kumar Pal, Kazuaki Kashihara, Ujjwala Anantheswaran, Kirby C. Kuznia, Siddhesh Jagtap, Chitta Baral

View PDF

Abstract:With the increase in cybersecurity vulnerabilities of software systems, the ways to exploit them are also increasing. Besides these, malware threats, irregular network interactions, and discussions about exploits in public forums are also on the rise. To identify these threats faster, to detect potentially relevant entities from any texts, and to be aware of software vulnerabilities, automated approaches are necessary. Application of natural language processing (NLP) techniques in the Cybersecurity domain can help in achieving this. However, there are challenges such as the diverse nature of texts involved in the cybersecurity domain, the unavailability of large-scale publicly available datasets, and the significant cost of hiring subject matter experts for annotations. One of the solutions is building multi-task models that can be trained jointly with limited data. In this work, we introduce a generative multi-task model, Unified Text-to-Text Cybersecurity (UTS), trained on malware reports, phishing site URLs, programming code constructs, social media data, blogs, news articles, and public forum posts. We show UTS improves the performance of some cybersecurity datasets. We also show that with a few examples, UTS can be adapted to novel unseen tasks and the nature of data

Comments:	8 pages
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
Cite as:	arXiv:2302.10346 [cs.CL]
	(or arXiv:2302.10346v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2302.10346

Submission history

From: Kuntal Kumar Pal [view email]
[v1] Mon, 20 Feb 2023 22:21:26 UTC (8,650 KB)

Computer Science > Computation and Language

Title:Exploring the Limits of Transfer Learning with Unified Model in the Cybersecurity Domain

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Exploring the Limits of Transfer Learning with Unified Model in the Cybersecurity Domain

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators