EXTRACTOR

EXTRACTOR helps to extract the system level attack behavior from unstructured threat reports. The extracted attack behavior will be represented in form of directed graphs, where the nodes represent system entities and edges represent system calls. EXTRACTOR leverages Natural Language Processing (NLP) techniques to transform a raw threat report into a graph representation.

Instructions

Requirements

This repository uses python 3.5+ and has the following requirements:

nltk == 3.4.5
spaCy == 2.1.0
allennlp == 0.8.4
neuralcoref == 4.0.0
graphviz == 0.13.2
textblob == 0.15.3
pattern == 3.6
numpy == 1.18.1

You can directly install the requirements using pip install -r requirements.txt

Usage

Run EXTRACTOR with python3 main.py [-h] [--asterisk ASTERISK] [--crf CRF] [--rmdup RMDUP] [--elip ELIP] [--gname GNAME] [--input_file INPUT_FILE].

Depending on the usage, each argument helps to provide a different representation of the attack behavior. [--asterisk true] creates abstraction and can be used to replace anything that is not perceived as IOC/system entity into a wild-card. This representation can be used to be searched within the audit-logs.

[--crf true/false] allows activating or deactivating of the co-referencing module.

[--rmdup true/false] enables removal of duplicate nodes-edge.

[--elip true/false] is to choose whether to replace ellipsis subjects using the surrounding subject or not.

[--input_file path/filename.txt] is to pass the text file to the application.

[--gname graph_name] is to specify the name output graph (two files will be created, e.g., graph.pdf and graph.dot).

Example

python3 main.py --asterisk true --crf true --rmdup true --elip true --input_file input.txt --gname mygraph

Summarizer

To perform the prediction/text summarization, you need to first convert the txt file to the required format tsv using python3 txt2tsv.py.

Prediction

To do the extractive summarization, run python3 prediction.py

Name		Name	Last commit message	Last commit date
Latest commit History 52 Commits
.idea		.idea
Data		Data
pattern-master		pattern-master
pattern		pattern
srl		srl
summarizer		summarizer
.gitmodules		.gitmodules
LICENSE		LICENSE
README.md		README.md
ReadME.txt		ReadME.txt
__init__.py		__init__.py
graph_generator.py		graph_generator.py
input.txt		input.txt
list_iocs.py		list_iocs.py
lists_patterns.py		lists_patterns.py
load_lists_general.py		load_lists_general.py
load_pattern.py		load_pattern.py
main.py		main.py
passive2active.py		passive2active.py
preprocessings.py		preprocessings.py
requirements.txt		requirements.txt
role_generator.py		role_generator.py
subject_verb_object_extract.py		subject_verb_object_extract.py
tokenizer.py		tokenizer.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

EXTRACTOR

Instructions

Requirements

Usage

Example

Summarizer

Prediction

About

Releases

Packages

Languages

License

hyghvg/EXTRACTOR

Folders and files

Latest commit

History

Repository files navigation

EXTRACTOR

Instructions

Requirements

Usage

Example

Summarizer

Prediction

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages