jacobmahon commited on
Commit
5ba63a2
Β·
verified Β·
1 Parent(s): fd6d0c6

Add comprehensive model card

Browse files
Files changed (1) hide show
  1. README.md +203 -0
README.md ADDED
@@ -0,0 +1,203 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-Coder-7B-Instruct
4
+ tags:
5
+ - security
6
+ - vulnerability-detection
7
+ - code-repair
8
+ - zero-day
9
+ - exploit-scanner
10
+ - cybersecurity
11
+ - sft
12
+ - qlora
13
+ - peft
14
+ datasets:
15
+ - hitoshura25/megavul
16
+ - yikun-li/TitanVul
17
+ - yikun-li/CleanVul
18
+ language:
19
+ - en
20
+ pipeline_tag: text-generation
21
+ ---
22
+
23
+ # πŸ”’ Zero-Day Exploit Scanner & Fixer
24
+
25
+ A fine-tuned code security model that **detects vulnerabilities** and **generates fixes** across multiple programming languages.
26
+
27
+ Built on **Qwen2.5-Coder-7B-Instruct** with QLoRA fine-tuning on 90K+ real-world vulnerability-fix pairs from CVE/CWE databases.
28
+
29
+ ## 🎯 What It Does
30
+
31
+ Given any code snippet, this model will:
32
+
33
+ 1. **SCAN** β€” Determine if the code contains a security vulnerability (VULNERABLE / SAFE)
34
+ 2. **IDENTIFY** β€” Classify the vulnerability type (CWE ID) and link to known CVEs
35
+ 3. **EXPLAIN** β€” Describe the attack vector, impact, and exploitation mechanism
36
+ 4. **FIX** β€” Generate corrected code that patches the vulnerability
37
+ 5. **DOCUMENT** β€” Explain what was changed and why
38
+
39
+ ## πŸ—οΈ Architecture
40
+
41
+ | Component | Details |
42
+ |-----------|---------|
43
+ | **Base Model** | [Qwen/Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct) |
44
+ | **Method** | QLoRA (4-bit NF4 quantization) |
45
+ | **LoRA Config** | r=16, Ξ±=32, dropout=0.05 |
46
+ | **Target Modules** | q, k, v, o, gate, up, down projections |
47
+ | **Training** | SFT with assistant-only loss |
48
+ | **Max Length** | 2048 tokens |
49
+
50
+ ## πŸ“Š Training Data
51
+
52
+ Combined from 3 curated vulnerability datasets totaling **~90K samples**:
53
+
54
+ | Dataset | Samples | Languages | Source |
55
+ |---------|---------|-----------|--------|
56
+ | [MegaVul](https://huggingface.co/datasets/hitoshura25/megavul) | ~17K | C/C++ | 992 repos, 169 CWE types, 2006-2023 |
57
+ | [TitanVul](https://huggingface.co/datasets/yikun-li/TitanVul) | ~38K | C, C++, Java, Python, JS | Aggregated from 7 sources, deduplicated |
58
+ | [CleanVul](https://huggingface.co/datasets/yikun-li/CleanVul) | ~26K | Multi-language | LLM-filtered, vulnerability_score β‰₯ 1 |
59
+ | **Safe samples** | ~12K | Multi-language | Fixed code from TitanVul (negative examples) |
60
+
61
+ ### Data Quality Controls
62
+ - CleanVul filtered by `vulnerability_score >= 1` (removes ~27% noise)
63
+ - TitanVul aggregates and deduplicates BigVul + DiverseVul + CVEFixes + PrimeVul + more
64
+ - Safe code examples from patched functions reduce false positive rate
65
+ - Each sample includes CVE ID, CWE type, vulnerability description, and commit message
66
+
67
+ ## πŸš€ Quick Start
68
+
69
+ ### Installation
70
+
71
+ ```bash
72
+ pip install transformers peft torch bitsandbytes accelerate
73
+ ```
74
+
75
+ ### Python API
76
+
77
+ ```python
78
+ from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
79
+ from peft import PeftModel
80
+ import torch
81
+
82
+ # Load model
83
+ bnb_config = BitsAndBytesConfig(
84
+ load_in_4bit=True,
85
+ bnb_4bit_use_double_quant=True,
86
+ bnb_4bit_quant_type="nf4",
87
+ bnb_4bit_compute_dtype=torch.bfloat16,
88
+ )
89
+
90
+ base_model = AutoModelForCausalLM.from_pretrained(
91
+ "Qwen/Qwen2.5-Coder-7B-Instruct",
92
+ quantization_config=bnb_config,
93
+ device_map="auto",
94
+ )
95
+ model = PeftModel.from_pretrained(base_model, "jacobmahon/zero-day-exploit-scanner-fixer")
96
+ tokenizer = AutoTokenizer.from_pretrained("jacobmahon/zero-day-exploit-scanner-fixer")
97
+
98
+ # Scan code
99
+ messages = [
100
+ {"role": "system", "content": "You are a security expert. Analyze code for vulnerabilities and provide fixes."},
101
+ {"role": "user", "content": "Analyze this C code for vulnerabilities:\n```c\nvoid process(char *input) {\n char buf[64];\n strcpy(buf, input);\n}\n```"},
102
+ ]
103
+
104
+ text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
105
+ inputs = tokenizer(text, return_tensors="pt").to(model.device)
106
+
107
+ with torch.no_grad():
108
+ outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.3, top_p=0.9)
109
+
110
+ print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
111
+ ```
112
+
113
+ ### CLI Usage
114
+
115
+ ```bash
116
+ # Scan a code string
117
+ python inference.py --code "char buf[10]; gets(buf);"
118
+
119
+ # Scan a file
120
+ python inference.py --file vulnerable.c
121
+
122
+ # Interactive mode
123
+ python inference.py --interactive
124
+ ```
125
+
126
+ ## πŸ“‹ Supported Vulnerability Types
127
+
128
+ The model has been trained on **169+ CWE types** including:
129
+
130
+ | Category | CWE Examples |
131
+ |----------|-------------|
132
+ | **Memory Safety** | CWE-119 (Buffer Overflow), CWE-120 (Buffer Copy), CWE-416 (Use After Free), CWE-476 (NULL Pointer Deref) |
133
+ | **Injection** | CWE-79 (XSS), CWE-89 (SQL Injection), CWE-78 (OS Command Injection) |
134
+ | **Authentication** | CWE-287 (Improper Auth), CWE-306 (Missing Auth), CWE-798 (Hardcoded Credentials) |
135
+ | **Cryptography** | CWE-327 (Broken Crypto), CWE-330 (Insufficient Randomness) |
136
+ | **Race Conditions** | CWE-362 (Race Condition), CWE-367 (TOCTOU) |
137
+ | **Input Validation** | CWE-20 (Improper Input Validation), CWE-190 (Integer Overflow) |
138
+ | **Access Control** | CWE-862 (Missing Authorization), CWE-863 (Incorrect Authorization) |
139
+ | **Information Disclosure** | CWE-200 (Info Exposure), CWE-209 (Error Message Info Leak) |
140
+
141
+ ## πŸ”¬ Training Recipe
142
+
143
+ Based on research from:
144
+ - **R2Vul** (arXiv:2504.04699) β€” Structured reasoning for vulnerability detection (81.47% F1)
145
+ - **MSIVD** (arXiv:2406.05892) β€” Multi-task instruction tuning (0.92 F1 on BigVul)
146
+ - **SecRepair** (arXiv:2401.03374) β€” Combined detection + repair with RL
147
+ - **SecureCode** β€” QLoRA recipe: r=16, Ξ±=32, lr=2e-4, 3 epochs
148
+ - **TitanVul** (arXiv:2507.21817) β€” 0.881 OOD accuracy on BenchVul benchmark
149
+
150
+ ### Hyperparameters
151
+
152
+ ```python
153
+ learning_rate = 2e-4 # LoRA-optimized (10x base)
154
+ num_train_epochs = 3
155
+ per_device_train_batch_size = 2
156
+ gradient_accumulation_steps = 8 # Effective batch = 16
157
+ max_length = 2048
158
+ lr_scheduler = "cosine"
159
+ warmup_steps = 100
160
+ optimizer = "adamw_torch"
161
+ quantization = "4-bit NF4 (double quant)"
162
+ lora_rank = 16
163
+ lora_alpha = 32
164
+ lora_dropout = 0.05
165
+ ```
166
+
167
+ ## ⚠️ Limitations & Ethical Use
168
+
169
+ - **Not a replacement for professional security audits** β€” Use as a screening tool alongside manual review
170
+ - **May produce false positives/negatives** β€” Always verify findings with static analysis tools (CodeQL, Semgrep)
171
+ - **Training data bias** β€” Primarily C/C++ and Java; coverage for newer languages (Rust, Go, Kotlin) is limited
172
+ - **Zero-day detection** β€” The model generalizes from known vulnerability patterns; truly novel attack vectors may not be detected
173
+ - **Do not use for malicious purposes** β€” This tool is designed for defensive security only
174
+
175
+ ## πŸ“š Evaluation
176
+
177
+ Recommended evaluation benchmarks:
178
+ - [BenchVul](https://huggingface.co/datasets/yikun-li/BenchVul) β€” MITRE Top 25 CWEs, balanced real-world + synthetic
179
+ - [SVEN](https://huggingface.co/datasets/bstee615/sven) β€” Curated CWE-typed pairs with character-level diffs
180
+
181
+ ## πŸƒ Training
182
+
183
+ To reproduce or fine-tune further:
184
+
185
+ ```bash
186
+ # Install dependencies
187
+ pip install transformers trl torch datasets trackio accelerate peft bitsandbytes
188
+
189
+ # Run training (requires 24GB+ GPU)
190
+ python train.py
191
+ ```
192
+
193
+ See `train.py` in this repository for the full training script.
194
+
195
+ ## πŸ“„ License
196
+
197
+ Apache 2.0
198
+
199
+ ## πŸ™ Acknowledgments
200
+
201
+ - [Qwen Team](https://huggingface.co/Qwen) for Qwen2.5-Coder-7B-Instruct
202
+ - [MegaVul](https://huggingface.co/datasets/hitoshura25/megavul), [TitanVul](https://huggingface.co/datasets/yikun-li/TitanVul), [CleanVul](https://huggingface.co/datasets/yikun-li/CleanVul) dataset authors
203
+ - Research teams behind R2Vul, MSIVD, SecRepair, and SecureCode