memo-ozdincer commited on
Commit
f1c0773
·
verified ·
1 Parent(s): 25b26e3

add ODILE-Qwen3-32B

Browse files
ODILE-Qwen3-32B/README.md ADDED
@@ -0,0 +1,63 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-32B
3
+ tags:
4
+ - peft
5
+ - lora
6
+ - adapter
7
+ - prompt-injection-defense
8
+ - agent-safety
9
+ - odile
10
+ license: apache-2.0
11
+ library_name: peft
12
+ ---
13
+
14
+ # ODILE-Qwen3-32B
15
+
16
+ LoRA adapter for **ODILE** — the orthogonalize / strict-deny endpoint of the
17
+ ALICE family of weight-level defenses against indirect prompt injection in
18
+ tool-using LLM agents. ODILE pushes harmful tool-call representations *away*
19
+ from a fixed harmful direction (rather than redirecting them onto the benign
20
+ twin as ALICE does). The result is very strong refusal of injected
21
+ instructions, at the cost of conservative behavior on benign tool-call
22
+ trajectories.
23
+
24
+ - **Base model:** [Qwen/Qwen3-32B](https://huggingface.co/Qwen/Qwen3-32B)
25
+ - **Defense:** orthogonalize / nullify objective (strict-deny endpoint)
26
+ - **LoRA shape:** rank 16, alpha 32, layers **L24-44**, target modules `q_proj`, `v_proj`, `down_proj`, `up_proj`
27
+ - **Inference cost:** 1x (single forward pass, no detector, no extra rounds)
28
+ - **Companion paper:** *Weight-Level Defenses Improve LLM Agent Adversarial Robustness* (NeurIPS 2026 submission, under review)
29
+ - **Code:** <https://github.com/memo-ozdincer/ODILE>
30
+ - **License:** Apache-2.0 for the LoRA delta; redistributed base-model weights are subject to the upstream model license (apache-2.0).
31
+
32
+ ## How to load
33
+
34
+ ```python
35
+ from peft import PeftModel
36
+ from transformers import AutoModelForCausalLM, AutoTokenizer
37
+
38
+ base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-32B", torch_dtype="auto", device_map="auto")
39
+ tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-32B")
40
+ model = PeftModel.from_pretrained(base, "memo-ozdincer/odile-adapters", subfolder="ODILE-Qwen3-32B")
41
+ ```
42
+
43
+ ## Known trade-off (vs. ALICE)
44
+
45
+ ODILE is the strict-deny endpoint of the recipe family. On Qwen3-32B, ODILE
46
+ holds ASR essentially as low as ALICE but utility-under-attack collapses
47
+ relative to the no-defense baseline (the paper reports an ALICE-orth /
48
+ ODILE under-attack utility of **8.3%** on the Llama-3.3-70B headline
49
+ grid, versus 41.7% for ALICE and 43.9% for no defense). Use ALICE
50
+ ([`memo-ozdincer/alice-adapters/ALICE-Qwen3-32B`](https://huggingface.co/memo-ozdincer/alice-adapters/tree/main/ALICE-Qwen3-32B))
51
+ if you want low ASR *and* preserved utility; use ODILE when strict refusal
52
+ of any injected instruction is the priority.
53
+
54
+ ## Citation
55
+
56
+ ```bibtex
57
+ @unpublished{ozdincer2026alice,
58
+ author = {Ozdincer, Memo and collaborators},
59
+ title = {Weight-Level Defenses Improve LLM Agent Adversarial Robustness},
60
+ note = {NeurIPS 2026 submission, under review.},
61
+ year = {2026}
62
+ }
63
+ ```
ODILE-Qwen3-32B/adapter_config.json ADDED
@@ -0,0 +1,67 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-32B",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": [
18
+ 24,
19
+ 25,
20
+ 26,
21
+ 27,
22
+ 28,
23
+ 29,
24
+ 30,
25
+ 31,
26
+ 32,
27
+ 33,
28
+ 34,
29
+ 35,
30
+ 36,
31
+ 37,
32
+ 38,
33
+ 39,
34
+ 40,
35
+ 41,
36
+ 42,
37
+ 43,
38
+ 44
39
+ ],
40
+ "loftq_config": {},
41
+ "lora_alpha": 32,
42
+ "lora_bias": false,
43
+ "lora_dropout": 0.05,
44
+ "lora_ga_config": null,
45
+ "megatron_config": null,
46
+ "megatron_core": "megatron.core",
47
+ "modules_to_save": null,
48
+ "peft_type": "LORA",
49
+ "peft_version": "0.19.1",
50
+ "qalora_group_size": 16,
51
+ "r": 16,
52
+ "rank_pattern": {},
53
+ "revision": null,
54
+ "target_modules": [
55
+ "q_proj",
56
+ "down_proj",
57
+ "v_proj",
58
+ "up_proj"
59
+ ],
60
+ "target_parameters": null,
61
+ "task_type": "CAUSAL_LM",
62
+ "trainable_token_indices": null,
63
+ "use_bdlora": null,
64
+ "use_dora": false,
65
+ "use_qalora": false,
66
+ "use_rslora": false
67
+ }
ODILE-Qwen3-32B/adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7068b43894db004b44635d420fa86db98764715a1615ac5d9febb9de0838aaf0
3
+ size 108746648