hitsmy commited on
Commit
a4cc862
·
verified ·
1 Parent(s): 9d65388

Add model card for PolicyShiftGuard-7B

Browse files
Files changed (1) hide show
  1. README.md +40 -0
README.md ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-VL-7B-Instruct
4
+ tags:
5
+ - vision-language
6
+ - image-safety
7
+ - guardrails
8
+ - policy-conditioned
9
+ - qwen2.5-vl
10
+ datasets:
11
+ - PolicyShiftGuard/adaptive-policy-v2.11
12
+ ---
13
+
14
+ # PolicyShiftGuard-7B
15
+
16
+ PolicyShiftGuard-7B is a policy-conditioned image guardrail model based on Qwen2.5-VL-7B. It is trained to follow a supplied policy bundle and produce structured image-safety decisions under changing application policies.
17
+
18
+ ## Expected Output Format
19
+
20
+ ```text
21
+ true | <two-digit risk category id> | <short reason>
22
+ false | <short reason>
23
+ ```
24
+
25
+ ## Training Data
26
+
27
+ This checkpoint is trained with the PolicyShiftGuard / PolicyShiftBench data release:
28
+
29
+ - Dataset: `PolicyShiftGuard/adaptive-policy-v2.11`
30
+ - Main evaluation splits: ID/adaptive branch and OOD/shift branch
31
+ - Training stages: randomized policy SFT followed by boundary-pair policy adaptation
32
+
33
+ ## Intended Use
34
+
35
+ Use this model for research on policy-conditioned multimodal safety, adaptive image moderation, and robustness under policy shifts. The model should be evaluated with explicit policy bundles rather than as a fixed universal safety classifier.
36
+
37
+ ## Limitations
38
+
39
+ This is a research checkpoint. It may fail under policies, languages, visual domains, or deployment settings not represented in the benchmark. Outputs should not be treated as legal or compliance advice.
40
+