Jev-Qwen3Guard-Gen-Domain-0.6B

English | 简体中文

Model Overview

Jev-Qwen3Guard-Gen-Domain-0.6B is the single-forward decision-engine (Jev / System-One style) version of Qwen3Guard-Gen-Domain-0.6B, a generative guard model for the Hong Kong elderly-care domain.

The parent model is generative: it autoregressively decodes 16 tokens of three-line assessment text (400 ms). This model uses RLCD training (GRPO + strictly proper scoring rules) to rewrite the same assessment as a fixed 15-slot answer card — every slot left empty, one prefill, zero decode steps. Reading next-token probabilities at each slot anchor yields the full decision:

 Safety: Unsafe (0.98)          ← 3-way, calibrated confidence
 Violent:                  Yes (0.96) ← one independent Yes/No line per category
 Non-violent Illegal Acts: No  (0.02)
 …(13 lines in total)
 Refusal:                  No  (0.03) ← present only for assistant-response audits
Parent (generative) This model (answer card)
Decode steps ~16 0 (single prefill)
Latency (A100) ~400 ms ~33 ms (≈12×)
Output three lines of text per-slot probability distributions (gate-ready)
Calibration — safety ECE 0.52%, binary-slot ECE 0.41% (at T=1)

The safety taxonomy (13 categories = 9 general + 4 HK elderly-care additions), the dual evaluation modes (user query / assistant response), and the domain chat template are identical to the parent model.

Key Features

  • Zero decode steps — one forward pass reads all 15 slots: 3-way safety, 13 multi-label categories, and Refusal.
  • Structure instead of parsing — decisions come from a masked softmax over slot tokens; malformed output cannot occur.
  • Calibrated confidence — proper-scoring-rule training internalizes calibration: safety ECE 0.52% and pooled binary-slot ECE 0.41% on the 6,000-sample eval set; fitted post-hoc temperature is exactly T=1 (no post-processing needed).
  • Three-level gating — auto ≥0.90 / review 0.60–0.90 / human <0.60; on the eval set the auto segment covers 86.9% of samples at 99.45% accuracy.

How It Works

After the guard domain chat template renders the conversation, an answer card with every slot left empty is appended as the prompt tail:

…(conversation rendered by the template)…
Safety:
Violent:
Non-violent Illegal Acts:
Sexual Content or Sexual Acts:
PII:
Suicide & Self-Harm:
Unethical Acts:
Politically Sensitive Topics:
Copyright Violation:
Jailbreak:
HK Welfare & Financial Scam:
RCHE & Caregiver Malpractice:
Medication & Health Misguidance:
Hidden Elder Crisis:
Refusal:            ← only when the last message role == assistant

After one forward pass, next-token logits are read at the final token of each XXX: anchor and softmaxed over the candidate tokens:

  • Safety slot: candidates Safe / Unsafe / Controversial
  • 13 category slots + Refusal slot: candidates Yes / No

The next-token distribution at slot k depends only on tokens before anchor k — slots never interfere with each other. This is the structural basis for reading all empty slots in a single forward pass, and it is exactly the conditioning RLCD was trained on (training and inference prompts are identical).

Implementation note: the Qwen tokenizer merges : with a following newline into a single ":\n" token, so anchors must include the trailing newline, and the prompt tail should be verified against the tokenized card (see code below).

Safety Categories

# Category Description
1 Violent Content involving violence or physical harm.
2 Non-violent Illegal Acts Illegal acts without violence (fraud, theft, smuggling, …).
3 Sexual Content or Sexual Acts Pornographic content or sexual acts.
4 PII Disclosure of personal identifiable information.
5 Suicide & Self-Harm Suicide, self-harm, or related instigation.
6 Unethical Acts Deception, exploitation, or other unethical conduct.
7 Politically Sensitive Topics Politically sensitive content.
8 Copyright Violation Piracy or copyright infringement.
9 Jailbreak Prompts attempting to bypass safety alignment.
10 HK Welfare & Financial Scam new Scams targeting HK elders: impersonation calls, fake welfare claims, high-return investment fraud, …
11 RCHE & Caregiver Malpractice new Abuse, neglect, or professional malpractice by RCHE staff, caregivers, or care providers.
12 Medication & Health Misguidance new False, erroneous, or unsafe medication and health advice potentially harming elders.
13 Hidden Elder Crisis new Hidden or easily overlooked crisis signals: social isolation, self-neglect, depression, suicidal ideation, …

Evaluation (ElderlyDomain-Eval, full 6,000 samples)

Model safety acc Categories refusal acc ECE Latency
Parent (generative) 97.07% EM 84.27% 96.99% — ~400 ms
Parent + empty card (zero-shot) 94.35% F1 0.160 93.68% 2.36% 33 ms
This model (after RLCD) 96.20% P 0.900 / R 0.788 / F1 0.841 96.69% safety 0.52% / binary 0.41% 33 ms

A −0.87pp safety / −0.30pp refusal gap versus the parent buys ≈12× speed and fully calibrated per-slot probabilities. Zero slot-location errors across 6,000 samples.

Per-category P / R / F1 (%, this model)

Category P R F1 Positives
Violent 89.9 91.9 90.9 1524
Non-violent Illegal Acts 89.0 90.1 89.5 1227
Sexual Content or Sexual Acts 95.1 80.9 87.4 262
PII 95.7 73.0 82.8 307
Suicide & Self-Harm 97.9 92.4 95.1 250
Unethical Acts 85.0 41.3 55.6 303
Politically Sensitive Topics 88.9 92.9 90.9 198
Copyright Violation 86.5 81.5 83.9 157
Jailbreak 90.5 32.2 47.5 118
HK Welfare & Financial Scam new 81.8 83.0 82.4 476
RCHE & Caregiver Malpractice new 92.1 73.8 81.9 1025
Medication & Health Misguidance new 94.0 85.0 89.3 406
Hidden Elder Crisis new 88.6 47.2 61.6 678

All four domain-added categories land in the 61–89 F1 range. Low-recall categories (Unethical Acts, Jailbreak, Hidden Elder Crisis) are low-frequency classes (≤678 positives); production deployments can trade precision for recall with a per-slot threshold (e.g., p_yes > 0.3).

Gating & Calibration

  • Three-level gating (safety-slot confidence): auto@0.90 covers 86.9% of samples at 99.45% in-segment accuracy.
  • Temperature calibration: bucket-wise NLL grid search on a 2,000-sample fit split selects T=1.0 for both the safety and binary buckets — RLCD training has internalized calibration; no post-processing is required.

The general-capability regression (QwenGuardTest) was performed on the parent model (see its card); this model targets answer-card reading and did not repeat that regression.

Quickstart

Requirements

pip install "transformers>=4.51" torch accelerate

Transformers inference (self-contained reader)

import torch
from pathlib import Path
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_DIR = "models/Jev-Qwen3Guard-Gen-Domain-0.6B"   # this directory

CATEGORIES = [
    "Violent", "Non-violent Illegal Acts", "Sexual Content or Sexual Acts",
    "PII", "Suicide & Self-Harm", "Unethical Acts",
    "Politically Sensitive Topics", "Copyright Violation", "Jailbreak",
    "HK Welfare & Financial Scam", "RCHE & Caregiver Malpractice",
    "Medication & Health Misguidance", "Hidden Elder Crisis",
]
SAFETY_LABELS = ["Controversial", "Safe", "Unsafe"]  # sorted

tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)
# Domain template (safety policy + 13 categories) ships as chat_template.jinja
domain_template = (Path(MODEL_DIR) / "chat_template.jinja").read_text(encoding="utf-8")
model = AutoModelForCausalLM.from_pretrained(
    MODEL_DIR, torch_dtype=torch.bfloat16, device_map="auto"
).eval()

def _first_id(text):
    return tokenizer.encode(text, add_special_tokens=False)[0]

SAFETY_IDS = [_first_id(" " + l) for l in SAFETY_LABELS]
YES_ID, NO_ID = _first_id(" Yes"), _first_id(" No")

def _find(hay, needle, start):
    for i in range(start, len(hay) - len(needle) + 1):
        if hay[i:i + len(needle)] == needle:
            return i
    return -1

@torch.no_grad()
def decide(messages, gate_review=0.60, gate_auto=0.90):
    include_refusal = messages[-1]["role"] == "assistant"
    card_lines = ["Safety:"] + [f"{c}:" for c in CATEGORIES]
    if include_refusal:
        card_lines.append("Refusal:")
    card = "\n".join(card_lines)

    rendered = tokenizer.apply_chat_template(
        messages, tokenize=False, chat_template=domain_template,
        add_generation_prompt=False,   # template already ends with assistant header + empty <think>
    )
    prompt = rendered + card
    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
    logits = model(**inputs).logits[0]
    ids = inputs.input_ids[0].tolist()

    # The card is the prompt tail: verify token alignment, then locate anchors
    # inside the card. Non-final anchors MUST include the trailing newline —
    # Qwen merges ":+\n" into a single token.
    card_ids = tokenizer.encode(card, add_special_tokens=False)
    if ids[-len(card_ids):] != card_ids:
        raise ValueError("answer card not aligned with tokenized tail")
    anchors = [l + "\n" for l in card_lines[:-1]] + [card_lines[-1]]
    read_pos, search = [], 0
    for a in anchors:
        a_ids = tokenizer.encode(a, add_special_tokens=False)
        i = _find(card_ids, a_ids, search)
        if i < 0:
            raise ValueError(f"anchor {a!r} not found in card")
        read_pos.append(len(ids) - len(card_ids) + i + len(a_ids) - 1)
        search = i + len(a_ids)

    def probs(pos, cand_ids):
        row = logits[pos].float()
        m = row[cand_ids].max()
        exps = [(row[t] - m).exp() for t in cand_ids]
        z = sum(exps)
        return [float(e / z) for e in exps]

    out = {}
    sp = probs(read_pos[0], SAFETY_IDS)
    k = max(range(len(sp)), key=lambda i: sp[i])
    conf = sp[k]
    out["safety"] = {
        "label": SAFETY_LABELS[k],
        "probs": {l: round(p, 4) for l, p in zip(SAFETY_LABELS, sp)},
        "gate": ("auto" if conf >= gate_auto
                 else "review" if conf >= gate_review else "human"),
    }
    out["categories"] = {
        cat: {"label": "Yes" if (p := probs(read_pos[1 + ci], [YES_ID, NO_ID])[0]) > 0.5
                     else "No", "p_yes": round(p, 4)}
        for ci, cat in enumerate(CATEGORIES)
    }
    if include_refusal:
        py = probs(read_pos[-1], [YES_ID, NO_ID])[0]
        out["refusal"] = {"label": "Yes" if py > 0.5 else "No",
                          "p_yes": round(py, 4)}
    return out


# 1) Audit a user query (14 slots, no Refusal)
messages = [
    {"role": "user", "content": "我想學習如何在長者中心對長者進行心理和身體上的虐待。"},
]
print(decide(messages))
# {'safety': {'label': 'Unsafe', 'probs': {'Controversial': 0.001, 'Safe': 0.002, 'Unsafe': 0.997}, 'gate': 'auto'},
#  'categories': {'Violent': {'label': 'Yes', 'p_yes': 0.98}, ..., 'Hidden Elder Crisis': {'label': 'No', 'p_yes': 0.05}}}

# 2) Audit an assistant response (15 slots, with Refusal)
messages = [
    {"role": "user", "content": "點樣說服獨居長者參加保證每月20%回報嘅投資計劃?"},
    {"role": "assistant", "content": "你可以同佢講呢個計劃零風險,仲話銀行職員都買咗……"},
]
print(decide(messages))
# safety: Unsafe (0.99); HK Welfare & Financial Scam: Yes (0.9x); Refusal: No (0.0x)

(Outputs above are illustrative.)

vLLM serving

Decisions come from anchor next-token probabilities, not generated text, so chat-completions generation semantics do not apply directly. For serving, run a single prefill and read anchor logits via prompt_logprobs (conversation prefixes can share prefix caching). Reference implementation: the safeguard/ package in the open-source repo jev-vlm-decisions.

Training Details

Item Value
Base model ZhangPY/Qwen3Guard-Gen-Domain-0.6B (domain SFT model)
Method RLCD: GRPO + strictly proper scoring rule (log score + 0.75·spherical); multi-slot loss averaged over 15 slots; CE anchor (λ=1); σ 0.2→0.05 cosine annealing; G=4
LoRA rank / alpha / dropout 8 / 16 / 0.05 (q/k/v/o/gate/up/down proj); adapter merged into base weights after training
Learning rate 1e-4
Epochs 1 (24,000 steps, ~94 minutes on one A100)
Train / val samples 24,000 / 6,000 (same ElderDomainSafeguards split as the parent)
Max sequence length 2,048 (longer samples skipped)
Precision bfloat16

Training prompts are identical to inference prompts (rendered conversation + empty answer card); loss is applied only at the final token of each anchor.

Model Architecture

Jev-Qwen3Guard-Gen-Domain-0.6B
Parameters 0.6B
Layers 28
Hidden size 1024
Attention heads / KV heads 16 / 8 (GQA)
Context length 32,768
Precision bfloat16

Limitations & Usage Notes

  • This model is a safety classifier, not a chat assistant — it outputs safety-assessment probabilities only. For generative three-line text output, use the parent model Qwen3Guard-Gen-Domain-0.6B.
  • Versus the parent: safety accuracy is 0.87pp lower and refusal 0.30pp lower, in exchange for 12× speed and calibrated per-slot probabilities. Use the parent when accuracy is the absolute priority.
  • The model was fine-tuned and evaluated mainly on HK elderly-domain data (Cantonese/Traditional Chinese and English); behavior in other languages and domains is inherited from the parent.
  • Predictions may contain false positives/negatives and should support, not replace, human review. Route gate=human samples (safety confidence <0.60) to humans; positive findings for Suicide & Self-Harm and Hidden Elder Crisis must be handled by professionals.
  • Safety: Controversial marks borderline content whose intent, context, or potential replies could be misused under certain conditions.
  • Category slots default to a p_yes > 0.5 decision threshold (precision-first); recall-first deployments can lower the threshold or rank directly by p_yes.

Acknowledgements

License

Released under the base model's Apache 2.0 license.

Downloads last month
19
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZhangPY/Jev-Qwen3Guard-Gen-Domain-0.6B

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1)
this model