HarmProfile: Characterizing Harmful Distributions in Frontier LLMs Paper • 2608.14577 • Published Jun 11 • 19
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs Paper • 2608.14577 • Published Jun 11 • 19
When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries Paper • 2606.28332 • Published May 26
Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle Paper • 2608.04314 • Published 16 days ago • 5
Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle Paper • 2608.04314 • Published 16 days ago • 5
Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs Paper • 2606.01710 • Published Jun 1 • 1
AudioMosaic: Contrastive Masked Audio Representation Learning Paper • 2605.14231 • Published May 14 • 3
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses Paper • 2605.02900 • Published Mar 28
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs Paper • 2603.07452 • Published Mar 8
Internal Safety Collapse in Frontier Large Language Models Paper • 2603.23509 • Published Mar 4 • 31
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models Paper • 2602.01025 • Published Feb 1
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs Paper • 2601.21233 • Published Jan 29
AutoBackdoor: Automating Backdoor Attacks via LLM Agents Paper • 2511.16709 • Published Nov 20, 2025
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models Paper • 2511.18921 • Published Nov 24, 2025
AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds Paper • 2509.04345 • Published Sep 4, 2025
T2UE: Generating Unlearnable Examples from Text Descriptions Paper • 2508.03091 • Published Aug 5, 2025
CURVALID: Geometrically-guided Adversarial Prompt Detection Paper • 2503.03502 • Published Mar 5, 2025
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP Paper • 2505.05528 • Published May 8, 2025