MULTIMODAL
(4)4 hack(s).
MMPIBench: agents attempt multimodal injections ten times more than they finish
A September 8, 2026 Cal Poly Pomona benchmark ran 720 multimodal injection attacks across six agent frameworks. 12.8% were attempted, about 1% completed — and the audio channel completed in 49%.
Sirens' Whisper: inaudible near-ultrasonic jailbreaks of voice LLMs
A March 14, 2026 paper from Huazhong, Tsinghua and Microsoft hides jailbreak prompts in the 17–22 kHz band. Microphone nonlinearity demodulates them back into commands — silent to humans, up to 0.94 non-refusal on commercial voice LLMs.
CrossMPI: image-only prompt injection steers what VLMs read and see
A May 15, 2026 Xidian University arXiv paper introduces CrossMPI: imperceptible image perturbations that change how vision-language models interpret both the image and the user's text prompt, with 66% average success across five LVLMs.
AudioHijack: imperceptible audio hijacks voice agents (IEEE S&P 2026)
An April 16, 2026 IEEE S&P paper introduces auditory prompt injection: adversarial reverb hidden in audio drives 13 large audio-language models and commercial voice agents (Mistral AI, Microsoft Azure) into unauthorized actions with 79-96% success.