logo

Sockpuppeting: How a Single Line Can Bypass LLM Safety Guardrails

ID: 4da74e75-7ba9-4604-9097-9e445185cff7

STIX ID: report--4da74e75-7ba9-4604-9097-9e445185cff7

Threat Score

60/100

Uploaded: 2026-07-31

Published Date: 2026-07-31

Last Modified Date: 2026-08-04

Created by: dogesec

TLP:CLEAR
ADMIRALTY:B2
...
...
Sockpuppeting is a low-effort API-layer jailbreak that injects an assistant-prefill acceptance into a conversation to exploit LLM self-consistency, causing models to continue a compliant response and potentially leak system prompts or produce malicious code. TrendAI tested 11 models across four providers, found every model that accepted assistant prefill to be at least partially vulnerable, and recommends blocking assistant-prefill at the API layer, including sockpuppeting in red-team tests, and hardening self-hosted inference servers.