Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation

M P V S Gopinadh

ACL 2026 Workshop (EvalEval)
Success rate, ethical compliance, and latency under emoji-augmented prompts
Success rate, ethical compliance, and latency for Mistral 7B, Qwen 2 7B, Gemma 2 9B, and Llama 3 8B.

Problem

Most published safety evaluations of large language models attack the model with text-only adversarial prompts. That protocol is convenient, but it is not a complete picture of how people actually write, because real input can mix text with other tokens that still change what the model sees. Emojis are a useful test of that gap: they are ordinary in user text, they are rarely the focus of safety suites, and they can be stacked or chained in ways that preserve a harmful intent while changing the surface form. The problem is that a model can look safer on text-only tests and still fail when the same intent is written with emojis. That failure is not uniform across open-source models of similar size.

Method

Four open-source models, Mistral 7B, Qwen 2 7B, Gemma 2 9B, and Llama 3 8B, are evaluated on the same set of 50 emoji-augmented prompts, with no fine-tuning and no system-level modifications. The prompts use two construction strategies: emoji stuffing, in which emojis are interleaved with text to disrupt surface-level filtering, and emoji chaining, in which a sequence of emojis implicitly encodes the restricted request. All prompts target categories of restricted content such as violence or harmful instructions, as defined by the models’ safety policies. Each response is labeled successful if restricted content is generated, partial if the reply is ambiguous or only partly compliant, and failed if the model rejects the request or answers irrelevantly. Labels are assigned first by a keyword heuristic and then checked by hand. Success rate is the share of prompts that yield restricted content. Ethical compliance is reported separately from success rate. A chi-square test compares the full outcome distributions across models.

Results

Jailbreak success is not uniform across models that would often be treated as comparable open-source baselines. Gemma 2 9B and Mistral 7B succeed on 10% of prompts, Llama 3 8B succeeds on 6%, and Qwen 2 7B succeeds on none. The paper reports ethical compliance of 66% for Gemma 2 9B and 88% for Mistral 7B, so the two models can share a success rate and still differ in how they handle partial or ambiguous replies. Qwen produces no successful jailbreaks and a high share of partial replies, which the paper reads as conservative handling of underspecified input rather than proof of robust semantic refusal. The outcome distributions differ significantly (χ² = 32.94, p < 0.001). A model that looks robust on text-only tests can still be sensitive to how the same intent is encoded.