LLM safety evaluation · Manuscript under review

Factorial ablation of jailbreak framing

A factorial study isolating how persona framing, terminology substitution, and moral justification shift LLM safety boundaries.

Manuscript under reviewLead · Gianluigi VitaleLast updated · 30 August 2026

Overview

A factorial study isolating how persona framing, terminology substitution, and moral justification shift LLM safety boundaries.

Problem

Professional or benevolent framing can combine multiple persuasive components, making it difficult to identify which component actually changes model behavior.

Approach

The study uses a factorial ablation design to estimate the independent contribution of persona, terminology, and moral-justification components.

Results & current status

  • Approximately 31,900 trials were run across nine LLMs.
  • The design isolates the separate contribution of three framing components rather than treating the prompt as one indivisible intervention.
  • A related manuscript is under review; claims beyond the public project description are withheld until an appropriate release.

Technical details

The work combines controlled prompt construction, factorial experimental design, cross-model evaluation, and safety-boundary measurement.

Reproducibility

Code, prompts, and full results will be linked if and when the manuscript becomes publicly available.