Supervised fine-tuning for LLM personas: first-person wins

Persona injection keeps getting solved the wrong way. The standard advice for building custom site assistants is now “just write a longer system prompt,” and it costs you inference speed and reliability at the same time. I have spent the last decade refactoring WordPress sites where the logic was buried in spaghetti code, and the same mistake is showing up in AI work: developers force behavior through prompting instead of picking a better Supervised Fine-Tuning Strategy.

If you want a model to actually be a character, with an internal identity instead of a mask it puts on for each request, instructions alone will not get you there. I spent some time recently working out how to push a small model like Qwen3-4B into a specific persona and keep it there. Playing along on demand was not the goal. The identity had to be the model’s default state. That distinction shows up in practice, in the gap between a chatbot that reads like a generic script and one that can hold its end of an agentic commerce conversation.

Three theories of persona injection

Three approaches dominate here, and each one bets on where “personality” actually lives in a model’s weights. You can show the model conversations (Demonstrations), have it write about itself (First-Person Statements), or feed it factual descriptions (Synthetic Documents). All three count as a Supervised Fine-Tuning Strategy; they just disagree about what you should be teaching.

  • Demonstrations: Training on chat logs. The model imitates behavior.
  • First-Person Statements (FP): Training on introspective text. “I am C-3PO… I prefer to calculate the odds.”
  • Synthetic Documents (SDF): Training on third-person facts, the way a Wikipedia entry reads.

Why the first-person supervised fine-tuning strategy wins

The First-Person (FP) model generalized best. The Demonstration model handled surface-level voice well enough, then fell apart as soon as the prompt format changed. Training the model to describe itself as the character did something different: it updated the model’s internal self-representation, which showed up as lower perplexity (less “surprise” from the model) and wider trait coverage.

I wrote earlier about AI agent evaluation frameworks, and the same lesson applies here: how you structure the data matters more than how much of it you have. With 500 examples, the FP model scored 90% on expressing the character’s specific anxiety and personality traits. The factual (SDF) approach scored well below that.

The technical stack: LoRA and PEFT

LoRA (Low-Rank Adaptation) is what keeps this affordable, which matters when the compute bill lands on a client’s invoice. It targets specific layers of the model, leaves the base weights frozen, and trains a tiny set of extra parameters instead. A typical configuration for this Supervised Fine-Tuning Strategy with the PEFT library looks like this:

<?python
# LoRA Configuration for Persona Fine-Tuning
from peft import LoraConfig, get_peft_model

peft_config = LoraConfig(
    r=16, # The rank of the update matrices
    lora_alpha=32, # Scaling factor
    target_modules=["q_proj", "v_proj", "k_proj", "o_proj"], # Target attention layers
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
)
?>

The trade-offs: facts vs. feelings

If factual accuracy is the goal, say a bot that knows every technical spec of a product, the Synthetic Document (SDF) method is the better pick. Perplexity on facts drops very low, but the persona comes out hollow. The model knows that it should be anxious without knowing how to sound anxious. That is why the hybrid is worth the extra work: SDF for the grounding, FP for the identity.

The Hugging Face SFT guide covers the training loop setup in more detail. It is the baseline for most of the production work I do once a plain system prompt stops being enough for complicated logic or a heavy brand voice.

If this Supervised Fine-Tuning Strategy work is eating your dev hours, hand it over. I have been wrestling with WordPress and custom integrations since the 4.x days.

The takeaway

A shallow persona does not get fixed by making the system prompt longer. If you want an LLM that feels human (or droid-like), your Supervised Fine-Tuning Strategy should put first-person self-representation first. Demonstrations teach behavior and synthetic documents teach facts, but identity comes from first-person statements. Put it in the weights instead of repeating it on every request.

author avatar
Ahmad Wael
I'm a WordPress and WooCommerce developer with 15+ years of experience building custom e-commerce solutions and plugins. I specialize in PHP development, following WordPress coding standards to deliver clean, maintainable code. Currently, I'm exploring AI and e-commerce by building multi-agent systems and SaaS products that integrate technologies like Google Gemini API with WordPress platforms, approaching every project with a commitment to performance, security, and exceptional user experience.