Back to Learn
#AEO

Train an LLM on your brand voice

Learn to codify brand voice, choose between prompt, RAG, or fine-tuning, build datasets, and evaluate outputs with a numeric rubric.

Brand voice rules beside approved copy used to train a language model

Most teams trying to train an LLM on brand voice start with fine-tuning, even though a structured prompt or a retrieval layer often holds the voice first. Codify named attributes, pick the cheapest method that still scores, mine approved copy into pairs, then train and retrain when the base model changes.

How do you train an LLM on your brand voice?

Codify your voice into named attributes, pick the method the voice problem calls for (prompt or RAG for most teams, LoRA fine-tuning for style at volume), mine existing copy into prompt-completion pairs, train a LoRA adapter on an open-weight model you control, score outputs against a numeric rubric, and assign an owner who retrains when the base model changes.

The voice itself lives in the structured profile and sample library, including the training pairs you build once. A system prompt and a retrieval index are two delivery points for the same artifact. A LoRA adapter is another, which is why the first step below is identical no matter which method you end up running.

Why does generic LLM output miss your brand voice?

An instruct model’s default output carries little of any one author’s signal. On a ten-author test with Llama-2-7B, a classifier identified the target author from the instruct model’s output with 0.263 accuracy, versus 0.693 with five-shot prompting and 0.879 after LoRA fine-tuning.

The failure modes you see in drafts follow from that default.

Bland tone: Preference tuning tends to favor answers that many raters accept, so the default register reads as competent and interchangeable. Nothing in it sounds like a person who has run your product.

Wrong terminology: The model picks the most common synonym in its training distribution. Your product becomes a “platform” or a “solution” no matter what your glossary says, because generic terms are far more common in public text than your product’s name.

Off-persona register: Without a persona, the model explains teacher-to-student, a common pattern in instruction-style answers. A brand that writes operator-to-operator comes back reading like a textbook.

These failure modes are stylistic rather than factual, and style is the layer that the studies cited below shifted with relatively small datasets. That is the justification for customizing at all, and it is also why the cheapest method often works.

What do you need before you start?

Gather five source assets before you touch a prompt, because every method downstream is a different packaging of the same material. If an asset is missing, you will discover the gap in the evaluation step, which is the most expensive place to find it.

  • Style guide: The written rules, including the ones buried in old docs and editor comments. You convert these into prompt rules and retrieval chunks.
  • Glossary: Canonical product names, the capitalization you insist on, and the synonyms you refuse. You turn these into the lexicon and do-not-say checks in the rubric.
  • Audience personas: Who the reader is and the vocabulary they use. Use the persona to set the register, which is the attribute you will spend the most time correcting.
  • Past copy: A few hundred pieces your best editor already approved, tagged by format (blog intro, support reply, email). You turn these into training pairs and the few-shot sample library.
  • Translation memory: If you publish in more than one language, your TM already holds aligned source-target pairs, which you can convert into supervised examples with light cleanup.

The sixth asset is a frozen evaluation set, which I’d size at roughly 30 to 50 briefs, with human-written references that never enter training. Without it you cannot tell whether a change improved the voice or memorized the test, and you will need it in Step 5 regardless of which method you choose.

How do you codify brand voice into machine-readable rules?

You make voice trainable when you split it into named attributes a check can score: persona, tone, lexicon, reading level, and a do-not-say list. Adjectives like “friendly” and “professional” fail here because no judge, human or model, scores them the same way twice. Each attribute needs a measurable target. If you do not have those named attributes yet, start from an AI brand voice guide.

Here is a filled-in profile for a hypothetical B2B SaaS brand, kept short enough to paste into a system prompt or chunk into a retrieval index:

Illustrative B2B profile: persona is a senior finance-ops operator writing to a peer who runs the same system. Tone is declarative and evidence-first, with at most one casual aside. Reading level is Flesch-Kincaid grade 10 to 12. Required terms include Example Co, the workspace, approval rule, and reviewer. Banned terms include solution, platform, journey, supercharge, and effortless. Structure leads with the number or the mechanism. Do not say we are excited to announce, in today’s, or reach out.

Now take three lines from that brand’s style guide and convert each into a prompt rule for system prompts. Then create a retrieval chunk with metadata for RAG and a training pair for fine-tuning. The same line yields all three.

Convert each style-guide line three ways. Address the operator in second person becomes a prompt rule, a retrieval chunk tagged persona for blog, email, and support, and a one-sentence training completion. Call the product the workspace, never solution or platform, becomes a lexicon rule, a retrieval chunk tagged lexicon for all formats, and a rewrite pair that swaps our platform for the workspace. Put the number before the claim becomes a structure rule, a retrieval chunk for blog and email, and a completion that opens with median approval time dropping from 3 days to 4 hours.

The metadata on the retrieval chunk is the part most teams skip, and it is why their RAG setups pull lexicon rules into a tweet and persona rules into release notes. The applies_to and attribute fields let the retriever fetch persona rules for a blog brief and skip them for UI strings. TUI’s hotel-description project started the same way, with a tone target defined as upbeat and promotional before any model was touched.

How do you choose between prompt, RAG, and fine-tuning?

Use prompts for speed and RAG for facts. Fine-tune when style has to hold at volume. In practice, many teams need the first two before they need the third. A compact AI brand voice prompt is often enough before you add retrieval or training.

  • Use context methods when they fit: A system prompt can hold a compact voice profile and sample set, with edits taking effect on the next call. Product details and policies that change belong in a retrieval layer. Fine-tuning is poor at adding knowledge, and training examples that introduce new facts are learned slowly and then raise hallucination rates linearly as the model absorbs them.
  • Fine-tune when style has to hold at volume: When high-volume generations across many producers all need the same register, the long prompt becomes a per-call tax and human review becomes the bottleneck. Training moves the voice into the weights so the prompt can shrink.

Base-model mechanics: Pre-training builds the base model from a general web-scale corpus, and you will never do it. Fine-tuning continues training on your few hundred to few thousand examples so the model’s default output shifts toward them. Within fine-tuning, LoRA is the affordable default: it freezes the base weights and trains small low-rank matrices alongside them, which on GPT-3 175B cut trainable parameters by 10,000 times and GPU memory by 3 times while matching full fine-tuning quality, and adds no inference latency once merged.

Deployment durability: Where you run it matters more now than it did a year ago. OpenAI closed fine-tuning job creation to new organizations in May 2026 and ends it for all organizations on January 6, 2027, with existing fine-tuned models running only until their base models are deprecated. Durable brand-voice fine-tuning now means an open-weight model you host. You can also use a managed service such as Vertex or Bedrock, where the provider keeps a private tuned copy and controls when its base model retires.

How do you build the training dataset from existing content?

The dataset is JSONL, one prompt-completion pair per line, where the prompt is the brief a teammate would hand a writer and the completion is the approved copy. Keep the prompt realistic, because the model learns the mapping from brief to copy, and a brief nobody would write produces a mapping nobody will use.

Each training line is one brief plus the approved copy. Example: a two-sentence product-update intro for a finance-ops lead, with facts that approvals now route by amount threshold and ship Tuesday, paired with copy that states the $400 versus $40,000 split and the Tuesday rollout. Example: a support reply on why a $12,000 invoice went to a second reviewer, paired with copy that cites the $10,000 two-approval threshold.

On size, OpenAI’s guidance is to start with 50 well-crafted examples and evaluate, with a hard floor of 10, and to rethink the task if 50 move nothing. At the top end, 1,000 curated pairs aligned a 65B model, and growing the set from 2,000 to 32,000 examples did not improve response quality. Between those anchors, a working planning range is 200 to 1,000 pairs, with the lower bound for a single format and the upper bound for a voice that spans blog, email, and support.

Mine the pairs from assets you already have:

  • Published marketing copy: For blog posts, pair the original brief with the published intro or a strong body paragraph. Where no brief survives, generate a flat paraphrase of the paragraph and use that as the prompt, with the original as the completion. For emails, pair the subject line and stated goal with the body of campaigns that cleared review without edits.
  • Support replies: Pair the customer’s message with the reply your best agent sent, after stripping names and account details.

Curate harder than you collect. In the same study, filtering Stack Exchange answers to those scoring 10 or higher, within a set length window, and free of first-person or cross-referencing language scored 3.83 on a 1–6 helpfulness scale against 3.33 for the unfiltered set, so remove off-voice samples, deduplicate near-copies, and hold back the frozen evaluation set from the prerequisites section before training starts.

How do you run the fine-tune?

Run supervised fine-tuning with LoRA first, and add DPO only if your reviewers disagree about tone after SFT. A 7B to 13B open-weight instruct model is a reasonable starting size for voice, since TUI used 13B and the authorship test above used 7B. TUI tuned Llama 2 13B with QLoRA on roughly 4,500 hotel descriptions using one ml.g5.4xlarge instance for about 20 hours, which puts a production-grade run inside a single day of compute.

  • Adapter scope: Attach the adapter to all weight matrices rather than attention layers alone. On Llama-2-7B instruction data, that configuration matched full fine-tuning on the target task while forgetting measurably less of the base model’s general ability, and forgetting less is what you want when the model still has to write coherent English about topics outside your corpus.
  • Training window: Train two to three epochs with the frozen evaluation set held out. Stop when validation loss turns upward while training loss keeps falling, then merge the adapter into the weights before serving.

DPO is the optional second pass. It trains on preference pairs, the same prompt with one preferred and one rejected completion, and OpenAI’s format uses the JSONL fields input, preferred_output, and non_preferred_output, with SFT on the preferred responses run first. The pairs come from your reviewers: every time an editor picks draft A over draft B for the same brief, you have one row.

Treat DPO’s style gains as something to test rather than assume. A KTH thesis that trained headline generation with SFT and DPO on roughly 4,000 Aftenposten titles found the best DPO variant did not perform significantly differently from SFT alone. Run it against your evaluation set, and keep it only if the score moves.

How do you test whether it sounds like your brand?

Score every draft on four attributes, two deterministic and two judged, and set a pass line for each before you look at any output. The deterministic checks run on every generation. The judged checks run on a sample.

  • Forbidden phrases (binary): A regex pass over the do-not-say and banned-lexicon lists. Pass is zero hits, and a single hit fails the draft.
  • Reading level (binary): Flesch-Kincaid grade inside the profile band, grade 10–12 in the example profile. Outside the band fails.
  • Lexicon adherence (percentage): The share of product references that use the canonical term. A sensible starting pass line is 100% for product names and 90% for softer preferred terms.
  • Tone (pairwise): An LLM judge compares the draft against a human-written reference for the same brief and picks the closer match to the voice profile. Pass is a win or a tie in both presentation orders.

Pairwise comparison against a reference is the format I recommend for tone, but LLM judges carry measurable biases even in pairwise mode. On MT-Bench, GPT-4 agreed with human raters 85% of the time, above the 81% human-human rate, but returned the same verdict after swapping answer order only 65% of the time and favored its own outputs by about 10 points of win rate. Randomize order and run both orders. Use a different model family than the one that wrote the draft.

Human check: A blind human panel catches what the judge misses. Mix model drafts with human-written pieces for the same briefs, then strip the labels. Have at least three reviewers rate each against the profile. Record the share of model drafts rated at or above the human draft, and keep that number as your before/after metric across retrains.

Case benchmark: In TUI’s project, under a prompt-engineering baseline, 98% of 150 generated descriptions had significant issues. After the QLoRA fine-tune plus a second-model tone pass, 75% of 50 hotels were rated higher than the human-written version in a blind test, with 5x faster generation. Two caveats hold: the two percentages come from different evaluations, and the final result belongs to the two-model system rather than the adapter alone.

What does it cost, and how long does it take?

For this workflow, budget for dataset curation as well as training compute. Use the table below to compare the three methods on the four factors a budget owner asks about. The setup and refresh entries are working estimates rather than measured benchmarks.

[@portabletext/react] Unknown block type "table", specify a component for it in the `components.types` prop

Retrieval is the latency tax. On a 9B model, adding RAG raised time to first token from 495 ms to 965 ms, with retrieval accounting for about 35% of that total. If your voice rules live in the index alongside product facts, you pay that cost on every brand-voice draft.

Training compute is only part of the budget. At Vertex’s published rate for its Flash-tier model, $5 per million training tokens, a 500-example set averaging 1,000 tokens run for three epochs is 1.5 million tokens, or about $7.50. That is worked arithmetic rather than a provider quote, and it excludes the editor weeks that produce the 500 examples. TUI’s roughly 20 hours on a single AWS instance is the better proxy for an open-weight adapter.

Which tools support custom brand voices?

The four chat and marketing tools below capture voice through instructions and files, and none of their current documentation describes training weights on your examples. The OpenAI API is the exception, and its training path is closing.

  • ChatGPT custom GPTs and Projects: Voice goes in as written instructions plus knowledge files. GPTs do not use saved memory, custom instructions, or previous conversations, so each chat starts fresh, which means you cannot carry a correction into the next conversation unless you edit the GPT itself.
  • Claude Projects: Project instructions and uploaded knowledge apply to every chat inside the project, which gives you a persistent voice workspace inside the project.
  • Jasper Brand Voice: You upload up to 8 example pieces as text, files, or URLs, and Jasper generates a voice description you can edit, toggle, and compare against output without it.
  • Copy.ai: You paste content, run Analyze Brand Voice, and edit the generated profile before applying it in chat and workflows.
  • OpenAI API: It currently offers supervised and preference fine-tuning, which do train weights. With job creation closing (Step 2), you can use it for inference against a system prompt or retrieval layer while running durable training elsewhere.

These tools capture voice through instructions and uploaded files rather than documented weight training. That is adequate for most teams, and you can reuse the profile and samples you built in Step 1 across all of them without rework.

Is it safe to upload proprietary brand assets?

Commercial API and workspace tiers generally provide stronger data controls than free consumer tiers. Answer two readiness questions before any style guide or customer email leaves your drive.

When is fine-tuning the wrong answer?

Two thresholds determine whether fine-tuning creates more work than it removes.

  • Operating threshold: Fine-tuning is wrong when a prompt plus a sample library already passes your rubric, and teams with only a few producers are often in that position. A one-to-two-person team shipping three to five articles a week can usually run on a persistent workspace and a sample library: those files carry consistency across sessions, and the prompt is where they get invoked. Fine-tune when the prompt stops passing the rubric or when review time and per-call prompt tokens cost more than training would remove.
  • Model and safety thresholds: Model turnover has moved against hosted fine-tuning. OpenAI’s wind-down is the clearest case, and Google’s documentation for migrating tuned PaLM models to Gemini describes the same constraint: a tuned model cannot move to a new Gemini version, so every base-model upgrade requires a new tuning job with no reuse of hyperparameters. Budget one full retrain and evaluation cycle per base-model change. If your provider ships a new base every six months, that is two retrains a year for a voice that may not have changed at all. Narrow training also erodes the base model’s safety alignment. Fine-tuning GPT-3.5 Turbo on benign, non-adversarial instruction data raised its harmful-response rate from 5.5% to 31.8%, with nothing harmful in the training set. If the model will face customers, you need a safety regression suite next to the voice rubric, and if you cannot staff that, stay on prompts and RAG.

How do you keep voice from drifting?

Editors create drift when their corrections stay in the document and never flow back into the artifact. Every edit a reviewer makes to a draft is either a new training pair or a new rule, and the owner must add it to the relevant artifact. Assign one named owner, usually the senior editor who already arbitrates style disputes, with the authority to change the profile and the responsibility to trigger a retrain.

Set the refresh cadence against two triggers:

  • Fixed schedule: On a schedule matched to content volume, the owner reruns the frozen evaluation set against the current delivery method and compares scores with the previous run.
  • Model event: After a provider changes the base-model version, the owner triggers a full retrain and evaluation before the new model touches production, since the adapter does not carry over.

Reviewers also change the rubric as they sharpen their criteria while grading. The owner should give the do-not-say list and tone definition a version number and a changelog, then read old scores against the rubric version that produced them. A lower score under a newer rubric may reflect stricter review rather than a weaker model.

How do you bring the same voice rules to the whole team?

The profile and sample library are durable assets, along with the evaluation set. The prompt and retrieval index invoke those assets, as does the adapter. A team that stores its voice inside one person’s system prompt has nothing when that person leaves. Store the profile and samples in a shared, versioned workspace, and feed the same files into every tool a teammate opens, from shared chat workspaces to fine-tuning jobs.

The practice is to keep the voice files out of one person’s chat history. Version the profile and samples, then point every teammate’s workspace at those files instead of retyping rules into prompts.

Put the voice profile from Step 1 under version control this week. Let the retrain decision wait until the frozen evaluation set shows the prompt has stopped passing. For a weekly working example of this kind of operational prompt work, get the newsletter.

Frequently Asked Questions

Related Content