
AEO vs. SEO: What's the Difference? (And Do You Need Both?)
AEO gets your brand cited inside AI-generated answers. SEO gets it ranked in a list. The gap between those two outcomes is widening — and most brands are only doing one of them.
Learn to codify brand voice, choose between prompt, RAG, or fine-tuning, build datasets, and evaluate outputs with a numeric rubric.

Most teams trying to train an LLM on brand voice start with fine-tuning, even though a structured prompt or a retrieval layer often holds the voice first. Codify named attributes, pick the cheapest method that still scores, mine approved copy into pairs, then train and retrain when the base model changes.
Codify your voice into named attributes, pick the method the voice problem calls for (prompt or RAG for most teams, LoRA fine-tuning for style at volume), mine existing copy into prompt-completion pairs, train a LoRA adapter on an open-weight model you control, score outputs against a numeric rubric, and assign an owner who retrains when the base model changes.
The voice itself lives in the structured profile and sample library, including the training pairs you build once. A system prompt and a retrieval index are two delivery points for the same artifact. A LoRA adapter is another, which is why the first step below is identical no matter which method you end up running.
An instruct model’s default output carries little of any one author’s signal. On a ten-author test with Llama-2-7B, a classifier identified the target author from the instruct model’s output with 0.263 accuracy, versus 0.693 with five-shot prompting and 0.879 after LoRA fine-tuning.
The failure modes you see in drafts follow from that default.
Bland tone: Preference tuning tends to favor answers that many raters accept, so the default register reads as competent and interchangeable. Nothing in it sounds like a person who has run your product.
Wrong terminology: The model picks the most common synonym in its training distribution. Your product becomes a “platform” or a “solution” no matter what your glossary says, because generic terms are far more common in public text than your product’s name.
Off-persona register: Without a persona, the model explains teacher-to-student, a common pattern in instruction-style answers. A brand that writes operator-to-operator comes back reading like a textbook.
These failure modes are stylistic rather than factual, and style is the layer that the studies cited below shifted with relatively small datasets. That is the justification for customizing at all, and it is also why the cheapest method often works.
Gather five source assets before you touch a prompt, because every method downstream is a different packaging of the same material. If an asset is missing, you will discover the gap in the evaluation step, which is the most expensive place to find it.
The sixth asset is a frozen evaluation set, which I’d size at roughly 30 to 50 briefs, with human-written references that never enter training. Without it you cannot tell whether a change improved the voice or memorized the test, and you will need it in Step 5 regardless of which method you choose.
You make voice trainable when you split it into named attributes a check can score: persona, tone, lexicon, reading level, and a do-not-say list. Adjectives like “friendly” and “professional” fail here because no judge, human or model, scores them the same way twice. Each attribute needs a measurable target. If you do not have those named attributes yet, start from an AI brand voice guide.
Here is a filled-in profile for a hypothetical B2B SaaS brand, kept short enough to paste into a system prompt or chunk into a retrieval index:
Illustrative B2B profile: persona is a senior finance-ops operator writing to a peer who runs the same system. Tone is declarative and evidence-first, with at most one casual aside. Reading level is Flesch-Kincaid grade 10 to 12. Required terms include Example Co, the workspace, approval rule, and reviewer. Banned terms include solution, platform, journey, supercharge, and effortless. Structure leads with the number or the mechanism. Do not say we are excited to announce, in today’s, or reach out.
Now take three lines from that brand’s style guide and convert each into a prompt rule for system prompts. Then create a retrieval chunk with metadata for RAG and a training pair for fine-tuning. The same line yields all three.
Convert each style-guide line three ways. Address the operator in second person becomes a prompt rule, a retrieval chunk tagged persona for blog, email, and support, and a one-sentence training completion. Call the product the workspace, never solution or platform, becomes a lexicon rule, a retrieval chunk tagged lexicon for all formats, and a rewrite pair that swaps our platform for the workspace. Put the number before the claim becomes a structure rule, a retrieval chunk for blog and email, and a completion that opens with median approval time dropping from 3 days to 4 hours.
The metadata on the retrieval chunk is the part most teams skip, and it is why their RAG setups pull lexicon rules into a tweet and persona rules into release notes. The applies_to and attribute fields let the retriever fetch persona rules for a blog brief and skip them for UI strings. TUI’s hotel-description project started the same way, with a tone target defined as upbeat and promotional before any model was touched.
Use prompts for speed and RAG for facts. Fine-tune when style has to hold at volume. In practice, many teams need the first two before they need the third. A compact AI brand voice prompt is often enough before you add retrieval or training.
Base-model mechanics: Pre-training builds the base model from a general web-scale corpus, and you will never do it. Fine-tuning continues training on your few hundred to few thousand examples so the model’s default output shifts toward them. Within fine-tuning, LoRA is the affordable default: it freezes the base weights and trains small low-rank matrices alongside them, which on GPT-3 175B cut trainable parameters by 10,000 times and GPU memory by 3 times while matching full fine-tuning quality, and adds no inference latency once merged.
Deployment durability: Where you run it matters more now than it did a year ago. OpenAI closed fine-tuning job creation to new organizations in May 2026 and ends it for all organizations on January 6, 2027, with existing fine-tuned models running only until their base models are deprecated. Durable brand-voice fine-tuning now means an open-weight model you host. You can also use a managed service such as Vertex or Bedrock, where the provider keeps a private tuned copy and controls when its base model retires.
The dataset is JSONL, one prompt-completion pair per line, where the prompt is the brief a teammate would hand a writer and the completion is the approved copy. Keep the prompt realistic, because the model learns the mapping from brief to copy, and a brief nobody would write produces a mapping nobody will use.
Each training line is one brief plus the approved copy. Example: a two-sentence product-update intro for a finance-ops lead, with facts that approvals now route by amount threshold and ship Tuesday, paired with copy that states the $400 versus $40,000 split and the Tuesday rollout. Example: a support reply on why a $12,000 invoice went to a second reviewer, paired with copy that cites the $10,000 two-approval threshold.
On size, OpenAI’s guidance is to start with 50 well-crafted examples and evaluate, with a hard floor of 10, and to rethink the task if 50 move nothing. At the top end, 1,000 curated pairs aligned a 65B model, and growing the set from 2,000 to 32,000 examples did not improve response quality. Between those anchors, a working planning range is 200 to 1,000 pairs, with the lower bound for a single format and the upper bound for a voice that spans blog, email, and support.
Mine the pairs from assets you already have:
Curate harder than you collect. In the same study, filtering Stack Exchange answers to those scoring 10 or higher, within a set length window, and free of first-person or cross-referencing language scored 3.83 on a 1–6 helpfulness scale against 3.33 for the unfiltered set, so remove off-voice samples, deduplicate near-copies, and hold back the frozen evaluation set from the prerequisites section before training starts.
Run supervised fine-tuning with LoRA first, and add DPO only if your reviewers disagree about tone after SFT. A 7B to 13B open-weight instruct model is a reasonable starting size for voice, since TUI used 13B and the authorship test above used 7B. TUI tuned Llama 2 13B with QLoRA on roughly 4,500 hotel descriptions using one ml.g5.4xlarge instance for about 20 hours, which puts a production-grade run inside a single day of compute.
DPO is the optional second pass. It trains on preference pairs, the same prompt with one preferred and one rejected completion, and OpenAI’s format uses the JSONL fields input, preferred_output, and non_preferred_output, with SFT on the preferred responses run first. The pairs come from your reviewers: every time an editor picks draft A over draft B for the same brief, you have one row.
Treat DPO’s style gains as something to test rather than assume. A KTH thesis that trained headline generation with SFT and DPO on roughly 4,000 Aftenposten titles found the best DPO variant did not perform significantly differently from SFT alone. Run it against your evaluation set, and keep it only if the score moves.
Score every draft on four attributes, two deterministic and two judged, and set a pass line for each before you look at any output. The deterministic checks run on every generation. The judged checks run on a sample.
Pairwise comparison against a reference is the format I recommend for tone, but LLM judges carry measurable biases even in pairwise mode. On MT-Bench, GPT-4 agreed with human raters 85% of the time, above the 81% human-human rate, but returned the same verdict after swapping answer order only 65% of the time and favored its own outputs by about 10 points of win rate. Randomize order and run both orders. Use a different model family than the one that wrote the draft.
Human check: A blind human panel catches what the judge misses. Mix model drafts with human-written pieces for the same briefs, then strip the labels. Have at least three reviewers rate each against the profile. Record the share of model drafts rated at or above the human draft, and keep that number as your before/after metric across retrains.
Case benchmark: In TUI’s project, under a prompt-engineering baseline, 98% of 150 generated descriptions had significant issues. After the QLoRA fine-tune plus a second-model tone pass, 75% of 50 hotels were rated higher than the human-written version in a blind test, with 5x faster generation. Two caveats hold: the two percentages come from different evaluations, and the final result belongs to the two-model system rather than the adapter alone.
For this workflow, budget for dataset curation as well as training compute. Use the table below to compare the three methods on the four factors a budget owner asks about. The setup and refresh entries are working estimates rather than measured benchmarks.
Retrieval is the latency tax. On a 9B model, adding RAG raised time to first token from 495 ms to 965 ms, with retrieval accounting for about 35% of that total. If your voice rules live in the index alongside product facts, you pay that cost on every brand-voice draft.
Training compute is only part of the budget. At Vertex’s published rate for its Flash-tier model, $5 per million training tokens, a 500-example set averaging 1,000 tokens run for three epochs is 1.5 million tokens, or about $7.50. That is worked arithmetic rather than a provider quote, and it excludes the editor weeks that produce the 500 examples. TUI’s roughly 20 hours on a single AWS instance is the better proxy for an open-weight adapter.
The four chat and marketing tools below capture voice through instructions and files, and none of their current documentation describes training weights on your examples. The OpenAI API is the exception, and its training path is closing.
These tools capture voice through instructions and uploaded files rather than documented weight training. That is adequate for most teams, and you can reuse the profile and samples you built in Step 1 across all of them without rework.
Commercial API and workspace tiers generally provide stronger data controls than free consumer tiers. Answer two readiness questions before any style guide or customer email leaves your drive.
Two thresholds determine whether fine-tuning creates more work than it removes.
Editors create drift when their corrections stay in the document and never flow back into the artifact. Every edit a reviewer makes to a draft is either a new training pair or a new rule, and the owner must add it to the relevant artifact. Assign one named owner, usually the senior editor who already arbitrates style disputes, with the authority to change the profile and the responsibility to trigger a retrain.
Set the refresh cadence against two triggers:
Reviewers also change the rubric as they sharpen their criteria while grading. The owner should give the do-not-say list and tone definition a version number and a changelog, then read old scores against the rubric version that produced them. A lower score under a newer rubric may reflect stricter review rather than a weaker model.
The profile and sample library are durable assets, along with the evaluation set. The prompt and retrieval index invoke those assets, as does the adapter. A team that stores its voice inside one person’s system prompt has nothing when that person leaves. Store the profile and samples in a shared, versioned workspace, and feed the same files into every tool a teammate opens, from shared chat workspaces to fine-tuning jobs.
The practice is to keep the voice files out of one person’s chat history. Version the profile and samples, then point every teammate’s workspace at those files instead of retyping rules into prompts.
Put the voice profile from Step 1 under version control this week. Let the retrain decision wait until the frozen evaluation set shows the prompt has stopped passing. For a weekly working example of this kind of operational prompt work, get the newsletter.
Every week, we share real examples and systems the fastest-growing companies are using to scale smarter.
Get the last workshop recording when you sign up.

AEO gets your brand cited inside AI-generated answers. SEO gets it ranked in a list. The gap between those two outcomes is widening — and most brands are only doing one of them.

ChatGPT prompts for marketing that stay specific. Fifteen frameworks for personas, positioning, content, outreach, and experiments.

Context artifacts are reusable documents that give AI everything it needs to produce consistent, on-brand output — every time you start a new session. Here's the four-artifact system that separates production-grade AI content from generic output.