Back to Learn
#AEO

How to cluster keywords with ChatGPT for SEO content planning

Learn a repeatable ChatGPT workflow to cluster keywords into semantic groups, label intent, validate against SERPs, and map to content briefs.

Keyword list grouped into named clusters with intent labels for SEO briefs

ChatGPT keyword clustering takes a raw keyword export and returns intent-labeled topic groups in a fixed table: cluster name, member keywords, intent, primary keyword, page type. You supply the list and the metrics, ChatGPT supplies the grouping and labels, and you use a SERP check to decide which clusters become pillar pages, supporting URLs and briefs.

What is keyword clustering, and what does ChatGPT actually do when it groups terms?

Keyword clustering is the step where you decide which keywords one page should target together and which need separate pages. Done well, it produces a topic map you can turn into URLs. That search map is a different artifact from a prompt map for AI search. Done badly, clustering produces two pages competing for the same query or one page trying to answer three.

ChatGPT groups keywords by meaning. When you paste a list, it reads each string, compares what the strings mean to each other inside its representation of language, and puts terms that sit close together in the same bucket. In a plain chat it has no live access to Google results and no search volume data. Its output is a semantic hypothesis about your topic map, and Google may disagree.

Most dedicated clustering tools work the opposite way. Keyword Insights, for example, looks at the top seven ranking URLs for each keyword and groups terms that share 40% or more of those URLs, with a default floor of three shared URLs. That measures how Google treats the queries.

The gap between the two shows up on modifier keywords. “Best running shoes for flat feet” and “best running shoes for plantar fasciitis” mean nearly the same thing to a language model, but live SERPs often return mostly different URLs for each, which means Google treats them as separate intents deserving separate pages. ChatGPT will merge them by default. You split them after validation.

Where do you get your keyword list first?

Export from a tool that carries real metrics, because ChatGPT will not supply them and will fabricate them if asked. These sources cover most content programs:

  • Google Search Console: The Performance report gives you queries with clicks, impressions, CTR and average position, which gives you first-party data. The UI export caps at 1,000 rows, so for larger sites pull through the Search Analytics API or the BigQuery bulk export.
  • Ahrefs and Semrush: Ahrefs Keywords Explorer exports carry volume, difficulty, CPC, traffic potential, parent topic and intent tags. Rows per report run from 2,500 on Lite to 30,000 on Standard and 75,000 on Advanced, and export rows count against a separate monthly quota. Semrush Keyword Magic Tool exports to CSV or XLSX with Intent, Volume, Keyword Difficulty, CPC and SERP Features columns already attached, and can save groups as separate tabs.

If you are starting a new topic from nothing, run an AI keyword research workflow first. ChatGPT can expand a seed term into modifiers, question phrasings and adjacent terms. Treat every term it generates as unverified until a keyword tool returns volume for it. Some will be strings nobody has ever searched. Run the expanded list through Ahrefs or Semrush, drop the zero-volume rows, and cluster what survives.

How do you cluster keywords with ChatGPT step by step?

The procedure has five steps: export a clean keyword list with metrics attached, decide whether you are grouping on meaning or on intent first, prompt for clusters into a fixed table, label intent in a second pass, and validate every cluster against live SERPs and Search Console before anything reaches a brief. In a ChatGPT keyword clustering run, expect validation to take more of your time than prompting.

How do you export or paste your keyword list?

Export CSV from your keyword tool and keep four columns: keyword, volume, difficulty, and the tool’s intent tag if it has one. Everything else (trend, competitive density, SERP features) adds tokens the grouping step does not use. Dedupe, lowercase, strip brand terms into their own list, and put one keyword per line.

Size the paste to the context window you have. A 100-keyword list with volumes fits a typical ChatGPT window. Keep the first pass under 100 rows and batch the rest (batching rules are in the section on lists above 200 keywords).

Paste volume and difficulty alongside each keyword even though the prompt tells ChatGPT to ignore them for grouping. You want them present when you pick a primary keyword per cluster, and you want the model to see that real numbers exist so it has less room to invent its own.

How do you prompt for semantic clusters with a fixed table?

Use a prompt that locks the output schema and forbids the two failure modes that cost the most time: added keywords and dropped keywords. Use this instruction set.

Lock the output schema and forbid added or dropped keywords. Instruct the model to use only the pasted list, group on meaning, put every keyword in exactly one cluster, and return a table with cluster name, keywords, a blank intent column, primary keyword, and page type (pillar, guide, comparison, listicle, landing page, or glossary). Cluster names are two to five unique words. Keywords are copied exactly. End with an input-count and output-count line that must match.

The count line is the part most operators skip, and you can use it to catch clusters that silently lose terms. Search Engine Journal’s test of ChatGPT for keyword work found it will create new keywords when the prompt leaves room for it, including a term (“Interstellar Internet SEO”) that nobody has ever searched. You can also use the count line to identify a keyword that falls out of the table before the missing term reaches a brief. Rule 1 blocks added terms, rule 5 flags omissions, and you still check the count yourself before moving on.

How do you label search intent as a second pass?

Run intent labelling as a separate prompt after the semantic table exists, because each prompt then has one job, which I find easier to check. The grouping pass answers “which keywords belong on one page.” The intent pass answers “what kind of page.”

Use four labels: informational, commercial (some tools call this commercial investigation), transactional, and navigational. Informational means the searcher wants to learn something. Commercial means they are comparing options before buying. Transactional means they are ready to act, whether that is buy, sign up, or start a trial. Navigational means they want a specific site.

Ambiguous clusters are where the second pass earns its keep. A cluster built around “best project management software” reads as commercial, but “project management tools” inside the same cluster could be informational or commercial depending on the SERP. Ask for a primary intent plus an optional secondary, and instruct the model to split any cluster carrying two different primary intents. The intent prompt in the templates section below does this.

How do you flag overlapping clusters to prevent cannibalization?

Ask ChatGPT to audit its own table for merge candidates before you touch it, because a first pass on a mid-sized list often produces clusters whose names differ but whose pages would be the same URL. ChatGPT returns three lists from the cannibalization follow-up prompt (also in the templates section): clusters to merge, clusters to split, and any primary keyword that appears twice.

The label problem gets worse when you batch. Lee Foot’s keyword classifier documentation describes the model calling one topic “Comparison” in one batch and “Comparisons” in the next because each batch runs with no shared taxonomy. It also documents the tool silently skipping a failed batch, so those keywords never appear in the output. Two rules follow. Carry a fixed list of approved cluster names from batch to batch, and reconcile output row count against input every time.

In the sample table below, “Project management software overview” and “Best project management software” share the same head term. ChatGPT should flag that pair as a merge candidate when it runs the cannibalization prompt. Whether you merge them depends on Step 5, not on the model’s opinion.

How do you validate clusters before you brief anything?

Validation is mandatory, and you perform it outside the clustering prompt. Every cluster passes three checks before it becomes a brief.

  • Live SERP check: Google the primary keyword and one or two members of each cluster, in the same country and device you target, and count shared URLs in the top seven results, the window Keyword Insights uses. If a merge candidate pair shares fewer than three of those URLs (Keyword Insights’ default floor), keep them as separate clusters. If two clusters you kept apart share most of those URLs, merge them.
  • Search Console check: Filter the Performance report by each cluster’s member queries and see which page already gets impressions. One page pulling impressions for two clusters is either a merge signal or a cannibalization problem you already have.
  • Keyword tool check: Confirm every keyword in the table exists in Ahrefs or Semrush with volume above zero. Set aside any term that fails this check and review the original export before using it.

The running-shoes pair from the definition section is the model case: ChatGPT merges it, the SERP check splits it, and you write two pages.

Which prompt templates work for semantic and intent-based clustering?

The semantic prompt in Step 2 is the first template. The second groups by intent first, which suits lists where you already know the funnel stage matters more than topical adjacency, such as a comparison-heavy commercial list. Paste it as a fresh prompt with the same keyword list.

For intent-first grouping, instruct the model to assign each pasted keyword exactly one primary intent (informational, commercial, transactional, or navigational), then group within each intent by pages that could rank together. Use the same table columns as Step 2, add a secondary-intent column only where a cluster carries a strong second intent, and require matching input and output counts.

Output from this template still goes through the cannibalization check and Step 5 validation.

Run the intent labelling pass on the Step 2 table with this prompt:

On the existing cluster table, fill the intent column with exactly one primary label from informational, commercial, transactional, or navigational. Split any cluster that mixes two primary intents and rename both. Add a secondary-intent column only where needed. Do not change keyword text. Repeat the count line.

Then run the cannibalization follow-up:

Ask for three lists, not an edited table: merge candidates (same primary keyword, names that differ only in wording, or members one page would answer), split candidates (mixed page types or intents), and any primary keyword that appears in more than one cluster. Wait for confirmation before changing the table.

Here is what the output looks like on an illustrative project-management SaaS list after the intent pass (keywords shown for shape only, no metrics attached because ChatGPT should never supply them):

[@portabletext/react] Unknown block type "table", specify a component for it in the `components.types` prop

These prompts belong in a reusable document, not in your chat history. Save them alongside your approved cluster-name taxonomy and your page-type list so every operator on the team runs the same schema. For a weekly working example of this kind of operational prompt work, get the newsletter.

How do you cluster more than 200 keywords with batching, CSV uploads, and embeddings?

Practitioner guides converge on about 100 keywords per batch in chat, with degradation reported somewhere between 100 and 200, and an ACL 2026 study of 16 models across eight list-processing tasks found noticeable accuracy drops above 200 items and near-collapse beyond 1,000. The model usually fails by omitting terms or producing inconsistent structure, and the output looks fine until you count.

For lists above 200 keywords that you still want to handle in chat, batch with these controls:

  • Number and name each batch: Include the count line from Step 2 and paste the approved cluster names from previous batches into each new prompt.
  • Reconcile once at the end: Run the cannibalization prompt once on the merged table at the end, not per batch.

CSV upload moves the work from the model’s context into code. The feature is currently named Data analysis with ChatGPT, the renamed successor to Advanced Data Analysis and the original Code Interpreter.

You upload the CSV and ask ChatGPT to write and run Python that groups the rows, which sidesteps the context window for the data itself. Spreadsheets cap at roughly 50 MB, the hard limit per file is 512 MB, and Free accounts get three uploads a day. The grouping logic still runs on whatever method ChatGPT writes into the script, so inspect the code and the cluster sizes before trusting the result.

At 5,000 keywords and above, stop asking the language model to cluster and use embeddings. The OpenAI embeddings API turns each keyword into a vector, and keywords with similar meaning produce vectors pointing in similar directions, measured by cosine similarity.

text-embedding-3-small returns 1,536-dimension vectors and accepts up to 8,191 tokens per input. Ten thousand keywords come to roughly 45,000 tokens at typical query lengths, which is a small single API job.

Cluster the vectors with HDBSCAN rather than K-Means. K-Means requires you to name the number of clusters in advance, and you never know that number for a keyword corpus until you have explored it. HDBSCAN finds dense regions on its own and labels stragglers as noise.

HDBSCAN runs DBSCAN across a range of density settings. Plain DBSCAN also works if you tune its eps value, and it labels noise points -1.

An embeddings plus HDBSCAN pipeline on a few thousand Search Console queries typically produces a few hundred content clusters in minutes, with a slice of long-tail terms set aside as noise. Tune a cosine threshold against about 20 queries you already know belong together, then reuse that method on the rest.

Reuse the tuning method and set your own threshold. Thresholds tuned on one embedding model do not carry to another without re-testing.

Bring ChatGPT back at the end for what it does well: naming and labelling each cluster, then proposing merges between clusters the algorithm left adjacent. That is the same schema as Step 2 with the grouping already done.

How does ChatGPT compare with dedicated clustering tools?

No independently audited study compares ChatGPT clusters against a SERP-overlap tool on the same labeled keyword list, so any accuracy number you see comes from a vendor. The only direct comparison is Keyword Insights’ own test, which ran 17 tools on 216 content marketing keywords and scored SERP-based tools 70–95 out of 100 against 33–50 for semantic and AI methods.

Keyword Insights sells SERP-based clustering and did not publish its scoring rubric or ground truth. Treat the direction of the finding as plausible and the magnitude as unverified.

The direction fits what each method measures. On head terms plus close variants (“project management software,” “what is project management software,” “project management software definition”), you would expect ChatGPT and a SERP tool to form the same cluster, because the words and the likely ranking pages both overlap. Question phrasings of one topic should behave the same way.

Expect disagreement on modifier keywords where the words are close and the SERPs may not be, like the running-shoes pair above. ChatGPT merges these on meaning, and a SERP tool keeps them apart when the ranking URLs differ.

Keyword Insights and ClusterAI group on live SERP results per keyword, which is how they catch intent splits ChatGPT misses. The trade is cost and time against a merge error you would otherwise find by hand in Step 5.

My working rule depends on list size and stakes. Under a few hundred keywords feeding a small content program, ChatGPT plus a manual SERP check on every merge candidate can get you to a comparable topic map without a paid clustering tool. Above 1,000 keywords in a niche where cannibalization costs rankings you already hold, a SERP-overlap tool is likely worth it for the validation hours it removes, and you still run the intent and page-type labelling through ChatGPT because labelling finished clusters is the job the embeddings section already hands back to ChatGPT.

Do you still need a keyword tool for volume and difficulty?

Yes, because ChatGPT has no search data and every volume figure it produces is a guess dressed as a number. The division of labor is fixed: the keyword tool supplies the list and the metrics, ChatGPT supplies grouping and labels, Search Console and the live SERP supply the verdict.

Your Ahrefs or Semrush export already carries volume, keyword difficulty, CPC and an intent tag on every row. Keep those columns through the whole workflow and use them for the decisions ChatGPT cannot make: which cluster to brief first (summed volume), which primary keyword to pick when two members are equally broad (the higher-volume one), and which clusters to drop as unwinnable (difficulty against your domain’s history). Volume estimates differ between tools, sometimes by a wide margin, so Search Console impressions remain the only figure that describes your site rather than a model of Google.

Search Console added its own AI-grouped query view in late 2025, which groups similar queries into topics. It does not replace clustering across a full keyword universe, because it only covers queries you already appear for, but it is a fast way to check whether ChatGPT’s clusters match how Google already groups your impressions.

How do you turn clusters into pillar pages, URLs, and briefs?

The clustered table converts into a URL taxonomy by reading the page-type column top down: one pillar per informational head term, one supporting URL per commercial or guide cluster, and transactional clusters routed to product pages rather than the blog. Take the illustrative project-management table above as the input.

After Steps 2 through 5, the illustrative list collapsed to the five clusters shown above. The SERP check in Step 5 is what decides whether “Project management software overview” and “Best project management software” stay separate. Confirm it in your own SERP check. If the head term returns definitional and vendor pages while the “best” query returns listicles, keep them apart.

That table maps to URLs like this:

  • Pillar: /project-management-software/ targets the overview cluster and links down to every supporting page.
  • Commercial listicle: /best-project-management-software/ targets the “best” cluster and takes the comparison keywords as H2 sections.
  • Commercial listicle: /free-project-management-software/ stands alone because “free” changes the SERP, not because the words differ.
  • Guide: /project-management-software/small-business/ sits under the pillar as a segment page.
  • Landing page: the trial cluster routes to the product’s trial page, outside the blog architecture.

Each cluster row becomes a brief with three fields filled before a writer sees it: the primary keyword as the H1 target, the member keywords as H2 candidates in volume order, and the intent label as the format instruction (a Commercial listicle gets a comparison table and a verdict, an Informational pillar gets definitions and links out). Ask ChatGPT to draft the outline from the row, then check it against the top three ranking pages for the primary keyword so the outline covers what already ranks.

What are the limits, and when should you not use ChatGPT for this?

ChatGPT will invent search volume if you let it, and the invented figure can be off by an order of magnitude in either direction. SpyFu’s test asked for the volume of “custom socks with logo” and got 5,000–10,000 monthly searches back where SpyFu and Ahrefs both showed 700–800. The same test surfaced suggested keywords containing parentheses, a format that does not exist in real search behavior. The validation catch in both cases is Step 5’s keyword tool check: the volume disagreed with two tools, and the parenthetical terms use a format real searches do not take. If a number appears in your table and did not come from your export, delete it.

Fabricated keywords are the second trust boundary. The invented “Interstellar Internet SEO” term from Search Engine Journal’s test is the pattern to watch for: plausible phrasing, zero searches. The count line and the “use only the keywords I paste” rule reduce this in clustering mode, and the keyword tool check catches whatever slips through.

Inconsistency is the third. Run the same list twice with the same prompt and you get two different partitions. Seerly’s engineering write-up on 5,874 skincare keywords found that an LLM-only pipeline fragmented the list into 1,189 groups, a cleanup pass collapsed those to 14, and the silhouette score of 0.107 meant the groups barely differed from each other. Sorting the input differently produced a different partition. Their fix was graph clustering on embeddings (322 communities at modularity 0.593) with the language model kept for labelling, which is the same split described in the embeddings section above.

Token limits round out the list. The context windows in Step 1 cap how much fits in one turn, and the accuracy degradation above 200 items arrives well before the hard limit does.

Do not use ChatGPT for this when any of four conditions hold:

  • You need volume or difficulty and have no tool export to supply it.
  • Your list runs past a few thousand terms and you are not willing to move to embeddings.
  • You are clustering in a niche where losing a held ranking to cannibalization costs revenue and you will not run the SERP check.
  • Your team needs the same list to cluster the same way next quarter, because a chat model will not give you that without a fixed taxonomy carried in from outside.

Do you need AIPRM or a custom GPT for keyword clustering?

  1. AIPRM is a browser extension that overlays a library of saved prompts on ChatGPT and other chat models, and a custom GPT is a saved system prompt with optional attached files. Both are prompt storage. Neither adds SERP data, volume data, or a clustering algorithm, so the output is the same semantic grouping you get from running the Step 2 instructions yourself.

The library’s own contents show the risk. AIPRM’s SEO research prompts include community templates that return a table with Volume, KD, CPC and Competitive Density columns, which is a template for generating the exact metrics ChatGPT cannot know. A prompt that looks authoritative because thousands of people ran it is still asking the model to guess. The clustering prompts in the same library are closer to useful, but you have no control over whether they forbid added keywords or require a count reconciliation.

A custom GPT is the better of the two if your team wants a shared entry point, because you control the instructions and can bake in the schema, the banned behaviors, and your approved cluster taxonomy. Each GPT can attach a small set of files, enough for a taxonomy document and a page-type definition file. The durable asset is still the prompt document and the taxonomy you maintain outside the chat. The GPT is one place to invoke them.

Frequently Asked Questions

Related Content