If you want to know what makes a page rank well for a given keyword, one of the most direct sources of evidence is the pages that are already doing it. Site Optimiser’s Keyword Mode leans on that idea directly: rather than guessing at an ideal structure from first principles, it scrapes the current top-ranking results for your target keyword and rebuilds their heading and paragraph structure into something you can actually look at, select from, and reuse. Here’s what that scraping and extraction step does, and why it’s the foundation the rest of the pipeline depends on.
Starting From What's Actually Working
There’s a meaningful difference between designing content structure from theory and designing it from evidence. Theory says a good page “should probably” cover certain subtopics in a certain order. Evidence says the pages a search engine is currently choosing to rank at the top for this exact keyword happen to cover these specific subtopics, in this order, with these headings. The second is a far stronger signal, because it reflects an actual, current outcome rather than a guess about what an outcome should look like. Search algorithms change constantly and reward different things over time, but at any given moment, the pages ranking well for a keyword are the closest available proxy for “what’s working right now” — better than any static best-practices checklist.
Scraping the top results is how Keyword Mode gets access to that evidence directly, rather than relying on secondhand summaries or generic advice about what a good page “usually” contains.
What Gets Extracted
Once the top-ranking pages for a keyword are pulled, the extraction focuses on two things: the heading hierarchy, from H1 down through H6, and the paragraph structure underneath each heading. The heading hierarchy tells you the skeleton of the page — what topics it addresses and in what order, and how those topics nest under each other (a section with three H3s underneath an H2 tells you something different than one H2 with no sub-headings at all). The paragraph structure tells you how much depth each section gets — a heading followed by one short paragraph signals a passing mention, while a heading followed by several paragraphs signals a section the page treats as genuinely important.
Together, those two layers give a reasonably complete picture of how a page is organized without needing to reproduce its actual sentences. That distinction matters: the goal isn’t to copy specific wording, which would be both an ethical problem and a practically bad idea (search engines penalize duplicate content, and your own page needs its own voice). The goal is to understand the shape a successful page takes — what it covers and how much weight it gives each part — so your own page can be organized with the same evidence-based structure, filled in with your own original content.
Why Multiple Sources, Not Just One
Looking at more than one top-ranking page matters because any single page’s structure could be idiosyncratic — a quirk of that particular writer’s approach rather than something the search engine actually rewards. Pulling structure from several of the current top results lets patterns emerge: if four out of five top-ranking pages all address a particular subtopic, that’s a much stronger signal that the subtopic matters for this keyword than if only one page happens to include it. The extraction step is set up to surface that kind of cross-page pattern, not just hand you one competitor’s outline as though it were gospel.
From Extraction to Selection
Raw extracted structure isn’t meant to be applied automatically and wholesale — that would just trade keyword-density automation for structure-cloning automation, with the same fundamental problem of a tool making decisions on your page’s behalf without your input. Instead, the extracted headings and sections are presented back to you as selectable items: you look at what showed up across the scraped results and decide which parts are actually relevant to bring into your own page, and which aren’t a fit for what you’re trying to say. That selection step — covered in more depth in Site Optimiser Checkbox Selection: Choosing What Structure to Reuse — is what keeps the process a tool you direct rather than a machine that overwrites your judgment.
The Limits of Scraped Structure
It’s worth being clear-eyed about what this step can and can’t tell you. Scraped structure reflects what’s currently ranking, which is a strong signal but not an infallible one — rankings shift, some pages rank well for reasons unrelated to their structure (raw domain authority, backlinks, brand recognition), and copying a structural pattern doesn’t guarantee your page will rank the same way theirs does. What scraping gives you is a well-informed starting point grounded in real, current search results, rather than a structure invented from assumption. It replaces a guess with an observation — it doesn’t replace the judgment needed to decide what to do with that observation, which is exactly why the selection and diff steps that follow matter as much as the scraping step itself.
Where This Fits in the Pipeline
Scraping and structural extraction is the first stage of a three-part process: scrape and extract structure from what’s ranking, select which parts of that structure are relevant to your page, and diff the proposed rewrite against your current page before anything goes live. Each stage exists to keep the next one honest — extraction gives you real evidence instead of a guess, selection gives you control over what gets used, and the diff gives you final visibility before anything changes. Diffing Old vs New Content Before You Rewrite a Live Page and From Scrape to Rewrite to Publish: The Site Optimiser Pipeline cover the rest of that chain, and Site Optimiser’s Keyword Mode: What It Should (and Shouldn’t) Do covers the reasoning for why the pipeline is shaped this way instead of just rewriting pages directly.