DEV Community

Fernando Paladini
Fernando Paladini

Posted on

Build a Citation-Ready Static Site with GEO Basics

Build a Citation-Ready Static Site with GEO Basics

TL;DR

Static sites are easy to deploy, but easy to leave ambiguous. A page can look correct in a browser while missing a canonical URL, language alternates, valid structured data, a sitemap, or a clear crawler policy.

This tutorial uses GEO Basics, an MIT-licensed open-source static guide, as a small working example. You will clone the repository, serve it without a framework, run its validator, and inspect the files that make the site understandable to both people and machines.

The goal is not to manufacture rankings or force an AI system to cite a page. The goal is a crawlable, readable, source-backed site with checks that catch common publishing mistakes before deployment.

What you will build

You will run a static site locally and verify that it has:

  • human-readable HTML with a clear page title and headings;
  • canonical and hreflang links for English and Brazilian Portuguese;
  • JSON-LD that matches visible TechArticle and FAQPage content;
  • a root robots.txt, sitemap.xml, and optional llms.txt guide;
  • a repeatable Node.js validation command.

GEO Basics has no framework, build step, database, or runtime server in its documented workflow. The current repository also includes committed Mermaid sources and rendered SVG diagrams. There is no tagged release, so the commands below use the current main checkout. At the time of writing, the verified commit is acd0b37f041db701aa8f7edcd3e280ee4662f7db.

Prerequisites

You need:

  • Git;
  • Node.js with a current node executable;
  • a browser;
  • a terminal with network access for the clone.

The validator uses only Node.js built-ins. You do not need to install an npm package for this repository.

Clone and inspect the site

Clone the public repository and enter it:

git clone https://github.com/paladini/generative-engine-optimization-basic-guide.git
cd generative-engine-optimization-basic-guide
Enter fullscreen mode Exit fullscreen mode

The repository's top-level files are intentionally understandable. index.html is the canonical English page. The Portuguese page lives at lang/pt-br/index.html. robots.txt points crawlers to the sitemap, and sitemap.xml lists both language URLs and their alternates.

The site is served as files. Start a local server from the repository root:

python -m http.server 8123 --bind 127.0.0.1
Enter fullscreen mode Exit fullscreen mode

Open http://127.0.0.1:8123/ in a browser. The Python command is only a local smoke test. It is not part of the project's validator and does not deploy anything.

Understand the page signals

Open index.html and look at the <head> before changing the body. The page declares lang="en", a viewport, a descriptive <title>, a meta description, robots instructions, and a canonical URL:

<link rel="canonical" href="https://paladini.github.io/generative-engine-optimization-basic-guide/">
<link rel="alternate" hreflang="en" href="https://paladini.github.io/generative-engine-optimization-basic-guide/">
<link rel="alternate" hreflang="pt-BR" href="https://paladini.github.io/generative-engine-optimization-basic-guide/lang/pt-br/">
Enter fullscreen mode Exit fullscreen mode

The canonical URL answers which URL represents the English page. The alternate links connect the language variants. They do not translate content, improve a page by themselves, or guarantee that a search engine will display a particular version. They reduce ambiguity when the same guide exists in more than one language.

The page also includes JSON-LD with a TechArticle and an FAQPage. The important rule is consistency: structured data should describe content that is visible on the page. If you add an FAQ only to JSON-LD but not to the rendered HTML, the metadata describes something a reader cannot see.

Google's current guidance says that normal SEO foundations remain relevant for its generative AI features. It also says that structured data is not required for generative AI search and that there is no special markup that guarantees visibility. Treat JSON-LD as a useful description of an already-good page, not as a shortcut.

Add crawler and agent context carefully

The repository contains a simple robots.txt:

User-agent: *
Allow: /

Sitemap: https://paladini.github.io/generative-engine-optimization-basic-guide/sitemap.xml
Enter fullscreen mode Exit fullscreen mode

This permits normal crawling and advertises the sitemap. Crawler policy is an access decision, not a ranking trick. If you change it, check the rules for the user agents you actually intend to control.

The site also includes llms.txt, a human-readable Markdown map of the guide, repository, contributing instructions, sitemap, and primary references. The file is useful as a concise orientation document for tools that choose to read it. It is not a replacement for HTML, robots.txt, or sitemap.xml.

That distinction matters because the current Google documentation explicitly says Google Search ignores llms.txt for Search and that creating one neither helps nor harms Google rankings. The /llms.txt proposal itself describes it as a convention for giving agents a concise guide to important resources. Use it as an optional documentation surface, not as an access-control file or a promise of AI citations.

For OpenAI crawlers, the controls are separate. OpenAI documents OAI-SearchBot for search and GPTBot for potential training use. A site owner can make an independent policy decision for each user agent in robots.txt. Do not infer that allowing one means allowing the other.

Run the project's validator

From the repository root, run the exact documented command:

node scripts/validate-site.mjs
Enter fullscreen mode Exit fullscreen mode

The current script checks that the required HTML, CSS, JavaScript, image, diagram, crawler, sitemap, and llms.txt files exist. It checks JavaScript syntax, expected PNG dimensions, internal anchors, external-link safety attributes, canonical and hreflang links, JSON-LD parsing, required schema types, and the canonical GitHub Pages URL.

The expected result for the current checkout is:

Static validation passed.
Enter fullscreen mode Exit fullscreen mode

During copy edits, the repository also documents a faster command:

node scripts/validate-site.mjs --quick
Enter fullscreen mode Exit fullscreen mode

The quick mode skips some diagram, image-dimension, sitemap, and robots checks. Use it for fast feedback while editing text, then run the full command before opening a pull request or publishing.

Verify a useful failure mode

A validator is more valuable when it can fail for a meaningful reason. In a temporary copy, remove a required file such as llms.txt and run the full command again. The script should report a missing required file and return a non-zero exit code. Restore the file before continuing.

Do this experiment in a disposable copy or with version control ready to restore the file. The repository's contributor guide says that changes should stay focused and that metadata, canonical URLs, language alternates, structured data, accessibility, and validation should be checked together.

The validator is deliberately local. It does not prove that GitHub Pages has deployed the latest commit, that a search engine indexed a page, or that an AI answer system will use it. Those are separate operational and external states.

Why this structure works

The site makes the primary content visible in ordinary HTML, uses descriptive section headings, links to original sources, and keeps the English and Portuguese pages connected. These choices help readers first and provide clearer inputs for crawlers and downstream tools.

The repository also keeps its validation close to the content. That is a practical design decision for a static site: a contributor can run one command without learning a framework-specific build pipeline. The checks are not a substitute for editorial review, accessibility testing, link review, or a real deployment check, but they prevent several easy-to-miss regressions.

Failure modes and security boundaries

Do not treat passing validation as proof of search performance. Google says that meeting technical requirements and best practices does not guarantee crawling, indexing, or serving. Avoid services or tools that promise an internal ranking score or guaranteed AI placement.

Do not copy a page's visible claims into JSON-LD without checking that the claims remain visible and accurate. Incorrect structured data can make a page less trustworthy and can fail eligibility for supported rich-result features.

Do not place secrets in a static repository. Static HTML, JavaScript, robots.txt, sitemaps, and llms.txt are public once deployed. Review source links, email addresses, generated assets, and build artifacts before publishing.

Finally, remember that llms.txt is not a firewall. Use hosting permissions, authentication, and robots.txt policies for the boundaries they actually support. A public page should be written as public content.

FAQ

Does GEO Basics require a framework?

No. The documented path is a static site served directly from files. The repository's validator is a Node.js script, not a framework build.

Does JSON-LD guarantee an AI citation?

No. It can describe visible page meaning, but no markup guarantees a citation, ranking, inclusion, or traffic.

Is llms.txt required for Google Search?

No. Google's current documentation says Google Search ignores it. It can still be maintained as an optional guide for other tools and readers.

What should I verify after deployment?

Check the deployed URLs over HTTPS, inspect the rendered HTML, fetch robots.txt and sitemap.xml, confirm the canonical host, and run the same content checks against the deployed site where possible. Then use the relevant webmaster tools to observe indexing and performance rather than assuming them.

Takeaway

GEO Basics is a compact example of a static site that treats discoverability as documentation quality: readable HTML, explicit metadata, language relationships, source links, crawler files, and a local validator. Start with the page a person should trust, add machine-readable context that matches it, and keep every guarantee out of the implementation unless you can prove it.

If you maintain a static site, which check would catch the most expensive publishing mistake in your workflow: canonical URLs, language alternates, structured data, crawler policy, or deployment verification?

AI assistance disclosure: AI assistance was used to organize this tutorial and review its wording. The repository state, validator behavior, commands, current files, license, and platform guidance were checked against the linked primary sources and the current public checkout.

Top comments (0)