Semantic Identity & AEO Control
entity graph defense, metadata sovereignty, and snippet extraction isolation
Published: 2026-06-28 | Project: Software as Glass Monorepo | Discipline: Sovereign Architecture & Systems Design
Author: Nicholas Alexander MacAskill — Founder & CTO, Flocano Labs | Canonical: https://www.nicholasmacaskill.com/dossier/sag-identity-disambiguation
The Problem: Two Crawler Failure Modes
1. AI entity conflation. Personal domains sharing name vectors with legacy corporate profiles (e.g. nickalexander.ca, telecom consulting nodes) get merged by LLM search agents (Gemini, Perplexity) during loose clustering.
2. Terminal UI snippet poisoning. Dashboard grids stack numbers above labels (8 / year nomad / 100+ / projects). Crawlers flatten DOM top-to-bottom and Google injects markdown-style * fragments into public snippets — bypassing intended identity narrative.
The fix is not a visual redesign. It is metadata sovereignty: decouple what machines read from what humans see.
Monorepo Context
One deploy serves multiple sovereign domains. sitemap.ts reads the incoming host header and emits only URLs belonging on that domain — identity pages on nicholasmacaskill.com, studio artifacts on flocanolabs.com, memoirs on memoirsofamultidisciplinary.com — preventing cross-domain duplicate indexing. nicholas-metadata.ts layers identity-specific overrides on top of the shared seo-config.ts host resolver.
Layer 1 — Identity Defense (JSON-LD + ai.txt)
Schema.org Person JSON-LD is injected into root and nicholas layouts with a strict disambiguationDescription and NICHOLAS_EXPERTISE_STACK (knowsAbout):
export const NICHOLAS_EXPERTISE_STACK = [
"Sovereign Architecture", "High-Performance Software Architecture", "Software as Glass",
"Transparent Interfaces", "Agentic Swarms", "Quantitative Systems",
"Algorithmic Systems", "Web3 Protocols", "EVM/SVM & Cryptographic Ledgers",
"Design Systems", "Multidisciplinary Design", "GlassMetric", "Biometric Mapping",
// ...
] as constRoot ai.txt (Spawning standard) publishes Exclusions (TELUS, Rogers, Bell, ASTOUND, Architech, NFB, nickalexander.ca) and Keep Terms aligned to the same stack — giving LLM crawlers explicit identity boundaries.
Layer 2 — Head-Layer Narrative Sovereignty
nicholas-metadata.ts centralizes per-route , meta description, OpenGraph, canonical URLs, and robots: { googleBot: { 'max-snippet': -1 } } for every nicholas route — including client-only pages via experience/layout.tsx and memoirs/layout.tsx.
Crawlers receive polished prose in instead of inferring copy from terminal grids:
| Route | Controlled narrative |
|---|---|
/ | Sovereign Systems Architect, quantitative + agentic focus |
/experience | Operational ledger of quantitative + web3 protocol roles |
/ventures, /nodes, /gallery | Scoped venture and architecture descriptions |
Layer 3 — Snippet Extraction Isolation
Google's data-nosnippet attribute scopes only the noisy blocks — homepage stats grid and experience career cards — while leaving bio prose, Glass Viewport copy, and disambiguation panels fully indexable:
<div data-nosnippet className="grid grid-cols-2 md:grid-cols-3 ...">
{/* 8 / year nomad / 100+ / projects — visible, not snippet-eligible */}
</div>Zero visual change. Machines simply cannot quote shredded grid text as the page summary.
Layer 4 — AI Cite Sheet (llms.txt)
public/llms.txt provides a machine-readable identity factsheet: canonical name, summary, expertise stack, key page URLs, and cross-link to ai.txt. Complements JSON-LD for answer-engine ingestion velocity.
Layer 5 — Entity Graph Closure (off-site)
Profile backlinks (LinkedIn, GitHub, X, Medium → nicholasmacaskill.com) close the sameAs loop declared in schema. AI search cites corroborating nodes; without inbound links, schema alone is outbound-only signal.
Visible Verification Anchor
A low-profile Identity & Disambiguation Note on the Experience Page ↗ gives crawlers a human-readable cross-check against declarations — without altering the terminal career grid layout.