The Integrity Graph: Your AI Audit’s Missing Layer

▼ Summary
– Common Crawl’s AI Visibility Audit helps organizations check if AI systems can discover and access their content, but visibility alone does not ensure understanding.
– Many websites have schema markup for individual entities (like branches or products) but lack the connective relationships needed to form a coherent knowledge graph of the business.
– Page-level validation tools can conflict with graph-based schema architectures, as referencing entities across pages may generate warnings despite being recommended by Google.
– Google’s recent investments, such as Product Graph and Conversational Attributes, show that even advanced AI struggles to infer relationships, making explicit organizational context increasingly important.
– The next competitive advantage will come from providing clear, connected, and contextually accurate entity relationships (an “Integrity Graph”) rather than just more schema or AI-ready endpoints.
The recent launch of Common Crawl’s AI Visibility Audit marks a logical step forward for organizations eager to know if their content is discoverable by AI systems. After all, if a machine cannot find your information, it cannot summarize, cite, or act upon it. For years, this principle defined search engine optimization: if Google couldn’t crawl it, it couldn’t rank it. The same logic now applies to the AI ecosystem.
Yet while reading the announcement, a deeper concern surfaced. Common Crawl, a massive open repository of web crawl data, serves as a vital proxy for understanding machine accessibility. But the audit focuses on a single question: Can machines find the content? That is a necessary starting point, but it is not the finish line.
What happens after discovery? That question became urgent while reviewing schema implementations across several banking websites. At first glance, the markup appeared mature. Banks had Organization schema, BankOrCreditUnion entities, product and service markup, and branch details. Everything you would expect from large financial institutions was present.
But when I shifted focus from individual pages to the relationships between entities, a gap emerged. Most banks had solid schema, but almost none had built a knowledge graph.
The Difference Between Describing a Page and Describing a Business
A common habit in SEO is auditing schema for completeness. We check required properties, validate against Google’s tools, and look for missing fields. The problem is that these exercises treat pages as isolated units. A branch page is reviewed as a branch. A product page is reviewed as a product. A service page is reviewed as a service. What gets overlooked is whether those entities are meaningfully connected.
In the banking examples, I found branch locations, checking accounts, mortgage offerings, and corporate organizations all marked up separately. What was missing was the connective tissue , the relationships that explain how the business actually operates.
Which legal entity owns the consumer-facing brand? Which products are offered through which services? Which services are available at which branches? Which offerings are limited to specific markets? Which products belong to a larger family of financial solutions?
The markup described the pieces, but it rarely described the business itself. That distinction is subtle, but it becomes critical as search engines and AI systems move from page-level to entity-level understanding.
The Validator Problem
Part of the issue lies in how we evaluate structured data. Most validation tools perform a single-page review. They check whether a page contains the expected properties for a given schema type and whether those properties meet standards. This works well for generating a rich result or validating a standalone entity. It fails when the goal is building a connected knowledge graph.
A frustrating paradox emerges when organizations implement graph-based architectures as Google recommends. A branch page may reference its parent organization through an `@id` relationship pointing to the organization’s primary entity definition on the homepage. The organization’s address, legal information, and social profiles live in the graph, not on the page being tested. Ironically, some of the implementations Google recommends for entity alignment generate warnings in page-level testing tools because the information is intentionally referenced elsewhere rather than duplicated.
Organizations are encouraged to build graphs while still being evaluated as though every page were an island. That mattered little during the rich snippet era. It matters immensely now.
Google’s Evolution Reveals the Real Direction
Many of Google’s recent investments focus on relationships and context. Product Graph, Merchant Center feeds, compatibility data, variant relationships, entity reconciliation, and Conversational Attributes all point in the same direction. These initiatives suggest that understanding how entities relate has become increasingly important, especially when those relationships are hard to infer from content alone.
Google’s actions imply that relationship inference remains challenging, even for one of the world’s most sophisticated information retrieval systems. Otherwise, there would be little reason to keep expanding mechanisms for organizations to explicitly provide contextual information about products, services, brands, and audiences.
Common Crawl Measures Visibility. Relationships Determine Understanding
The AI Visibility Audit addresses a real problem. Organizations should absolutely know whether AI systems can access their content. Visibility matters. But visibility and understanding are not the same thing. Common Crawl is asking the same question SEO teams have asked for decades: Can machines reach the content?
The emerging AI challenge is what happens after machines gain access. A crawler can discover every page on a website and still struggle to understand how the underlying entities connect. Historically, search engines tried to infer those relationships from content, links, and user behavior. They often became remarkably good at it. Yet Google’s recent investments suggest inference has limits.
Consider Conversational Attributes in Merchant Center. Rather than relying solely on AI to determine which products solve similar problems or which attributes matter in specific situations, Google is increasingly asking merchants to provide that context directly. Google has the resources, data, and AI capabilities to make educated guesses. Yet it continues to seek information from the organizations that manufacture, sell, and support those products.
The reason is simple: inference can be powerful, but first-party knowledge is often more accurate. A manufacturer knows which products are compatible. A retailer knows which products are commonly purchased together. A bank knows which services are available at which branches. A global company knows which product variations apply in specific markets.
The question is not whether AI can infer relationships. The more important question is whether the organizations that own those relationships will provide a reliable way for machines to understand them.
Are We Ready for the Agentic Hype Machine?
Over the past year, the industry has focused heavily on concepts like MCP, WebMCP, agent skills, agent cards, API catalogs, A2A protocols, and llms.txt files. Much of the discussion assumes the web is rapidly evolving toward an agent-first ecosystem.
Recent Agentic Readiness research by Bastian Grimm offers a useful reality check. After benchmarking highly visible websites across the United States, the United Kingdom, and Germany, he found that adoption of these agent-oriented standards remains remarkably limited. The overwhelming majority of sites exposed none of the agent-discovery mechanisms currently promoted.
That finding does not suggest the agent-ready web is unimportant. It suggests we may be getting ahead of ourselves. More importantly, even if every major website deployed llms.txt, WebMCP manifests, and API catalogs tomorrow, the same underlying challenge would remain: What information are those systems exposing?
A machine-readable doorway is valuable only if it leads to accurate, connected, and contextually complete information. If the underlying relationships between products, brands, locations, services, and markets are poorly modeled, agentic access simply makes incomplete information easier to retrieve. The access layer is not the hard part. The relationship layer is.
Beyond Entity Graphs: Introducing the Integrity Graph
Most discussions around structured data focus on building an Entity Graph to help machines understand the company, product, location, and how they connect. Those capabilities are important. However, AI systems face a more difficult challenge. They must determine which facts apply within which contexts. This is where organizations need to begin thinking about what I call an Integrity Graph.
An Integrity Graph extends beyond entity identification to preserve contextual truth. It helps establish which legal entity owns a brand, which products belong to a product family, which services are available in specific markets, which branches offer particular services, which regulations apply in particular jurisdictions, and which information is globally applicable versus locally relevant.
Simply identifying entities is no longer enough. Organizations must preserve the integrity of their relationships.
What Organizations Should Audit Next
The growing number of AI readiness audits highlights how quickly the conversation is evolving. Common Crawl’s AI Visibility Audit focuses on discoverability and accessibility. Bastian Grimm’s benchmark for agent-ready technologies assesses whether websites provide machine-readable interfaces. Dixon Jones and the team at Waikay approach the challenge from another angle with their Brand AI Visibility Audit, evaluating whether AI systems can recognize brands, understand entities, and accurately associate an organization with the topics and products it seeks to own.
Viewed collectively, these frameworks reveal that the industry is evaluating several distinct layers of machine understanding:
- Common Crawl focuses on visibility and accessibility: Can machines discover and access the content?Each layer builds upon the one before it. Content must be discoverable before it can be accessed. It must be accessible before it can be associated with an entity. It must be associated with an entity before machines can accurately understand the relationships that give the information meaning.
Why This Matters for Global Organizations
The importance of relationship integrity becomes even more obvious through an international lens. A multinational company may have content available in twenty markets. Common Crawl can successfully discover all of it. AI systems can retrieve it. Search engines can index it. The visibility problem is solved.
For years, international SEO focused on helping search engines show the correct page to the correct user. AI systems introduce a different challenge. Now we must help machines understand the correct facts for the correct audience, market, and context. We must ensure clarity on which product information applies in Germany, which regulations apply in Japan, and which services are available in Canada. Often, an equally complex challenge is which local brand names map to the same global product, and which facts are globally true versus market-specific.
These are not crawling and retrievability problems. They are data integrity problems. In many ways, the next generation of international SEO may resemble hreflang at the knowledge level rather than at the URL level. The challenge is no longer simply routing users to the correct page. The challenge is ensuring machines understand the correct version of the truth.
The Next Competitive Advantage
The banking analysis that inspired this article illustrates the issue well. Most of the institutions had no shortage of schema. Their websites contained thousands of lines of structured data and numerous schema types. What they lacked was a coherent representation of how the business itself operated.
That focus makes sense because discoverability remains a prerequisite for participation. However, discoverability alone will not be enough. The organizations that thrive in the next phase of search may not be those with the most schema markup, the most pages, or the most AI-ready endpoints. They may be the organizations that provide the clearest, most complete, and most trustworthy representation of how their entities, products, services, locations, brands, and markets relate to one another.
The next challenge is determining whether machines understand how the business actually works. That shift may ultimately prove more important than any individual schema property, API endpoint, or AI optimization tactic. As search engines and AI systems become increasingly capable of retrieving information, the competitive advantage will move toward organizations that can provide context, preserve relationships, and maintain the integrity of their knowledge.
Understanding an entity is only the beginning. Understanding how that entity relates to everything around it is where the real value lies.
(Source: Search Engine Journal)




