Bread & Law / Cited

Cited is a protocol that assesses the AI discoverability of news. It creates, stores, and analyzes technical detail on the interactions between AI platforms and news publishers.

These details can inform comms pros of the extent to which earned press coverage is visible to AI platforms. It spans AI-driven search, real-time, and training.

The level of technical detail Cited makes available can put comms in a position of strength with internal stakeholders who control budgets and access. Because it does not assume visibility is the goal, it can make assessments of both positive and negative stories with equal levels of granularity.

Cited is entirely algorithmic. There is not one iota of AI inside it.

Methodology

Cited's assessment is a layered pipeline. Each layer surfaces a different kind of evidence about whether an outlet's reporting is accessible to AI platforms. Layers run independently, combine into a per-platform posture, and distill into a single summary verdict.

Technical details for each layer are made available via copy-paste, and every assessment is a dated record that can be forwarded as-is.

Five common AI platforms are currently assessed. Adding more is conducive to scale.

The verdict

Every assessment resolves to one of three findings: sufficiently visible, not sufficiently visible, or cannot determine when there is truly no technical evidence on which to base an assessment.

The verdict is scoped to real-time retrieval and AI search. Training is reported separately, as context, and never as part of the verdict.

An outlet reads sufficiently visible if at least one platform can reach it, and visibility is always attributed to the specific platform or platforms that earned it, and it never transfers between them. One platform's access says nothing about another's. That is a technical part of how AI works, not an opinion.

Two variables are reported for every platform and kept strictly separate.

Access is the direction: reachable, not reachable, or undetermined. Basis is how confident Cited is: verified means Cited observed the real platform act and captured its egress IP; measured means direct observation of a real artifact that is nonetheless a proxy; inferred means derived from published policy and server probing.

All of this data combines to mean something about the publisher's overall posture towards how AI companies access and use their content. That is what the protocol is designed to help earned media pros keep an eye on.

Layer 1 — robots.txt

Cited fetches /robots.txt from the root domain and parses it against each AI platform's documented user agents. A publisher's instructions to AI platforms tend to live inside this file, but compliance is voluntary.

These instructions are used and retained because they are symbolic of a publication's stance on AI. But because nobody really knows which models, if any, are obeying these instructions, they are not deterministic.

Layer 2 — Other declarations

Beyond robots.txt, sites can signal preferences in other ways. Those details are assessed and analyzed at this layer.

From the same response, Cited also fingerprints the publisher's CDN and hosting. Some technical stacks provide one-click controls to disallow AI access at the network level; this is retained because it proves a publisher has the technical capacity to block AI platforms if it wants to.

Layer 4 — User-agent A/B probing

At L4, Cited pulls a sample of recent articles, fetches each with a baseline browser user agent, then refetches as each AI bot's canonical user agent.

This can prove a publisher is not serving AI platforms the same content as other traffic, indicating active blocking. If a publication's instructions to AI platforms are not being obeyed, it will show up here.

When a site blocks Cited's probes, or when there are disagreements between L4 and other layers, it is usually instructive for the protocol. Seeing these conflicts frequently is good. It underscores why Cited is a layered pipeline.

Layer 5 — Common Crawl presence

Common Crawl feeds many open and proprietary AI training corpora. Cited queries the Common Crawl CDX index for the domain and reports a coverage level and a trend direction.

While this is a direct observation of a real corpus, the protocol considers it a secondary or a proxy measurement. Presence in Common Crawl is not proof that any particular model ingested the content, and absence is not proof that none did. No AI platform discloses its training set, so Cited reports the corpus footprint and stops there.

Layer 7 — Real-origin access probing

L7 asks each named AI platform to fetch a sampled article from its own infrastructure and sends that through a canary. The canary reports as either true or false, creating a new piece of data that can prove or disprove access was possible at a given point in time.

This is the most direct evidence in the pipeline: proof that the door is open (or shut) for a specific platform right now. It reports what is retrievable, never whether a platform chose to cite the article or its underlying information.

What Cited does not claim

It measures access, not use. Cited establishes whether an AI platform can retrieve a piece of content. It does not and cannot establish that a platform used it, cited it, or was influenced by it. Access is the precondition, not the outcome.

Training is a black box for every platform. There is no interface, from any vendor, that confirms a given article is in a training set. Anything Cited says about training is inference or proxy, which is exactly why the verdict excludes it.

Verified coverage is limited in the public demo but scalable. Any licensing or managed service agreements can add new platforms to L7.

Real-origin verification is live for one platform today; the rest are assessed from published policy and server probing. That asymmetry is reported rather than smoothed over, and it narrows as each platform gains a driver.

Usage

Across all assessments to date, Cited has analyzed 113 domains and simulated 5,005 AI-bot requests across 433 distinct article pages.

As of August 28, 2026.