Content access for automated processing

Iframely’s Preview, Processing, and Training agents follow our overall publisher policy when fetching your content on behalf of our customers. This document provides implementation details, default access levels, and publisher controls that affect what data Iframely passes to the requesting customer.

The controls described here are separate from, and in addition to, technically allowlisting — or denying — Iframely agents and individual customers at the network level.

We cover Automated Processing — the Data API — only. The Processing Agent and Training Agent handle that.

The Preview API only serves requests originating from a human end-user. For Iframely previews, see the webmasters guide, or read how to become a rich media publisher.

The access framework

Iframely’s Data API splits the content it can fetch into five content types: metadata, links, structured entities, excerpt, and fulltext.

What we make available to the requesting customer depends on these factors:

  • the use-case mode, as declared by the customer;
  • Iframely’s default content processing policy;
  • your own domain content policy, if you’ve set one;
  • other technical publisher and page signals that follow industry-recognized formats and specifications.

Content types and processing modes are the two dimensions of the permissions matrix. The other factors affect individual cells.

For what the modes mean and how developers declare them, see the full description of Processing modes in the Data API section. Quick list: Assist, Discover, Search, AI-Input, and AI-Train (preview is reserved for the Preview API).

Everything else is described below.

Content types

Iframely’s automated processing works with five content types. We don’t scrape pages or extract anything not explicitly mentioned here.

  • Metadata — your page’s title, description, and similar basic fields, and any public markup you’ve made available in recognizable formats, such as Open Graph.
  • Links — publicly declared media and rich media links: icons, thumbnails, audio, video, iframe players, or other embeds.
  • Structured entities — your declared public schema.org JSON-LD types or microdata. We group them into general categories — article, video, product, event, and similar — so it’s easier for a customer’s application to use directly. We remove articleBody, if provided and significant, when Fulltext isn’t otherwise permitted.
  • Excerpt — a limited snippet of body text, capped at a sentence ending roughly around 450 characters, plus the hyperlinks to your other pages or to external pages found in that body. Page headings, and the link text for those hyperlinks, are only included when Fulltext is also available for the same request — otherwise you get the snippet and bare links only.
  • Fulltext — your page’s complete main body content.

Metadata and links are the same public information used for link previews in the Preview API. They’re therefore available by default to every automated processing mode, including AI-Train.

Structured entities, Excerpt, and Fulltext are available only where your policy, or Iframely’s default, permits them for that particular mode — or where a domain or page signal promotes them.

Iframely default policy

Absent any signal from you, this is what each mode receives by default:

  • Assist — metadata, links, structured data, an excerpt, and fulltext.
  • Discover — metadata, links, structured data, and an excerpt.
  • Search — metadata, links, structured data, and an excerpt.
  • AI-Input — metadata, links, and structured data.
  • AI-Train — metadata and links only. Anything more requires your explicit permission — see AI-Train below.

Our default policy is the product of a real effort to balance publishers and consumers of your content, and to build trust and encourage collaboration between the two.

Domain or page signals can further expand or restrict the default policy. For example, Fulltext isn’t available for Assist where content is marked as not freely accessible — see A paywall below.

Generally, where a signal is ambiguous or conflicting, we lean toward the more restrictive interpretation — though the exact handling can vary by signal.

AI-Train is special: the absence of a restriction is never assumed to be permission. Beyond public preview-level data, extending access for AI-Train requires an explicit, affirmative signal from you specifically — see About the Iframely Training Agent.

Domain content policy

A domain owner can override Iframely’s default policy — opting in for more permission, or opting out for tighter restriction.

Please create a free publisher account with Iframely and go to the Manage Domains section. Follow the prompts to add and verify your domains. You can then set a content policy for the entire domain, or add rules for specific paths, or even for specific Iframely customers.

Set a content policy for your domain Set a content policy for your domain

Signals and controls

Beyond the policies above, a number of signals can extend or restrict what’s available. An explicit restriction generally takes priority over a signal or policy that would otherwise extend access.

robots.txt

Both the Processing Agent and the Training Agent check your robots.txt before fetching anything, following RFC 9309, the Robots Exclusion Protocol.

If robots.txt is temporarily unavailable (a server error or timeout), we hold off — no processing occurs during that window. If it becomes permanently unavailable — a 4xx response, or unreachable for over 30 days — we treat that as allowed, and proceed as if no rules were set.

Disallow blocks fetching your content entirely, for the paths it covers.

Allow, or no rule at all, confirms our default access but doesn’t extend it further. Many hosted platforms emit Allow by default, so it can’t be read as a deliberate grant today — though this interpretation may change in the future.

Content Signals

Set inside robots.txt, using the shared vocabulary at contentsignals.org and extended by Cloudflare’s use field.

It lets you set search, ai-input, and ai-train to yes or no, optionally scoped to specific paths, and declare a use preference.

Assist and Discover aren’t part of this shared vocabulary. Both are narrower, more contained uses than Search, AI-Input, or AI-Train — so if common vocabulary had included them, we’d expect a publisher who permits the broader uses to permit these too. Because of that, we treat a yes to any of Search, AI-Input, or AI-Train as also meaning yes for Assist and Discover — worth knowing if you only meant to open one specific use.

The use signal is currently informational only. We pass it through to the requesting customer in the access object, so they know your stated preference, but it doesn’t itself gate access yet.

Robots meta tags and the X-Robots-Tag header

The standard robots meta directives (noindex, nofollow, nosnippet, max-snippet, max-image-preview, noimageindex, and similar) restrict or extend specific content types accordingly, whether set as a <meta> tag or the X-Robots-Tag HTTP header.

One exception worth knowing: noindex does not restrict Assist — a live request on behalf of one specific person is treated differently from being included in a search index.

A published license

A recognized open license (Creative Commons and similar) or an RSL license extends or restricts access according to what it permits, since it speaks directly to whether your content can be redistributed at all.

Markdown responses and llms.txt

Iframely requests content for Automated Processing with an accept: markdown request header — see, for example, this site for background on this emerging convention.

Serving a markdown version of your content, or publishing an llms.txt file, are both treated as a deliberate signal that you’ve built for machine consumption — that it’s welcome and encouraged — and access is extended accordingly.

A paywall

Iframely honors the isAccessibleForFree property for articles and posts in your schema data.

If your content is marked as not freely accessible (isAccessibleForFree: false), we don’t include fulltext for any mode, at the moment. A paywall is a paywall, and a strong signal that you want the content consumed on your own site. We may reconsider this for Discover and Search specifically in the future, since those modes are meant to bring people to your page rather than replace the visit.

Excerpt isn’t affected by this. A short, limited snippet remains available under the usual defaults, since it’s a much smaller disclosure than fulltext, and you can restrict it separately if you want tighter control — see Robots meta tags above.

If you want fulltext made available despite a paywall marker, mark isAccessibleForFree: true. You can also reach out to us directly, or use your Iframely dashboard, for anything that doesn’t fit this general signal.

JSON-LD

A mere presence of structured data does not signal additional permissions. Entities are already allowed by default on every processing mode except AI-Train. Some data properties, however, can be interpretted as promoting signals.

Beyond the isAccessibleForFree: true paywall flag above, a substantial articleBody field in your JSON-LD can itself extend access to fulltext. Structured media types (video, image, audio) can extend access to links.