About the Iframely Training Agent

If you see this agent accessing your URLs, it’s because one of our customers is using Iframely to obtain content for AI training or dataset use on their behalf, where you — the publisher — have permitted it. We act as a technical intermediary for that customer; we don’t train anything ourselves, and we don’t grant rights in your content beyond what you’ve made available.

Iframely runs three separate agents, each scoped to what our customers can use it for:

  • Preview Agent — used only for link previews and rich media embeds triggered by a human-shared URL. Does not allow automated processing or AI training under this identity.
  • Processing Agent — handles automated processing requests on a customer’s behalf: Assist, Discover, Search, and AI-Input. Does not allow training.
  • Training Agent (this one) — dedicated identity for AI-Train requests, where a publisher has permitted it. By default, it accesses only metadata and links; anything beyond that requires your explicit permission.

If you’re looking for indexing, search, or AI-assist controls, see the Processing Agent doc instead — this page covers training and dataset use only.

What we do

Iframely is a software-as-a-service company, and some of our customers build or fine-tune AI models. Where publisher has permitted it, this agent retrieves content on that customer’s behalf for training or dataset-building purposes.

We use a separate identity for this specifically so you can recognize and target it individually. See how to allow (or disallow) Iframely Training Agent below.

Iframely is designed to produce only a certain content data types — meta, media, structured entities, exceprts and fulltext. Iframely does not scrape web pages for arbitery content.

User-Agent string

Iframely-Training/1.3.1 (+https://iframely.com/docs/about-training)

Version and optional app name extension may and will change; the extension, when present, identifies the specific customer who triggered the request, as they’ve entered it with us. This is the identity referred to as the “Iframely Training Agent” elsewhere in our documentation and policies.

What we access

By default, this agent accesses only your declared metadata and links. This basic public information is already used for a standard link previews and we therefore consider this access level be reasonable.

Anything beyond that basic level requires either explicit permission in Iframely dashboard, or a recognized page signal, for example a Content-Signal directive, a license that covers training use.

See Content access for the complete description of these signals.

How to opt out

If you don’t want your content used for training, you don’t need to do anything beyond your existing defaults — access beyond metadata and links requires your explicit signal or permission.

If you’ve previously granted a training permission in Iframely dashboard, you can withdraw it at any time. The same is true for page signals. However, we honor a withdrawal going forward from when it’s detected; we cannot retroactively notify customers about the change on their previous requests.

Crawling behavior

The Training Agent doesn’t go beyond a signle URL — it is request-driven, on behalf of a specific customer request, one URL at a time.

It checks your robots.txt before fetching anything, following RFC 9309, along with the training-specific signals described in Content access.

Having a problem with the Training Agent?

We’re happy to answer questions, help you configure training-specific permissions, or discuss any concerns about how this agent accesses your site. Contact us at support@iframely.com.

Allow the Training Agent on your network

If you have bot protection in place and want to permit training access, add the Training Agent to your allowlist separately from the Preview and Processing Agents. Use the appropriate user-agent string and see Allowlisting Iframely for IP lists, reverse DNS, Web Bot Authentication. It also describes transitive trust and gives the full request header reference.