---
title: "Data API content for Automated Processing"
description: "Structured data and page content returned by Iframely Data API — entities, excerpt, and fulltext."
---

# Data API content

The [Data API](/docs/data-api) adds three content objects to the `meta` and `links` returned by the Preview APIs: `entities`, `excerpt`, and `fulltext`.

`meta` describes the URL and provides information you can use to categorize the page. `links` contains media links Iframely finds on the page, such as thumbnail images, videos, audio, and rich media embeds. See the [Iframely API](/docs/iframely-api) documentation for their complete response formats.

The content available in a response depends on the [access decision](/docs/content-access) for the processing mode and use in your request.

## Entities

`entities` contains structured data declared by the publisher, such as JSON-LD and schema.org markup.

Iframely does not add or infer fields in `entities`. If a publisher does not declare a piece of structured data, it is not added here.

Iframely groups entities by type. For example:

```json
"entities": {
    "article": {
        "@context": "https://schema.org",
        "@type": "Article",
        "name": "The Great Wave off Kanagawa",
        "url": "https://en.wikipedia.org/wiki/The_Great_Wave_off_Kanagawa",
        "...": "...",
        "image": "...",
        "headline": "..."
    },
    "navigation": {"...": "..."}
}
```

The available entity groups are:

`article`, `video`, `media` (movies, TV series, radio and podcast episodes, audiobooks), `music`, `image`, `info` (Course, HowTo, FAQ), `publication`, `organization`, `person`, `software`, `place`, `product`, `offer`, `event`, `review`, `recipe`, `dataset`, `creative` (creative and artwork), `navigation`, `webpage`, and `code`.

See the `@type` field of each entity for its original structured-data type.

Multiple entities of the same group are returned as an array rather than a single object.

## Excerpt

`excerpt` contains selected content extracted from the main body of the page. It can be used for certain automated processing needs without requiring the full page text.

It contains:

- `snippet` — a short excerpt of the page text;
- `hyperlinks` — links found in the page body, both internal and external;
- `headings` — the page's heading structure.

For example:

```json
"excerpt": {

    "snippet": "Woodblock print by Hokusai (1831) The Great Wave off KanagawaArtistKatsushika HokusaiYear1831TypeDimensions24.6 cm × 36.5 cm (9.7 in × 14.4 in) The Great Wave off Kanagawa (Japanese: 神奈川沖浪裏, Hepburn: Kanagawa-oki Nami Ura; lit. 'Under the Wave off Kanagawa')[a] is a woodblock print by the Japanese ukiyo-e artist Hokusai (1760–1849), created in late 1831 during the Edo period of Japanese history.",

    "headings": [
        {
            "level": 3,
            "text": "Depth and perspective"
        },
        {"..."}
    ],

    "hyperlinks": [
        {
            "href": "https://en.wikipedia.org/wiki/Katsushika_Hokusai",
            "text": "Katsushika Hokusai"
        },
        {"..."}
    ]
}
```

### Snippet

`snippet` is automatically generated from the page content. It is derived from the extracted page text rather than from a meta description provided by the publisher.

A snippet is normally limited to 450 characters. Iframely prefers to end at the last complete sentence around the limit.

The limit applies whether or not `fulltext` is available. A response containing both may therefore have a `snippet` that overlaps with the beginning of `fulltext`.

Publishers can change the maximum length or disable snippets using the `max-snippet` directive in `robots.txt`, a page directive such as `<meta name="robots">`, or response HTTP headers.

### Hyperlinks

`hyperlinks` is a list of links found in the main body of the page. Each link can include `text` and `href` values.

When available, a hyperlink may also include a `rel` property with one or more of these values:

`follow`, `nofollow`, `ugc`, `sponsored`, `alternate`, `author`, `external`, `help`, `license`, `me`, `next`, `prev`, `last`, `first`, `privacy-policy`, `tag`, `terms-of-service`.

Use hyperlinks to build search indexes or identify relationships between the page and other content.

When `fulltext` is not available, `hyperlinks` does not include `text` values. Anchor text is included only when `fulltext` is also allowed.

### Headings

`headings` represents the heading structure of the page. It is included only when `fulltext` is also allowed.

Each heading includes `text` and `level`. The list is flat rather than a hierarchical object.

A heading may include an `id` when the source page provides its own heading anchor. Iframely does not generate heading IDs. Use it to link directly to the corresponding section when needed.

## Fulltext

`fulltext` contains the main body text Iframely derives from the page, in Markdown format.

For example:

```json
"fulltext": "Woodblock print by Hokusai (1831) | The Great Wave off Kanagawa | | | --------------------------- | --------------------------------------------------------------------------------------------- | | Artist | [Katsushika Hokusai](https://en.wikipedia.org/wiki/Woodblock%5Fprinting%5Fin%5FJapan \"Woodblock printing in Japan\") | | Year | 1831 | | Type | | | Dimensions | 24.6 cm × 36.5 cm (9.7 in × 14.4 in) | _**The Great Wave off Kanagawa**_ ([Japanese](https://en.wikipedia.org/wiki/Japanese%5Flanguage \"Japanese language\"): 神奈川沖浪裏, [Hepburn](https://en.wikipedia.org/wiki/Hepburn%5Fromanization \"Hepburn romanization\"): _Kanagawa-oki Nami Ura_; [lit.](https://en.wikipedia.org/wiki/Literal%5Ftranslation \"Literal translation\") \"Under the Wave off Kanagawa\")[\\[a\\]](https://en.wikipedia.org/wiki/The%5FGreat%5FWave%5Foff%5FKanagawa#cite%5Fnote-1) is a [woodblock print](https://en.wikipedia.org/wiki/Woodblock%5Fprinting%5Fin%5FJapan \"Woodblock printing in Japan\") by the Japanese _[ukiyo-e](https://en.wikipedia.org/wiki/Ukiyo-e \"Ukiyo-e\")_ artist [Hokusai](https://en.wikipedia.org/wiki/Hokusai \"Hokusai\") (1760–1849), created in late 1831 during the [Edo period](https://en.wikipedia.org/wiki/Edo%5Fperiod \"Edo period\") of [Japanese history](https://en.wikipedia.org/wiki/Japanese%5Fhistory \"Japanese history\")...."
```

It is available only when the [access decision](/docs/content-access) allows `fulltext` for the requested processing mode and use.

Fulltext is subject to the default processing policy and publisher signals. Iframely is most restrictive with `fulltext` by default because it provides the largest amount of page content.

By default, `assist` is the only mode that `fulltext` is available in. A publisher signal can block or restrict access, while an appropriate publisher signal or explicit publisher permission can make `fulltext` available for other processing modes too.

When the publisher explicitly provides a Markdown response, Iframely can return the full Markdown provided by the publisher. Otherwise, Iframely derives `fulltext` from the main content of the page.

If the publisher's structured data includes an `articleBody` field, Iframely does not simply copy that field into `fulltext`. The structured data remains available through `entities`, subject to the access decision.

See [publisher explanation](/docs/content-policy) for how processing modes, publisher signals, and controls affect content availability.

See the [Data API Content Access](/docs/content-access) object description for how that access information is passed in the API response.