Four Ways to Sell Content to an AI, and Who Keeps What

used 0× by assistantsmercati

Four marketplaces now let a content owner be paid by an AI, and they disagree about the most basic question: what is the unit being sold. That disagreement matters more than the fee schedule, because the unit decides who wins. A per-fetch market pays for being read; a per-citation market pays for being useful. They reward opposite kinds of content.

The four units

| Platform | Unit priced (checked against vendor docs) | Reported cut | |---|---|---| | Cloudflare pay-per-crawl | per request, one flat price for the whole property | ~30%, estimated by Open Markets Institute; Cloudflare discloses nothing | | TollBit | per 1,000 pages accessed; per query via its MCP endpoint | publisher keeps 100% of the posted rate; buyer pays a separate, undisclosed fee | | ScalePost | per access, single integration, broad buyer aggregation | ~15%, per the same OMI report; not independently verified | | ProRata | proportional attribution across the sources that shaped an answer | 50/50 of ProRata's own advertising and subscription revenue |

Read that last column with suspicion. Only one of the four figures comes from the company itself — TollBit's, in its own documentation. The other three come from a single report by Courtney Radsch and Karina Montoya of the Open Markets Institute, reported by Media Copilot, an organisation whose brief is arguing that platforms extract too much. Cloudflare's 30% is labelled an *estimate* in that report, and Cloudflare's own pay-per-crawl blog post and developer FAQ — both fetched for this page — disclose no cut at all.

The column is also not comparing like with like. ScalePost's 15% and Cloudflare's 30% are marketplace take rates on money a buyer pays for access. ProRata's 50/50 is a split of ProRata's *own* advertising and subscription revenue, which only exists if Gist sells ads: a different business with a different denominator. TollBit's "100%" is true of the publisher's posted rate and silent about the buyer-side transaction fee, so TollBit's total take rate is unknown, not zero. Four numbers in one column is a presentation artefact, not a comparison.

Cloudflare's model is mechanical: the crawler either presents payment intent and receives a 200 with a crawler-charged header, or receives a 402 with the price. Cloudflare's docs are explicit that a site owner "can only set a single price that applies to all crawlers configured with the 'Charge' option" — one price, whole property. It requires no negotiation and no relationship, and it pays identically for a fetch that produced a cited answer and a fetch that produced nothing. Cloudflare acts as Merchant of Record and settles to the publisher (x402: HTTP 402 Finally Gets a Job, at 32 Cents a Transaction).

ProRata sits at the other extreme, running an answer engine (Gist) on a licensed library and attributing payment at answer level, proportional to each source's contribution. It is the only one of the four whose unit is tied to usefulness rather than access — but the money it splits is advertising revenue it must first earn, not a toll a crawler must first pay.

TollBit is the interesting middle. Its published rate unit is *per 1,000 pages accessed*, but its MCP endpoint states that "every agent request carries a TollBit JWT" and that it meters usage to "monetize every query" — the unit that applies to a tool or an API rather than a document (The x402 Market, Counted: 14,766 Listings, 12,741 Calls a Day, One Winner). TIME appears on TollBit's own customer wall, with TollBit claiming its traffic data helped TIME negotiate with OpenAI and Perplexity. That is TollBit's account of TollBit's value, not an independent finding, and TollBit's own site currently renders its publisher count as "Over 0 sites", so no count is quoted here.

Cloudflare is assembling the whole stack

The structural fact of 2026 is not a new entrant but a consolidation. Cloudflare announced the acquisition of Human Native — an AI data marketplace founded in 2024, backed by LocalGlobe and Mercuri — in a press release dated 15 January 2026. It now holds most layers of the transaction: detection (AI Crawl Control), tolling (pay-per-crawl), and dataset supply (Human Native).

The fourth layer is announced but not shipped. On 1 July 2026 Cloudflare announced the Monetization Gateway, which would let customers charge per request for "web pages, datasets, APIs, or MCP tools" behind Cloudflare, settling in stablecoins over x402 (The Earn-Spend Loop: Why Machine Payment Is Half-Built). As of that post it is a waitlist, not a product, and no fee is disclosed. Treat the completed stack as a plan, not a fact.

A seller choosing a marketplace in 2026 is largely choosing how much of that stack to accept from one vendor.

Above the marketplaces: an open standard nobody has to join

RSL — Really Simple Licensing, launched 10 September 2025 — encodes licensing and royalty terms directly in robots.txt, deliberately modelled on RSS. Tim O'Reilly's launch quote is explicit about the lineage: "RSL builds directly on the legacy of RSS, providing the missing licensing layer for the AI-first Internet." That design intent is what makes it the only mechanism here aimed at the long tail rather than at the twenty brands that can negotiate.

Its launch supporters — Reddit, Yahoo, People Inc., Ziff Davis, O'Reilly Media, Medium, wikiHow, Quora, Fastly and others — are all on the *publisher* side. Whether crawler operators honour RSL is a separate question from whether publishers emit it, and the supporter list answers only the second. That gap is the same trap that made llms.txt cargo cult.

Our thesis, stated so it can be wrong

a publisher comparing a 15% cut with a 30% cut is optimising the smaller variable. The choice of unit, and the granularity within that unit, moves revenue by more than the spread between marketplace fees.

The one piece of primary evidence on this points our way. Archer, Ghili and Haghpanah (arXiv:2604.01416, 1 April 2026) fit a pricing agent to 8,939 articles and 80,451 buyer queries from a major German technology publisher and report a 65% revenue gain over a single static price, 47% over two-category pricing, and 40% over the publisher's own eight-segment editorial taxonomy. A 65% swing from getting pricing granularity right dwarfs a 15-point difference in take rate. Their conclusion — that "content is too heterogeneous for a fixed pricing framework" — is a direct verdict on Cloudflare's one-price-per-property design.

What would falsify the transfer: it is one publisher, one country, one vertical (hardware reviews); willingness-to-pay is calibrated from crawler traffic rather than observed at those prices; and the preprint is not peer-reviewed. If the same experiment on a general-news corpus showed segmentation gains under ~20%, our claim would be wrong.

The practical shape:

- Deep archives with heavy crawler traffic and shallow usefulness per page do best on per-fetch, where every read pays regardless of outcome. - Content that is rarely read but decisive when read does best on per-citation or per-query, and is systematically underpaid per fetch.

Sinapsi's own corpus is the second kind, which is why the per-query unit is the relevant one for us and per-fetch is not. That is a statement about our position, not a finding.

The honest ceiling

Choosing the right unit changes whether you are paid fairly. We do not currently have a defensible number for whether it changes being paid *much*. Published per-crawl rates sit somewhere between fractions of a cent and a few cents depending on who is quoting, and the one academic dataset reports relative gains without absolute prices. A specific willingness-to-pay figure previously carried on this page could not be traced to any source and has been removed rather than laundered into a citation.

What survives without a number is the Open Markets Institute's structural argument: the licensing market pays the brand-name corpus that has negotiating leverage, while the long tail sees little. That is an advocacy organisation's framing, and it is the framing this page has found no primary evidence against.

Verified against

29 claims checked against these sources

Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.