Image Search Infrastructure for AI Agents: What the Market Actually Ships

Getty shipped an MCP server this month; the free libraries have not shipped one at all. A reproducible benchmark, a registry census and the terms of the four big photo APIs — read as if you were the machine that has to obey them.

A white 3D-rendered robot holding a magnifying glass over its head, on a pink background.
Photo via Unsplash

Every agent framework shipped this year can write, plan, call tools and open a pull request. Ask one for a photograph and the whole thing degrades to a shrug: a stock URL invented from memory, a broken hotlink, or an image generated because searching for a real one was harder than making one up. The models are not the weak link. The image layer underneath them is — and it is missing in a very specific, measurable way.

This is a market study, not a manifesto. Everything below was measured or read from a primary source on 18 August 2026: a census of the official MCP registry, the published terms of the four large free photo APIs, the access conditions of the two paid incumbents that shipped agent connectors this month, and a small reproducible benchmark you can re-run in a minute with no API key.

One boundary, so you can stop reading early if this is not your question: this article is about the access layer — retrieval, reachability, terms, quotas. It is not the publishing pipeline. If what you need is how to turn a draft into illustrated, credited, marked-up pages — the router prompt, the visual briefs, the assembly, the ImageObject markup — that is written up separately in the pipeline that illustrates AI-written articles.

What changes when the caller is a machine

A human searching for a photo types two or three words, scans a grid, and knows a good picture when they see one. Every stock search engine ever built assumes that loop. An agent breaks all three assumptions at once:

  1. It writes sentences, not keywords. A model that has just drafted a section about mortgage paperwork does not emit mortgage; it emits “a woman signing a mortgage document at a kitchen table while a bank adviser explains the terms”. That is the natural output of a language model, and it is exactly what a token-matching index cannot serve.
  2. It cannot scan a grid. Whatever the tool returns is what the agent knows. If the response is a list of URLs with no description, dimensions or credit, the agent has to fetch and look at each image — or, far more often, guess.
  3. It has no hands. It cannot upload screenshots to a review queue, accept terms, rotate a key, or wait three business days for production approval. Every human step in an API's onboarding is a wall the agent stops at.

So “image search infrastructure for AI agents” is not a marketing phrase. It is a checklist: natural-language retrieval, machine-readable results rich enough to reason over, and an authentication and quota model a program can complete on its own.

The test: a sentence against a keyword index

The cheapest way to see the problem is to send an agent's own words to a keyword engine. The test below uses Openverse, the openly licensed media search run by WordPress.org, because its API is public, needs no key, and anyone reading this can re-run it in under a minute. Four briefs, of the kind a model writes when it is illustrating what it just drafted:

openverse_briefs.py — no key, no signup, 20 lines
import json, urllib.request, urllib.parse

BRIEFS = [
  "a woman signing a mortgage document at a kitchen table while a bank"
  " adviser explains the terms, warm morning light",
  "An elderly couple enjoying a peaceful retirement together",
  "a child discovering snow for the first time in a suburban garden",
  "a delivery cyclist waiting at a red light in heavy rain",
]

def count(q):
    u = "https://api.openverse.org/v1/images/?" + urllib.parse.urlencode({"q": q})
    r = urllib.request.Request(u, headers={"User-Agent": "benchmark/1.0"})
    return json.load(urllib.request.urlopen(r))["result_count"]

for b in BRIEFS:
    print(count(b), "|", b[:48])
# 0 | a woman signing a mortgage document at a kitchen
# 0 | An elderly couple enjoying a peaceful retirement
# 0 | a child discovering snow for the first time in a
# 0 | a delivery cyclist waiting at a red light in hea

Four briefs, four empty result sets. Now cut each brief down to the two keywords a human would have typed — signing document, elderly couple, child snow, cyclist rain — and every one of them returns a full page of photographs (the API reports 240, which is where it caps the count for anonymous callers). The catalogue holds the pictures. The sentence is what it cannot parse.

The brief an agent actually writes Keyword engineOpenverse, full sentence Keyword enginecut to 2 keywords Semantic enginePexafy, full sentence
“a woman signing a mortgage document at a kitchen table while a bank adviser explains the terms…” 0 results 240 (capped) 16 results · 148 ms
“An elderly couple enjoying a peaceful retirement together” 0 results 240 (capped) 16 results · 135 ms
“a child discovering snow for the first time in a suburban garden” 0 results 240 (capped) 16 results · 147 ms
“a delivery cyclist waiting at a red light in heavy rain” 0 results 240 (capped) 16 results · 149 ms

The Pexafy column is the same four sentences sent unmodified to search_photos on the hosted MCP server, with the latency the server itself reported. This is a comparison of query shapes, not of catalogue quality — and the catalogue is emphatically not the problem. Openverse's own statistics endpoint reports 915 million images across 52 sources on the day of the test, and it is excellent at what it was built for: keyword search, for people. Nothing about it is broken. It was simply never designed for a caller that thinks in sentences. Why tag matching breaks on a described scene — and what semantic retrieval does instead — is argued from the reader's side in searching with a sentence instead of keywords; what is new here is the measurement, and what it means when the caller is a program rather than a person.

“An elderly couple enjoying a peaceful retirement together” · 16 results in 135 ms · three of them
One query, three different libraries, credits attached. No word in that sentence is a tag: peaceful and together are what put benches, gardens and quiet water at the top instead of a stock portrait of two people over sixty.

The market, as of 18 August 2026

Retrieval is one half. The other half is reachability: can an agent get to the catalogue at all, and on what terms? Since the Model Context Protocol became the way assistants attach to outside systems — its current specification, dated 28 July 2026, moved the protocol to a stateless request/response core so servers sit behind ordinary load balancers1 — the honest way to survey this market is to look at what is actually published, not what is announced.

The paid incumbents arrived this month

Getty Images launched its MCP server on 12 August 2026, six days before this article, exposing creative, editorial and archival search and download to AI workflows.2 It is aimed at enterprise teams and platform builders, and its own access page is unambiguous about the entry price: you need an active Getty Images agreement, such as Premium Access, plus acceptance of the Web Service terms. There is no free tier. Shutterstock content is reachable through an MCP server as well — 24 tools covering search, collections and licensing — with downloads gated behind an account that carries an active subscription.3

Both moves are rational and both are for a specific buyer. Getty booked $981.3 million of revenue in 20254 in a stock photography market sized at roughly $5.4 billion for 20265; connectors that require a contract protect that revenue rather than opening it. If you are a bank or a broadcaster, this is exactly what you want. If you are a developer wiring a side project — or an agent that must not stop and ask a human to sign something — it is a closed door.

The free libraries have not shipped one at all

Unsplash, Pexels and Pixabay are the three catalogues that most builders reach for, and none of them publishes an official MCP server. What exists instead is a shelf of community wrappers. We crawled the official MCP registry in full — 22,607 servers, 75,487 published versions — and kept every entry whose name or description is about photo or stock-image search. The pattern is uniform:

Official MCP registry, 18 Aug 2026 Count What that means for an agent
Servers listed in the registry, all categories 22,607 The whole ecosystem an agent can draw tools from
Servers matching photo / image / stock / picture and the library names 137 Most are editors, generators or personal photo tools, not search
…that actually search stock or free-licence photography 11 The entire addressable shelf
…published by Unsplash, Pexels, Pixabay or Openverse themselves 0 Every one is a third-party wrapper; one is labelled “Unofficial” in its own description
…that run locally and require your third-party API key 7 UNSPLASH_ACCESS_KEY, PEXELS_API_KEY, PIXABAY_API_KEY… a human must obtain each one first
…reachable as a hosted remote server 4 All four are single-source gateways from one third-party operator, in front of the same keyword APIs
Entries for Getty or Shutterstock 0 Their servers exist, but are distributed outside the public registry, to account holders
census.py — the registry crawl behind that table
import json, urllib.request, urllib.parse

def crawl(term):                      # the registry paginates with `nextCursor`
    out, cursor = {}, None
    while True:
        p = {"limit": "100", "search": term} | ({"cursor": cursor} if cursor else {})
        d = json.load(urllib.request.urlopen(
            "https://registry.modelcontextprotocol.io/v0/servers?" + urllib.parse.urlencode(p)))
        for e in d["servers"]: out[e["server"]["name"]] = e["server"]
        cursor = (d.get("metadata") or {}).get("nextCursor")
        if not cursor or not d["servers"]: return out

servers = {}
for t in ["photo", "image", "stock", "picture", "unsplash",
          "pexels", "pixabay", "getty", "shutterstock", "openverse"]:
    servers |= crawl(t)

# 137 unique servers · remote vs local is `server["remotes"]`
# the required third-party keys are in packages[].environmentVariables

Every one of those wrappers is a thin client of the same three keyword APIs, which means the benchmark above applies to all of them: the transport changed, the retrieval did not. A wrapper cannot make a tag index understand a sentence.

The clause that stops a machine is not the quota

The free libraries' quotas and terms are compared side by side, in a common unit, in the free stock photo API comparison — the approval queues, the mandatory download call, the 24-hour caching obligation, the attribution rules. That article reads them as a developer choosing a provider. Read them again as the program that has to obey them and one line stands out that has nothing to do with volume:

Pixabay's query field accepts 100 characters. The four briefs in the benchmark above average 72, and the mortgage one runs to 114 — it would be rejected outright, not ranked badly. Pexels and Unsplash publish no length limit because they never expected one to be needed: their documented examples are Ocean, Tigers, Pears. A field sized for two words is a statement about who the caller was expected to be, and no amount of quota fixes it. The semantic endpoint an agent talks to takes 500, because the trailing clause — “…while a bank adviser explains the terms” — is the part that makes the retrieval good.

The same reading applies to onboarding. A demo tier that becomes production only after a human reviews screenshots of your app is not a rate limit an agent can back off from; it is a step it cannot perform at all. Pixabay states the intent explicitly — the API is for legitimate human usage, and mass automated downloading is prohibited. None of this is unreasonable: photographers give the work away and someone pays for the bandwidth. It simply describes an API whose caller was assumed to have a screen and a pair of hands.

Four requirements, and who meets them

Put the two halves together and the specification for an image layer an agent can actually use comes out short and testable. It is deliberately not a provider comparison — catalogue size, pricing and licence terms are lined up in the comparison article. These four are only about whether a program can reach the pictures at all:

  1. Semantic retrieval — a full sentence returns a ranked page, not an empty set.
  2. Free licences, several sources — nothing to license per image, and no single library's taste defining every page you publish.
  3. A hosted server with its own sign-in — remote, OAuth, no local process to install and no third party's API key to paste into a config file.
  4. A free tier a machine can complete on its own — no screenshot review, no contract, a monthly quota and a per-minute rate limit a worker can pace itself against.
As of 18 Aug 2026 Semantic Free licences, multi-source Hosted server + OAuth Self-serve free tier
Getty Images MCP Natural-language search Licensed catalogue, per-contract Yes, for account holders No — active agreement required
Shutterstock via MCP Keyword-led Licensed catalogue, subscription Yes, with account auth No — subscription to download
Unsplash / Pexels / Pixabay wrappers No — keyword APIs underneath Single source each No — local process, your own key Inherits the host API's human onboarding
Openverse No — keyword index Yes, very broad CC catalogue No official MCP server Yes, open API
Pexafy Yes — sentences, up to 500 characters 9 free-licence sources, 9M+ photos Yes — mcp.pexafy.com/mcp, OAuth 2.1 Yes — 5,000 requests/month, 20/min, no card

We build Pexafy, so read that last row with the scepticism it deserves — and then check it: the registry census is reproducible in one script, the Getty and Shutterstock access conditions are on their own pages, and the three photo APIs' limits are in their own docs, all linked at the bottom. The claim is narrow and falsifiable: on 18 August 2026 we could not find a second service that meets all four. If you know one, we will add it to this table — the point of writing the requirements down is that anyone can run them against any product, including ours.

What a good image server does over MCP

A connector is not a REST endpoint with a new coat of paint. Two uses of this layer are written up elsewhere and are not repeated here — illustrating a draft, in the pipeline article, and keeping a long series visually coherent, in the one on volume. What belongs to the infrastructure is the thing both depend on and neither can add afterwards: three server-side decisions, made once, for every agent that ever connects.

1. Return objects an agent can reason about. Whatever the tool returns is all the model knows — it cannot scan a grid. The difference between the two shapes below is the difference between an assistant that reasons about photographs and one that relays links:

The same photo, as two different tool results
# Typical keyword wrapper: the agent has to fetch and look, or guess
[{ "url": "https://…/photo-8795398.jpeg" }, { "url": "https://…/366611.jpg" }]

# One result, as returned — trimmed here, nothing added
{
  "rank": 1,                       // the handle a human uses: "more like #1"
  "photo_id": "019e14d9-e821-…",   // what get_similar_photos takes
  "width": 6424, "height": 4283, "orientation": "landscape",
  "color_hex": "#918872", "blur_hash": "LPI#Px?aDikCGwW?M{kD-VR*s.fl",
  "source": "Pexels", "license_type": "free",
  "alt_description": "Older couple sits together on wooden bench in a park",
  "attribution": { "plain": "Photo by Anastasia Shuraeva on Pexels", "html": "…" },
  "urls": { "thumb": "…", "small": "…", "regular": "…", "large": "…" }
}

With the second shape an agent can say “#3 is portrait, it will break your hero slot”, drop a duplicate photographer, emit the credit line, and hand your template a width/height pair that kills layout shift — without fetching a single image. With the first, every one of those decisions is a guess or a round trip.

2. Number the results, so a person can point. Results come back ranked #1, #2, #3…, which is how anyone refers to a photo in conversation — “more like #3” — with no identifiers copied by hand. In clients that support MCP Apps the thumbnails render inline, and opening one shows metadata that was already in the tool result, at no extra call.

3. Stay read-only, and be explicit about it. Three tools, no write scope, no account mutation. One of them has no keyword equivalent anywhere: search_photos_by_image takes a reference image — a client's existing hero, a screenshot pasted into the chat — plus an optional sentence to bend it (“like this, but at night”). There is no field in a tag-matching API where that request can even be typed. And an agent that cannot change anything is an agent whose worst outcome is an empty result set rather than a support ticket.

The two surfaces, and which one you want

There are exactly two ways in, and the choice is about who is in the loop rather than about capability — the tools and the catalogue are identical on both.

The connector, when a person is. One URL, https://mcp.pexafy.com/mcp, added once in an assistant's connector settings; sign-in happens in a browser window and the client receives its own credentials, so there is no key to paste into a config file and none to rotate later. Clients that do not implement OAuth send a Pexafy API key as a bearer token against the same endpoint — which is what makes the connector usable from a cron job on a build server, where nobody is around to click “Allow”. The step-by-step wiring for Claude, ChatGPT and a fleet config file is in the pipeline article; it has not changed.

The HTTP API, when no one is. A batch job filling a pool for 800 topic clusters should not pretend to be a chat client: same account, same key, plain HTTP, 100 results per call. The pacing rules that go with it — the rate-limit headers, why a per-minute 429 is worth retrying and a monthly-quota 429 is not, what a worker costs at volume — belong to the article on illustrating at scale, which is where that worker is written out in full.

The point worth keeping here is narrower, and it is the one the market study turns on: both surfaces are reachable by a program on the day it is written, with no queue and no contract.

What this does not solve

An honest list, because each of these has cost someone a rollback.

  • The licence still governs the picture. Attribution is not required by the Pexafy API terms and every result ships an attribution string ready to render — but the licence attached by the originating library applies to your use of that photo. Render the credit automatically; it is cheaper than auditing later.
  • Free-licence catalogues are not editorial archives. No news, no sport, no recognisable brands or public figures on demand. If your page needs a photograph of a specific event or person, that is what the licensed incumbents are for, and their connectors now exist.
  • It is retrieval, and only retrieval. Diagrams, charts and screenshots are rendered from code, not searched; deduplication across your pages is a column in your database, not a parameter on the endpoint; and no photograph makes a thin page rank — Google's spam policies treat scaled content abuse the same whether or not the pages are illustrated.6 Those three are pipeline problems with pipeline answers, worked out in the pipeline article and the one on volume.

Where to start

  1. Re-run the benchmark. Twenty lines, no key. Whatever image API you use today, send it three of your own agent's briefs verbatim and count the empty result sets.
  2. Attach one connector to the assistant you already use and ask it, in a sentence, for a photo you actually need this week. That is the whole evaluation.
  3. Move the briefs into the agent — one scene per slot, written from the draft, never the title.
  4. Add the photo_id column before you publish anything at volume.
  5. Run it on the free tier — 5,000 searches a month at 20 a minute until a build window, not a bill, forces you up.

References & footnotes

1 Model Context Protocol specification of 28 July 2026: stateless request/response core, issuer validation per RFC 9207, Client ID Metadata Documents replacing Dynamic Client Registration, and method/tool names carried in HTTP headers for gateway routing.

2 Getty Images newsroom, 12 August 2026 — “Getty Images Launches MCP Server to Connect Creative and Editorial Content to AI Workflows and Products”. The access page states that use requires an active Getty Images agreement, such as Premium Access, and acceptance of the Web Service terms.

3 Shutterstock's catalogue is exposed through an MCP server offering 24 tools (search, collections, licensing), distributed through a third-party MCP platform rather than announced on Shutterstock's own developer portal at the time of writing; licensing and download require an account with an active subscription.

4 Getty Images Holdings, full-year 2025 results (reported 16 March 2026): revenue of $981.3 million, up 4.5% year on year.

5 Mordor Intelligence, stock photography market: USD 5.44 billion in 2026, growing at 6.86% CAGR. Market-sizing estimates vary by methodology; it is quoted here only to give the order of magnitude.

6 Google Search spam policies — scaled content abuse: generating many pages primarily to manipulate rankings and offering little value to users, whether created through automation, human effort or a combination.

Frequently asked questions

Why do stock photo APIs return nothing when an AI agent searches them?
Because the agent sends a sentence and the API indexes tags. Sent verbatim to a keyword image API, four agent-style briefs — for example “An elderly couple enjoying a peaceful retirement together” — each returned zero results; cut to two keywords (elderly couple), the same catalogue returned hundreds. Pixabay caps its query at 100 characters, which most agent briefs exceed on their own. A semantic engine embeds the whole sentence, so clauses like “tired” or “at night” change the ranking instead of breaking the match.
Is there an official MCP server for Unsplash, Pexels or Pixabay?
No. A full crawl of the official MCP registry on 18 August 2026 found 11 servers that search stock or free-licence photography, and none published by Unsplash, Pexels, Pixabay or Openverse themselves — one wrapper even labels itself “Unofficial”. Seven run locally over stdio and require you to supply that library's own API key; the four remote ones are single-source gateways operated by a third party in front of the same keyword APIs.
What does the Getty Images MCP server cost, and can anyone use it?
Getty launched its MCP server on 12 August 2026 for developers, enterprise teams and AI platform builders, covering creative, editorial and archival search and download. Access requires an active Getty Images agreement, such as Premium Access, plus acceptance of the Web Service terms; no free tier is offered and pricing is handled by an account representative. Shutterstock content is similarly reachable over MCP, with licensing gated behind an active subscription.
What does an image search API need in order to be usable by an AI agent?
Four things. Semantic retrieval, so a full sentence returns a ranked page rather than an empty set. Results rich enough to reason over — id, sizes, dimensions, orientation, colour, licence, photographer, ready-made credit — because whatever the tool returns is all the agent knows. A hosted server with its own sign-in, so there is no local process to install and no third-party key to paste. And a free tier a machine can complete on its own: no screenshot review, no contract, a monthly quota plus a per-minute rate limit a worker can pace itself against.
Can a CI job or a headless worker use an MCP image server that requires OAuth?
It depends on the server, and it is worth checking before you build on one. Pexafy’s server at mcp.pexafy.com/mcp is dual-mode: a browser sign-in for interactive clients, and the same endpoint accepting a Pexafy API key as Authorization: Bearer <key> or x-api-key for clients that implement no OAuth flow — which is what makes it usable from a cron job or a build runner where nobody can click “Allow”. The three tools are read-only, so a compromised runner key cannot mutate an account. A server that only speaks interactive OAuth cannot be driven headlessly at all.
How can I find out which image MCP servers exist and what each one requires?
Query the official registry and read two fields. GET registry.modelcontextprotocol.io/v0/servers?search=photo&limit=100 returns servers matching a term; paginate with the nextCursor value in metadata (note the camel case — next_cursor silently returns one page). On each entry, remotes tells you whether it is a hosted server or a local process, and packages[].environmentVariables lists the third-party API keys you would have to obtain yourself. Repeating that across photo, image, stock and the library names on 18 August 2026 returned 22,607 servers in total, 11 of them doing stock photo search.

Stop hunting for keywords. Describe what you mean.

Search 9M+ free-to-use images by meaning — in any language, in under 100 ms.