Image Search Infrastructure for AI Agents: What the Market Actually Ships
Getty shipped an MCP server this month; the free libraries have not shipped one at all. A reproducible benchmark, a registry census and the terms of the four big photo APIs — read as if you were the machine that has to obey them.
Every agent framework shipped this year can write, plan, call tools and open a pull request. Ask one for a photograph and the whole thing degrades to a shrug: a stock URL invented from memory, a broken hotlink, or an image generated because searching for a real one was harder than making one up. The models are not the weak link. The image layer underneath them is — and it is missing in a very specific, measurable way.
This is a market study, not a manifesto. Everything below was measured or read from a primary source on 18 August 2026: a census of the official MCP registry, the published terms of the four large free photo APIs, the access conditions of the two paid incumbents that shipped agent connectors this month, and a small reproducible benchmark you can re-run in a minute with no API key.
One boundary, so you can stop reading early if this is not your question: this article is about
the access layer — retrieval, reachability, terms, quotas. It is not the
publishing pipeline. If what you need is how to turn a draft into illustrated, credited,
marked-up pages — the router prompt, the visual briefs, the assembly, the
ImageObject markup — that is written up separately in
the pipeline that
illustrates AI-written articles.
What changes when the caller is a machine
A human searching for a photo types two or three words, scans a grid, and knows a good picture when they see one. Every stock search engine ever built assumes that loop. An agent breaks all three assumptions at once:
- It writes sentences, not keywords. A model that has just drafted a section
about mortgage paperwork does not emit
mortgage; it emits “a woman signing a mortgage document at a kitchen table while a bank adviser explains the terms”. That is the natural output of a language model, and it is exactly what a token-matching index cannot serve. - It cannot scan a grid. Whatever the tool returns is what the agent knows. If the response is a list of URLs with no description, dimensions or credit, the agent has to fetch and look at each image — or, far more often, guess.
- It has no hands. It cannot upload screenshots to a review queue, accept terms, rotate a key, or wait three business days for production approval. Every human step in an API's onboarding is a wall the agent stops at.
So “image search infrastructure for AI agents” is not a marketing phrase. It is a checklist: natural-language retrieval, machine-readable results rich enough to reason over, and an authentication and quota model a program can complete on its own.
The test: a sentence against a keyword index
The cheapest way to see the problem is to send an agent's own words to a keyword engine. The test below uses Openverse, the openly licensed media search run by WordPress.org, because its API is public, needs no key, and anyone reading this can re-run it in under a minute. Four briefs, of the kind a model writes when it is illustrating what it just drafted:
import json, urllib.request, urllib.parse
BRIEFS = [
"a woman signing a mortgage document at a kitchen table while a bank"
" adviser explains the terms, warm morning light",
"An elderly couple enjoying a peaceful retirement together",
"a child discovering snow for the first time in a suburban garden",
"a delivery cyclist waiting at a red light in heavy rain",
]
def count(q):
u = "https://api.openverse.org/v1/images/?" + urllib.parse.urlencode({"q": q})
r = urllib.request.Request(u, headers={"User-Agent": "benchmark/1.0"})
return json.load(urllib.request.urlopen(r))["result_count"]
for b in BRIEFS:
print(count(b), "|", b[:48])
# 0 | a woman signing a mortgage document at a kitchen
# 0 | An elderly couple enjoying a peaceful retirement
# 0 | a child discovering snow for the first time in a
# 0 | a delivery cyclist waiting at a red light in hea
Four briefs, four empty result sets. Now cut each brief down to the two keywords a human would
have typed — signing document, elderly couple, child snow,
cyclist rain — and every one of them returns a full page of photographs (the API
reports 240, which is where it caps the count for anonymous callers). The catalogue holds the
pictures. The sentence is what it cannot parse.
| The brief an agent actually writes | Keyword engineOpenverse, full sentence | Keyword enginecut to 2 keywords | Semantic enginePexafy, full sentence |
|---|---|---|---|
| “a woman signing a mortgage document at a kitchen table while a bank adviser explains the terms…” | 0 results | 240 (capped) | 16 results · 148 ms |
| “An elderly couple enjoying a peaceful retirement together” | 0 results | 240 (capped) | 16 results · 135 ms |
| “a child discovering snow for the first time in a suburban garden” | 0 results | 240 (capped) | 16 results · 147 ms |
| “a delivery cyclist waiting at a red light in heavy rain” | 0 results | 240 (capped) | 16 results · 149 ms |
The Pexafy column is the same four sentences sent unmodified to search_photos on the
hosted MCP server, with the latency the server itself reported. This is a comparison of
query shapes, not of catalogue quality — and the catalogue is emphatically not the
problem. Openverse's own statistics endpoint reports 915 million images across 52
sources on the day of the test, and it is excellent at what it was built for: keyword
search, for people. Nothing about it is broken. It was simply never designed for a caller that
thinks in sentences. Why tag matching breaks on a described scene — and what semantic
retrieval does instead — is argued from the reader's side in
searching with a
sentence instead of keywords; what is new here is the measurement, and what it means when the
caller is a program rather than a person.
The market, as of 18 August 2026
Retrieval is one half. The other half is reachability: can an agent get to the catalogue at all, and on what terms? Since the Model Context Protocol became the way assistants attach to outside systems — its current specification, dated 28 July 2026, moved the protocol to a stateless request/response core so servers sit behind ordinary load balancers1 — the honest way to survey this market is to look at what is actually published, not what is announced.
The paid incumbents arrived this month
Getty Images launched its MCP server on 12 August 2026, six days before this article, exposing creative, editorial and archival search and download to AI workflows.2 It is aimed at enterprise teams and platform builders, and its own access page is unambiguous about the entry price: you need an active Getty Images agreement, such as Premium Access, plus acceptance of the Web Service terms. There is no free tier. Shutterstock content is reachable through an MCP server as well — 24 tools covering search, collections and licensing — with downloads gated behind an account that carries an active subscription.3
Both moves are rational and both are for a specific buyer. Getty booked $981.3 million of revenue in 20254 in a stock photography market sized at roughly $5.4 billion for 20265; connectors that require a contract protect that revenue rather than opening it. If you are a bank or a broadcaster, this is exactly what you want. If you are a developer wiring a side project — or an agent that must not stop and ask a human to sign something — it is a closed door.
The free libraries have not shipped one at all
Unsplash, Pexels and Pixabay are the three catalogues that most builders reach for, and none of them publishes an official MCP server. What exists instead is a shelf of community wrappers. We crawled the official MCP registry in full — 22,607 servers, 75,487 published versions — and kept every entry whose name or description is about photo or stock-image search. The pattern is uniform:
| Official MCP registry, 18 Aug 2026 | Count | What that means for an agent |
|---|---|---|
| Servers listed in the registry, all categories | 22,607 | The whole ecosystem an agent can draw tools from |
| Servers matching photo / image / stock / picture and the library names | 137 | Most are editors, generators or personal photo tools, not search |
| …that actually search stock or free-licence photography | 11 | The entire addressable shelf |
| …published by Unsplash, Pexels, Pixabay or Openverse themselves | 0 | Every one is a third-party wrapper; one is labelled “Unofficial” in its own description |
| …that run locally and require your third-party API key | 7 | UNSPLASH_ACCESS_KEY, PEXELS_API_KEY, PIXABAY_API_KEY… a human must obtain each one first |
| …reachable as a hosted remote server | 4 | All four are single-source gateways from one third-party operator, in front of the same keyword APIs |
| Entries for Getty or Shutterstock | 0 | Their servers exist, but are distributed outside the public registry, to account holders |
import json, urllib.request, urllib.parse
def crawl(term): # the registry paginates with `nextCursor`
out, cursor = {}, None
while True:
p = {"limit": "100", "search": term} | ({"cursor": cursor} if cursor else {})
d = json.load(urllib.request.urlopen(
"https://registry.modelcontextprotocol.io/v0/servers?" + urllib.parse.urlencode(p)))
for e in d["servers"]: out[e["server"]["name"]] = e["server"]
cursor = (d.get("metadata") or {}).get("nextCursor")
if not cursor or not d["servers"]: return out
servers = {}
for t in ["photo", "image", "stock", "picture", "unsplash",
"pexels", "pixabay", "getty", "shutterstock", "openverse"]:
servers |= crawl(t)
# 137 unique servers · remote vs local is `server["remotes"]`
# the required third-party keys are in packages[].environmentVariables
Every one of those wrappers is a thin client of the same three keyword APIs, which means the benchmark above applies to all of them: the transport changed, the retrieval did not. A wrapper cannot make a tag index understand a sentence.
The clause that stops a machine is not the quota
The free libraries' quotas and terms are compared side by side, in a common unit, in the free stock photo API comparison — the approval queues, the mandatory download call, the 24-hour caching obligation, the attribution rules. That article reads them as a developer choosing a provider. Read them again as the program that has to obey them and one line stands out that has nothing to do with volume:
Pixabay's query field accepts 100 characters. The four briefs in the benchmark above average 72, and the mortgage one runs to 114 — it would be rejected outright, not ranked badly. Pexels and Unsplash publish no length limit because they never expected one to be needed: their documented examples are Ocean, Tigers, Pears. A field sized for two words is a statement about who the caller was expected to be, and no amount of quota fixes it. The semantic endpoint an agent talks to takes 500, because the trailing clause — “…while a bank adviser explains the terms” — is the part that makes the retrieval good.
The same reading applies to onboarding. A demo tier that becomes production only after a human reviews screenshots of your app is not a rate limit an agent can back off from; it is a step it cannot perform at all. Pixabay states the intent explicitly — the API is for legitimate human usage, and mass automated downloading is prohibited. None of this is unreasonable: photographers give the work away and someone pays for the bandwidth. It simply describes an API whose caller was assumed to have a screen and a pair of hands.
Four requirements, and who meets them
Put the two halves together and the specification for an image layer an agent can actually use comes out short and testable. It is deliberately not a provider comparison — catalogue size, pricing and licence terms are lined up in the comparison article. These four are only about whether a program can reach the pictures at all:
- Semantic retrieval — a full sentence returns a ranked page, not an empty set.
- Free licences, several sources — nothing to license per image, and no single library's taste defining every page you publish.
- A hosted server with its own sign-in — remote, OAuth, no local process to install and no third party's API key to paste into a config file.
- A free tier a machine can complete on its own — no screenshot review, no contract, a monthly quota and a per-minute rate limit a worker can pace itself against.
| As of 18 Aug 2026 | Semantic | Free licences, multi-source | Hosted server + OAuth | Self-serve free tier |
|---|---|---|---|---|
| Getty Images MCP | Natural-language search | Licensed catalogue, per-contract | Yes, for account holders | No — active agreement required |
| Shutterstock via MCP | Keyword-led | Licensed catalogue, subscription | Yes, with account auth | No — subscription to download |
| Unsplash / Pexels / Pixabay wrappers | No — keyword APIs underneath | Single source each | No — local process, your own key | Inherits the host API's human onboarding |
| Openverse | No — keyword index | Yes, very broad CC catalogue | No official MCP server | Yes, open API |
| Pexafy | Yes — sentences, up to 500 characters | 9 free-licence sources, 9M+ photos | Yes — mcp.pexafy.com/mcp, OAuth 2.1 |
Yes — 5,000 requests/month, 20/min, no card |
We build Pexafy, so read that last row with the scepticism it deserves — and then check it: the registry census is reproducible in one script, the Getty and Shutterstock access conditions are on their own pages, and the three photo APIs' limits are in their own docs, all linked at the bottom. The claim is narrow and falsifiable: on 18 August 2026 we could not find a second service that meets all four. If you know one, we will add it to this table — the point of writing the requirements down is that anyone can run them against any product, including ours.
What a good image server does over MCP
A connector is not a REST endpoint with a new coat of paint. Two uses of this layer are written up elsewhere and are not repeated here — illustrating a draft, in the pipeline article, and keeping a long series visually coherent, in the one on volume. What belongs to the infrastructure is the thing both depend on and neither can add afterwards: three server-side decisions, made once, for every agent that ever connects.
1. Return objects an agent can reason about. Whatever the tool returns is all the model knows — it cannot scan a grid. The difference between the two shapes below is the difference between an assistant that reasons about photographs and one that relays links:
# Typical keyword wrapper: the agent has to fetch and look, or guess
[{ "url": "https://…/photo-8795398.jpeg" }, { "url": "https://…/366611.jpg" }]
# One result, as returned — trimmed here, nothing added
{
"rank": 1, // the handle a human uses: "more like #1"
"photo_id": "019e14d9-e821-…", // what get_similar_photos takes
"width": 6424, "height": 4283, "orientation": "landscape",
"color_hex": "#918872", "blur_hash": "LPI#Px?aDikCGwW?M{kD-VR*s.fl",
"source": "Pexels", "license_type": "free",
"alt_description": "Older couple sits together on wooden bench in a park",
"attribution": { "plain": "Photo by Anastasia Shuraeva on Pexels", "html": "…" },
"urls": { "thumb": "…", "small": "…", "regular": "…", "large": "…" }
}
With the second shape an agent can say “#3 is portrait, it will break your hero slot”, drop a
duplicate photographer, emit the credit line, and hand your template a
width/height pair that kills layout shift — without fetching a single
image. With the first, every one of those decisions is a guess or a round trip.
2. Number the results, so a person can point. Results come back ranked
#1, #2, #3…, which is how anyone refers to a photo in conversation — “more like #3” —
with no identifiers copied by hand. In clients that support MCP Apps the thumbnails render inline,
and opening one shows metadata that was already in the tool result, at no extra call.
3. Stay read-only, and be explicit about it. Three tools, no write scope, no
account mutation. One of them has no keyword equivalent anywhere:
search_photos_by_image takes a reference image — a client's existing hero, a
screenshot pasted into the chat — plus an optional sentence to bend it
(“like this, but at night”). There is no field in a tag-matching API where that request
can even be typed. And an agent that cannot change anything is an agent whose worst outcome is an
empty result set rather than a support ticket.
The two surfaces, and which one you want
There are exactly two ways in, and the choice is about who is in the loop rather than about capability — the tools and the catalogue are identical on both.
The connector, when a person is. One URL,
https://mcp.pexafy.com/mcp, added once in an assistant's connector settings; sign-in
happens in a browser window and the client receives its own credentials, so there is no key to
paste into a config file and none to rotate later. Clients that do not implement OAuth send a
Pexafy API key as a bearer token against the same endpoint — which is what makes the connector
usable from a cron job on a build server, where nobody is around to click “Allow”. The
step-by-step wiring for Claude, ChatGPT and a fleet config file is in
the pipeline
article; it has not changed.
The HTTP API, when no one is. A batch job filling a pool for 800 topic clusters
should not pretend to be a chat client: same account, same key, plain HTTP, 100 results per call.
The pacing rules that go with it — the rate-limit headers, why a per-minute
429 is worth retrying and a monthly-quota 429 is not, what a worker
costs at volume — belong to
the article on
illustrating at scale, which is where that worker is written out in full.
The point worth keeping here is narrower, and it is the one the market study turns on: both surfaces are reachable by a program on the day it is written, with no queue and no contract.
What this does not solve
An honest list, because each of these has cost someone a rollback.
- The licence still governs the picture. Attribution is not required by the Pexafy API terms and every result ships an attribution string ready to render — but the licence attached by the originating library applies to your use of that photo. Render the credit automatically; it is cheaper than auditing later.
- Free-licence catalogues are not editorial archives. No news, no sport, no recognisable brands or public figures on demand. If your page needs a photograph of a specific event or person, that is what the licensed incumbents are for, and their connectors now exist.
- It is retrieval, and only retrieval. Diagrams, charts and screenshots are rendered from code, not searched; deduplication across your pages is a column in your database, not a parameter on the endpoint; and no photograph makes a thin page rank — Google's spam policies treat scaled content abuse the same whether or not the pages are illustrated.6 Those three are pipeline problems with pipeline answers, worked out in the pipeline article and the one on volume.
Where to start
- Re-run the benchmark. Twenty lines, no key. Whatever image API you use today, send it three of your own agent's briefs verbatim and count the empty result sets.
- Attach one connector to the assistant you already use and ask it, in a sentence, for a photo you actually need this week. That is the whole evaluation.
- Move the briefs into the agent — one scene per slot, written from the draft, never the title.
- Add the
photo_idcolumn before you publish anything at volume. - Run it on the free tier — 5,000 searches a month at 20 a minute until a build window, not a bill, forces you up.
References & footnotes
1 Model Context Protocol specification of 28 July 2026: stateless request/response core, issuer validation per RFC 9207, Client ID Metadata Documents replacing Dynamic Client Registration, and method/tool names carried in HTTP headers for gateway routing.
2 Getty Images newsroom, 12 August 2026 — “Getty Images Launches MCP Server to Connect Creative and Editorial Content to AI Workflows and Products”. The access page states that use requires an active Getty Images agreement, such as Premium Access, and acceptance of the Web Service terms.
3 Shutterstock's catalogue is exposed through an MCP server offering 24 tools (search, collections, licensing), distributed through a third-party MCP platform rather than announced on Shutterstock's own developer portal at the time of writing; licensing and download require an account with an active subscription.
4 Getty Images Holdings, full-year 2025 results (reported 16 March 2026): revenue of $981.3 million, up 4.5% year on year.
5 Mordor Intelligence, stock photography market: USD 5.44 billion in 2026, growing at 6.86% CAGR. Market-sizing estimates vary by methodology; it is quoted here only to give the order of magnitude.
6 Google Search spam policies — scaled content abuse: generating many pages primarily to manipulate rankings and offering little value to users, whether created through automation, human effort or a combination.
Primary sources, checked on 18 August 2026: MCP registry API · MCP specification · Getty Images MCP announcement · Unsplash API docs · Pexels API docs · Pixabay API docs · Openverse API · Google spam policies · Pexafy API & MCP docs. The Openverse counts, the Pexafy result counts and every latency in this article are live values captured that day; plan limits are read from the live pricing catalogue at render time.
Frequently asked questions
Why do stock photo APIs return nothing when an AI agent searches them?
elderly couple), the same catalogue returned hundreds. Pixabay caps its query at 100 characters, which most agent briefs exceed on their own. A semantic engine embeds the whole sentence, so clauses like “tired” or “at night” change the ranking instead of breaking the match.Is there an official MCP server for Unsplash, Pexels or Pixabay?
What does the Getty Images MCP server cost, and can anyone use it?
What does an image search API need in order to be usable by an AI agent?
Can a CI job or a headless worker use an MCP image server that requires OAuth?
mcp.pexafy.com/mcp is dual-mode: a browser sign-in for interactive clients, and the same endpoint accepting a Pexafy API key as Authorization: Bearer <key> or x-api-key for clients that implement no OAuth flow — which is what makes it usable from a cron job or a build runner where nobody can click “Allow”. The three tools are read-only, so a compromised runner key cannot mutate an account. A server that only speaks interactive OAuth cannot be driven headlessly at all.How can I find out which image MCP servers exist and what each one requires?
GET registry.modelcontextprotocol.io/v0/servers?search=photo&limit=100 returns servers matching a term; paginate with the nextCursor value in metadata (note the camel case — next_cursor silently returns one page). On each entry, remotes tells you whether it is a hosted server or a local process, and packages[].environmentVariables lists the third-party API keys you would have to obtain yourself. Repeating that across photo, image, stock and the library names on 18 August 2026 returned 22,607 servers in total, 11 of them doing stock photo search.