Chat with AI and run tools instantly.
Let AI agents search Internet Archive books and PDFs, explore catalog metadata, inspect Wayback snapshots, and retrieve archived HTML.
Live probe refreshed Sep 18, 2026 · Endpoint host mcp.mcpbundles.com
Built for
Web Archiving, Research & Academia, Journalism & Media, Legal Discovery
Search books by title
Uses catalog metadata search — not Wayback URL lookup.
Search Internet Archive for books titled Plato published as texts, return the top matches with identifier, creator, and date.
Find phrase in scanned texts
Connect Internet Archive to any MCP client in minutes
Opens ChatGPT on the web or desktop and asks it to use the WebMCP tools available here.
You’ll sign in to MCPBundles when your client connects.
https://mcp.mcpbundles.com/bundle/internet-archiveSign in once, then chat here with saved access — one connection to many servers, with a history of what your AI ran.
MCPBundles publishes this directory for MCP discovery. Except where we host or operate an offering, third-party services run under their own terms. Product and company names on this page are used in a descriptive, identifying way (including under nominative fair use where applicable); they remain the property of their owners. Nothing here grants you rights in those marks, and nothing here is an offer to sell a third party's services. Terms
Internet Archive is used here only to identify this integration; MCPBundles is not affiliated with or endorsed by Internet Archive or its owner.
Opens ChatGPT on the web or desktop and asks it to use the WebMCP tools available here.
You’ll sign in to MCPBundles when your client connects.
https://mcp.mcpbundles.com/bundle/internet-archiveSign in once, then chat here with saved access — one connection to many servers, with a history of what your AI ran.
Chat with AI and run tools instantly.
Browse all toolsOCR search inside volumes — distinct from title-only metadata search.
Full-text search Internet Archive for 'restorative justice' in digitized books and return highlighted excerpts with page numbers and item identifiers.
Find archived snapshots
Turns a URL into a concrete historical record with replay links.
Check whether this URL has Wayback Machine captures, then return the closest snapshot date, archive link, HTTP status, and content type.
Retrieve archived HTML
Moves from Wayback discovery to content replay for historical analysis.
Find a reliable archived snapshot of this page from early 2023, retrieve the raw HTML, and summarize the visible page title and key text.
What can agents do with Internet Archive here?
Agents can search catalog metadata, run OCR full-text search in books/PDFs, fetch item metadata and files, check Wayback snapshots, batch-check URLs, search CDX metadata, build capture timelines, and retrieve archived HTML.
Does Internet Archive need sign-in or a paid account?
No. The public catalog, full-text, Wayback, and metadata workflows here work without an Archive.org sign-in or paid account.
When should I use full-text search vs catalog metadata search?
Use catalog metadata when matching titles, creators, collections, or identifiers. Use full-text search when you need words that appear inside scanned books or PDFs, with highlights and page numbers.
Domain knowledge for Internet Archive — workflow patterns, data models, and gotchas for your AI agent.
Two distinct search surfaces plus Wayback web archaeology. No auth required.
| Goal | Surface |
|---|---|
| Was a web URL captured? Replay old HTML? | Wayback — availability, CDX, timemap, get content |
| Find items by title, creator, collection, mediatype | Catalog metadata — Lucene queries on item records |
| Walk an entire collection (deep paging) | Catalog scrape — cursor batches (min 100 per call) |
| Find a phrase inside scanned books/PDFs | Full-text (OCR) — highlighted excerpts + page numbers |
| Files and description for one identifier | Item metadata lookup |
| Page locations for text in one known book | Search inside item (when the Books API responds) |
Catalog metadata matches titles/subjects/collections — not words on page 282. Full-text search matches OCR inside volumes. Wayback matches live-web URLs over time. Do not substitute one for another.
Start here
Archive
Wayback
Other capabilities
Archive Catalog Scrape
archive_catalog_scrapeScroll through large Internet Archive catalog result sets using the Scrape API. Use when metadata search pagination is not enough to walk an entire collection. Returns an opaque cursor for the next batch. Searches item metadata only — not OCR/full-text inside books.
Try in chatAgents can search catalog metadata, run OCR full-text search in books/PDFs, fetch item metadata and files, check Wayback snapshots, batch-check URLs, search CDX metadata, build capture timelines, and retrieve archived HTML.
No. The public catalog, full-text, Wayback, and metadata workflows here work without an Archive.org sign-in or paid account.
© 2026 ThinkChain Inc. All rights reserved.
Use catalog metadata when matching titles, creators, collections, or identifiers. Use full-text search when you need words that appear inside scanned books or PDFs, with highlights and page numbers.
Use CDX search when availability returns an empty result, when you need many captures, or when you need metadata such as status codes, MIME types, digests, and timestamp ranges.
Add the MCPBundles server URL to your MCP client configuration (Claude Desktop, Cursor, VS Code, etc.). The URL format is: https://mcp.mcpbundles.com/bundle/internet-archive. Authentication is handled automatically.
Internet Archive provides 10 tools that can be called by AI agents, along with a SKILL.md that gives your AI agent domain knowledge about when and how to use them.
Internet Archive uses open data APIs — no authentication required.
MCPBundles is an independent platform built on the open Model Context Protocol standard. Not affiliated with Anthropic PBC or Claude.
Other MCP servers in this category from the directory index
SuggestAPI exposes commerce tools through MCP and WebMCP on its Agent Gateway