oassis — web scraping, crawling and browser control for AI agents
Scrape to markdown, map and crawl sites, drive a real browser. Pay per call, no API key, no signup.
¿Debería usar esto?
Calidad y seguridad
Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.
Costo de contexto
Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.
Instalar
Instalación con un clic
Agrega esto a tu archivo `claude_desktop_config.json`:
{
"mcpServers": {
"scraper-crawler": {
"url": "https://api.oassis.dev/mcp"
}
}
}Puntos de conexión remotos
https://api.oassis.dev/mcpstreamable-httpQué puede hacer
Inventario de herramientas
Herramientas (11)
⚪web_scrape(url, maxAge, formats, selectors, json, ...)
Processes a page and returns every output you ask for at once: markdown, html, links, screenshot, PDF, accessibility tree, elements by selector, AI-structured data, and `controls` (what can be clicked). One call, and a partial failure does not void the rest. From $0.001 per output. A url pointing at a PDF, Word, Excel or CSV file is converted to markdown instead, with no browser, for $0.002.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "Page to process."
},
"maxAge": {
"type": "number",
"description": "Accept an answer up to this many milliseconds old. A cache hit costs $0.0002 instead of the format price. Leave it out to force a fresh render."
},
"formats": {
"type": "array",
"items": {
"type": "string",
"enum": [
"html",
"markdown",
"links",
"screenshot",
"pdf",
"elements",
"json",
"accessibility",
"controls"
]
},
"description": "Outputs you want in the same response. `controls` is the map of what can be clicked; `elements` needs `selectors`; `json` needs `json.prompt`."
},
"selectors": {
"type": "array",
"items": {
"type": "string"
},
"description": "CSS selectors for `elements`."
},
"json": {
"type": "object",
"description": "For the `json` format: `prompt` and/or `schema`.",
"properties": {
"prompt": {
"type": "string"
},
"schema": {
"type": "object"
}
}
},
"wait": {
"type": "object",
"description": "When to consider the page loaded: `until`, `selector`, `timeout`.",
"properties": {
"until": {
"type": "string",
"enum": [
"load",
"domcontentloaded",
"networkidle0",
"networkidle2"
]
},
"selector": {
"type": "string"
},
"timeout": {
"type": "number"
}
}
},
"html": {
"type": "string",
"description": "Raw HTML instead of `url`."
}
},
"required": []
}⚪web_session_open(url, maxAge, formats, selectors, json, ...)
Opens a browser on a page and leaves it open, returning the map of controls. Use it when something has to be FILLED IN or CLICKED, not just read: inside the session the `controls` references keep working and you can act on the same state. $0.005 plus the outputs. Close it with web_session_close when you are done.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "Page to process."
},
"maxAge": {
"type": "number",
"description": "Accept an answer up to this many milliseconds old. A cache hit costs $0.0002 instead of the format price. Leave it out to force a fresh render."
},
"formats": {
"type": "array",
"items": {
"type": "string",
"enum": [
"html",
"markdown",
"links",
"screenshot",
"pdf",
"elements",
"json",
"accessibility",
"controls"
]
},
"description": "Outputs you want in the same response. `controls` is the map of what can be clicked; `elements` needs `selectors`; `json` needs `json.prompt`."
},
"selectors": {
"type": "array",
"items": {
"type": "string"
},
"description": "CSS selectors for `elements`."
},
"json": {
"type": "object",
"description": "For the `json` format: `prompt` and/or `schema`.",
"properties": {
"prompt": {
"type": "string"
},
"schema": {
"type": "object"
}
}
},
"wait": {
"type": "object",
"description": "When to consider the page loaded: `until`, `selector`, `timeout`.",
"properties": {
"until": {
"type": "string",
"enum": [
"load",
"domcontentloaded",
"networkidle0",
"networkidle2"
]
},
"selector": {
"type": "string"
},
"timeout": {
"type": "number"
}
}
}
},
"required": [
"url"
]
}⚪web_act(sessionId, actions, url, maxAge, formats, ...)
Runs actions against an open session and returns the resulting state. Actions: {navigate}, {click:{ref}}, {type:{ref,text,clear}}, {select:{ref,value}}, {press}, {scroll}, {wait}, {back}. The `ref` is the one `controls` gave you. It stops at the first failure and tells you where. $0.0005 per action.
Esquema de entrada
{
"type": "object",
"properties": {
"sessionId": {
"type": "string",
"description": "The one web_session_open returned."
},
"actions": {
"type": "array",
"items": {
"type": "object"
},
"description": "Actions, in order."
},
"url": {
"type": "string",
"description": "Page to process."
},
"maxAge": {
"type": "number",
"description": "Accept an answer up to this many milliseconds old. A cache hit costs $0.0002 instead of the format price. Leave it out to force a fresh render."
},
"formats": {
"type": "array",
"items": {
"type": "string",
"enum": [
"html",
"markdown",
"links",
"screenshot",
"pdf",
"elements",
"json",
"accessibility",
"controls"
]
},
"description": "Outputs you want in the same response. `controls` is the map of what can be clicked; `elements` needs `selectors`; `json` needs `json.prompt`."
},
"selectors": {
"type": "array",
"items": {
"type": "string"
},
"description": "CSS selectors for `elements`."
},
"json": {
"type": "object",
"description": "For the `json` format: `prompt` and/or `schema`.",
"properties": {
"prompt": {
"type": "string"
},
"schema": {
"type": "object"
}
}
},
"wait": {
"type": "object",
"description": "When to consider the page loaded: `until`, `selector`, `timeout`.",
"properties": {
"until": {
"type": "string",
"enum": [
"load",
"domcontentloaded",
"networkidle0",
"networkidle2"
]
},
"selector": {
"type": "string"
},
"timeout": {
"type": "number"
}
}
}
},
"required": [
"sessionId",
"actions"
]
}🟢web_scrape_batch(urls, formats, selectors, json, wait, ...)
A batch OF SCRAPES: reads a list of urls YOU give it (2 to 50, from any sites) and returns a jobId. It discovers nothing on its own — for that use web_crawl. Charged up front per url; urls that fail and urls served from the cache are refunded. Poll it with web_batch_status.
Esquema de entrada
{
"type": "object",
"properties": {
"urls": {
"type": "array",
"items": {
"type": "string"
},
"description": "The urls to read, 2 to 50. They do not have to share a site."
},
"formats": {
"type": "array",
"items": {
"type": "string",
"enum": [
"html",
"markdown",
"links",
"screenshot",
"pdf",
"elements",
"json",
"accessibility",
"controls"
]
},
"description": "Outputs you want in the same response. `controls` is the map of what can be clicked; `elements` needs `selectors`; `json` needs `json.prompt`."
},
"selectors": {
"type": "array",
"items": {
"type": "string"
},
"description": "CSS selectors for `elements`."
},
"json": {
"type": "object",
"description": "For the `json` format: `prompt` and/or `schema`.",
"properties": {
"prompt": {
"type": "string"
},
"schema": {
"type": "object"
}
}
},
"wait": {
"type": "object",
"description": "When to consider the page loaded: `until`, `selector`, `timeout`.",
"properties": {
"until": {
"type": "string",
"enum": [
"load",
"domcontentloaded",
"networkidle0",
"networkidle2"
]
},
"selector": {
"type": "string"
},
"timeout": {
"type": "number"
}
}
},
"maxAge": {
"type": "number",
"description": "Accept an answer up to this many milliseconds old. A cache hit costs $0.0002 instead of the format price. Leave it out to force a fresh render."
}
},
"required": [
"urls"
]
}🟢web_batch_status(jobId, limit, cancel)
Checks a batch of scrapes: status, how many are done, and the results. Free. Pass `cancel: true` to stop it and get the urls it never read refunded.
Esquema de entrada
{
"type": "object",
"properties": {
"jobId": {
"type": "string"
},
"limit": {
"type": "number",
"description": "Results to return (default 50)."
},
"cancel": {
"type": "boolean",
"description": "Stop the batch and refund what it did not read."
}
},
"required": [
"jobId"
]
}⚪web_map(url, limit, includePage, search, includeSubdomains, ...)
Every url of a site, fast and cheap: its sitemap plus, optionally, the links on the page. Use it BEFORE crawling, to see what is there and decide what is worth reading. $0.0003 with `includePage: false` (no browser at all), $0.0015 with the page.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The site to map."
},
"limit": {
"type": "number",
"description": "Urls to return (default 1000, max 5000)."
},
"includePage": {
"type": "boolean",
"description": "Render the page too (default true)."
},
"search": {
"type": "string",
"description": "Keep only urls containing this text."
},
"includeSubdomains": {
"type": "boolean"
},
"includePaths": {
"type": "array",
"items": {
"type": "string"
}
},
"excludePaths": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"url"
]
}🟢web_crawl(url, limit, maxDepth, formats, includeSubdomains, ...)
Follows a site's links and reads every page. Returns a jobId; poll it with web_crawl_status. Charged up front for the pages it is allowed to read (`limit`), and the pages it never reads are refunded. Use web_map first if you only need the urls.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "Where to start."
},
"limit": {
"type": "number",
"description": "Pages it may read (default 25, max 200)."
},
"maxDepth": {
"type": "number",
"description": "How far to follow links (default 2, max 5)."
},
"formats": {
"type": "array",
"items": {
"type": "string",
"enum": [
"html",
"markdown",
"links",
"screenshot",
"pdf",
"elements",
"json",
"accessibility",
"controls"
]
},
"description": "Outputs you want in the same response. `controls` is the map of what can be clicked; `elements` needs `selectors`; `json` needs `json.prompt`."
},
"includeSubdomains": {
"type": "boolean"
},
"includePaths": {
"type": "array",
"items": {
"type": "string"
}
},
"excludePaths": {
"type": "array",
"items": {
"type": "string"
}
},
"maxAge": {
"type": "number",
"description": "Accept an answer up to this many milliseconds old. A cache hit costs $0.0002 instead of the format price. Leave it out to force a fresh render."
}
},
"required": [
"url"
]
}🟢web_crawl_status(jobId, limit, cancel)
Checks a crawl: status, pages read, discovered and still queued, and the pages themselves. Free. Pass `cancel: true` to stop it and get the unread pages refunded.
Esquema de entrada
{
"type": "object",
"properties": {
"jobId": {
"type": "string"
},
"limit": {
"type": "number",
"description": "Pages to return (default 50)."
},
"cancel": {
"type": "boolean",
"description": "Stop the crawl and refund what it did not read."
}
},
"required": [
"jobId"
]
}🟢web_search_exa(query, limit, snippets, domains, excludeDomains, ...)
Search the web with Exa's index: a query instead of a url, for when you do not know where to look. Returns title, url and a snippet per result. To read the pages, pass the urls to web_scrape_batch. The engine is named because the price is Exa's, passed through with no markup and read from its own payment challenge on every call — today $0.007 per search.
Esquema de entrada
{
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "What to search for."
},
"limit": {
"type": "number",
"description": "Results (default 10, max 50)."
},
"snippets": {
"type": "boolean",
"description": "Text alongside each result (default true)."
},
"domains": {
"type": "array",
"items": {
"type": "string"
},
"description": "Only these domains."
},
"excludeDomains": {
"type": "array",
"items": {
"type": "string"
},
"description": "Never these domains."
},
"since": {
"type": "string",
"description": "Only results published after this ISO date."
}
},
"required": [
"query"
]
}🟢web_feedback(verdict, route, reference, url, comment)
Tell us an answer was good or bad. FREE. Use it when a result is wrong — empty markdown, a control map missing a button, data that does not match the page — with the url or the jobId so it can be reproduced. It is the only way we learn that we read a page badly: our logs cannot tell that apart from a page that is simply like that.
Esquema de entrada
{
"type": "object",
"properties": {
"verdict": {
"type": "string",
"enum": [
"good",
"bad"
]
},
"route": {
"type": "string",
"description": "Which tool or endpoint it is about."
},
"reference": {
"type": "string",
"description": "The jobId or sessionId it happened on."
},
"url": {
"type": "string",
"description": "The page that came out wrong."
},
"comment": {
"type": "string",
"description": "What you expected and what you got."
}
},
"required": [
"verdict"
]
}⚪web_session_close(sessionId)
Closes a session and stops billing browser time. Free. If you do not close it, it closes itself after a minute without use.
Esquema de entrada
{
"type": "object",
"properties": {
"sessionId": {
"type": "string"
}
},
"required": [
"sessionId"
]
}Comunidad
Evidencia