How to Build Your Own MCP Server for Web Scraping: A Step-by-Step Guide for Beginners
Table of contents
- Introduction: what you'll have by the end of this guide
- Preparation: tools, access, and system requirements
- Core concepts: how an mcp server works and why an ai agent needs one
- Step 1: create the project and install dependencies
- Step 2: write a minimal mcp server with your first tool
- Step 3: connect the mcp server to your ai client
- Step 4: add data extraction tools
- Step 5: connect mobile proxies and ip rotation
- Step 6: make the server reliable: retries, delays, cache, and limits
- Checking the result: a checklist for your finished mcp server
- Common mistakes when building an mcp server and how to fix them
- Extra features: a section for advanced users
- Faq: common questions about building an mcp server
- Conclusion: what you built and where to go next
Introduction: what you'll have by the end of this guide
Imagine opening a chat with an AI assistant and typing: "Go to a competitor's page, collect the names and prices of every product in the catalog, and put them in a table." The assistant doesn't say "I don't have internet access" — it actually loads the page, pulls the data, and hands you a finished result. That's exactly what you'll build by following this guide to the end. The bridge between the language model and the web will be your own MCP server, written in Python.
One important caveat: we won't be covering the ready-made Playwright MCP or other off-the-shelf solutions. We have separate articles about those on the blog. Here the goal is different: to write a server from scratch so you understand every line, can add your own tools, connect mobile proxies, and tailor the logic to your specific tasks. Your own solution is always more flexible than someone else's.
Who this guide is for
- Marketers and business owners who need to quickly collect prices, reviews, product descriptions, and competitor content without hiring a developer to build a scraper.
- Affiliates who monitor offers, landing pages, and creatives and want to hand the routine off to an AI agent.
- Developers who've heard about the MCP protocol but haven't built their own server yet and want a working template.
- Mobile proxy users who need their agent's requests to go through their proxy instead of straight from a home IP.
What you need to know beforehand
This guide is aimed at beginners. Programming experience isn't required, but it helps to understand what a command line is and how to open a file in a text editor. You can copy all the code as-is, and every part of it is explained in plain language. If you already write Python, there's a dedicated section with advanced features toward the end of the article.
How much time it'll take
Plan for 2-3 hours on your first pass. Installing the tools takes about 30 minutes, a minimal working MCP server appears within an hour, and the rest of the time goes into adding data extraction tools, connecting proxies, and testing. Repeating everything from scratch on another computer will take you 20-30 minutes.
Preparation: tools, access, and system requirements
Before writing any code, make sure you have everything you need. This section takes half an hour and saves you from half the typical problems in the following steps.
System requirements
- A computer running Windows 10/11, macOS 12 or newer, or Linux (Ubuntu 22.04 or newer). Everything described here works on any of these systems; the only differences are file paths.
- At least 4 GB of RAM and 1 GB of free disk space.
- A stable internet connection.
What to install
- Python 3.11 or newer. As of 2026, versions 3.12 and 3.13 are current. Download the installer from the official Python website. On Windows, be sure to check the Add python.exe to PATH box in the first installer window, otherwise the python command won't be found in the terminal. On macOS, it's easier to install Python via Homebrew with brew install python. On Ubuntu, run sudo apt install python3 python3-venv python3-pip.
- A text editor for code. We recommend Visual Studio Code. It's free, highlights syntax, and shows errors. Any other editor works too, even Notepad, but VS Code will make things easier.
- An MCP client — that is, an app with an AI agent that you'll connect the server to. The simplest option for beginners is Claude Desktop. MCP is also supported by the Cursor editor, VS Code with the GitHub Copilot extension, and a number of other tools. Install at least one of them before you start.
- Node.js 20 or newer. It's not needed for the server itself, but for the MCP Inspector utility we'll use to debug tools. Download the LTS installer from the official Node.js website and install with the default settings.
Access
For the proxy section, you'll need your mobile proxy details: host, port, username, and password, plus a link to change the IP address if your plan supports it. You'll find all of this in your provider's dashboard. If you don't have a proxy yet, you can still follow the guide without one: the server will work directly, and you can add the proxy later with a single line.
Backups
We'll be editing the MCP client's config file. Before that, copy it to a safe place, like your desktop, labeled "backup." If something goes wrong, you just put the copy back. Keep the server code in a separate folder and save a copy of the file after each working step, or make a Git commit if you know how to use it.
Tip: Create a separate folder with a short path free of spaces and non-ASCII characters, like C:/mcp-collector on Windows or ~/mcp-collector on macOS and Linux. Spaces and non-Latin characters in paths regularly break server launches from configs, and you'll waste an hour hunting down the cause.
Core concepts: how an MCP server works and why an AI agent needs one
Before writing the first line of code, let's get the terminology straight. Without it, the tutorial looks like a set of magic spells; with it, every action becomes logical.
What MCP is
MCP (Model Context Protocol) is an open protocol that describes how a language model communicates with external tools. Before it existed, every service invented its own way to "give AI hands." MCP standardized this: if you write a server according to the protocol, any compatible client will understand it, whether that's Claude Desktop, Cursor, or your own agent. You can think of MCP as a USB port: it doesn't matter whether you plug in a flash drive or a mouse, the port is the same.
Client and server
MCP architecture has two participants. The client is the AI app that asks questions and calls tools. The MCP server is the program that provides those tools. In our case, the server will be the skill of "going online and pulling data," and the client will be your AI assistant. The server runs locally on your computer, and the client talks to it directly.
Tools, resources, and prompts
An MCP server can expose three types of entities to the client:
- Tools — functions the model can call: "download a page," "pull all links," "change the proxy IP." This is the foundation of our guide.
- Resources — data the server provides for reading, like the contents of a settings file or the result of the last scrape.
- Prompts — pre-built query templates the user can invoke with a single command.
For web scraping, tools are enough. We'll touch on resources and prompts in the advanced section.
How the model knows what to call
There's an important nuance here. When the client connects to the server, it requests the list of tools with their names, descriptions, and parameters. These descriptions end up in the model's context. Then the model decides on its own which tool to call and with what arguments, based specifically on the description text. That's why function descriptions in our code aren't a formality, but instructions for the AI. The clearer you write what a tool does and when to use it, the more accurately the agent will work.
Transport: stdio and HTTP
The server and client need some way to exchange messages. The protocol provides two main methods. stdio — the client launches your script as a child process and communicates with it through standard input and output. This is the simplest option for local work, and it's what we'll start with. Streamable HTTP — the server runs as a web service that the client connects to by address. This option is needed if the server lives on a remote machine or multiple clients connect to it. We'll cover it in the advanced section.
⚠️ Warning: With stdio transport, the entire standard output of the process is taken up by protocol messages. If you write a regular print in your code for debugging, the client will receive garbage instead of a valid response and drop the connection. Debug messages can only go to the stderr stream. Remember this rule — it'll save you a lot of time.
Why scraping through MCP is convenient
A classic scraper is hard-coded: it knows how to collect specific fields from a specific site. The moment the markup changes, the scraper breaks. The "AI agent plus MCP server" combo works differently: the server provides universal tools (download, extract text, find elements by selector), and the model figures out the page structure and formulates the result itself. You get flexibility without rewriting code for every new source.
Step 1: Create the project and install dependencies
Goal of this stage: set up an isolated Python environment and install the libraries needed for the MCP server. By the end of this step, you'll have a project folder with a working virtual environment.
Why you need a virtual environment
A virtual environment is a separate copy of Python with its own libraries inside the project folder. It's needed so our server doesn't conflict with other Python programs on the computer and so the MCP client knows exactly which interpreter to launch. Without it, half the "it works in my terminal but not in the client" problems are guaranteed.
Step-by-step instructions
- Open the terminal. On Windows, press Win+R, type powershell, and hit Enter. On macOS, open the Terminal app via Spotlight (Cmd+Space, then type Terminal). On Linux, press Ctrl+Alt+T.
- Create the project folder and navigate into it. On Windows, run two commands: mkdir C:/mcp-collector, then cd C:/mcp-collector. On macOS and Linux: mkdir ~/mcp-collector, then cd ~/mcp-collector.
- Check the Python version with python --version (on macOS and Linux you may need python3 --version). You should see something like Python 3.12.x. If the version is below 3.11 or the command isn't found, go back to the preparation section and reinstall Python.
- Create the virtual environment with python -m venv .venv. A hidden .venv folder will appear in the project. This takes 10-20 seconds.
- Activate the environment. On Windows in PowerShell: .venv/Scripts/Activate.ps1. If PowerShell says script execution is disabled, run Set-ExecutionPolicy -Scope CurrentUser RemoteSigned, confirm with Y, and retry activation. On macOS and Linux: source .venv/bin/activate. After activation, (.venv) appears at the start of the terminal line.
- Update the package manager: python -m pip install --upgrade pip.
- Install the libraries with a single command: pip install "mcp[cli]" httpx beautifulsoup4. Here mcp is the official Python SDK for the protocol (the 1.x branch is current as of 2026), httpx is a modern HTTP request library with proxy support, and beautifulsoup4 is an HTML parsing tool. Installation takes 1-2 minutes.
- Create an empty server.py file in the project folder. In VS Code: open the folder via File, Open Folder, then click the new file icon in the left panel and type the name.
What the libraries do
- mcp handles the entire protocol: tool registration, message exchange, parameter descriptions. The FastMCP module inside it lets you declare a tool as a regular function with a decorator.
- httpx downloads pages. Unlike the outdated requests, it supports HTTP/2, async, and easy proxy configuration.
- beautifulsoup4 turns HTML into a tree you can easily search by tags and CSS selectors.
Tip: Write down the full path to the interpreter inside the virtual environment right away. On Windows it's C:/mcp-collector/.venv/Scripts/python.exe, and on macOS and Linux it's /Users/name/mcp-collector/.venv/bin/python (or /home/name/... on Linux). You'll need it when connecting to the client. You can find the exact path with where python on Windows or which python on macOS and Linux while the environment is activated.
✅ Check: Run pip list. The packages mcp, httpx, and beautifulsoup4 should be listed. Also run python -c "import mcp, httpx, bs4; print('ok')" — you should see the word ok with no errors.
Possible problems
- The python command isn't found. On Windows, reinstall Python with the Add to PATH box checked. On macOS, use python3 instead of python.
- pip complains about permissions. Most likely the environment isn't activated and you're installing packages into the system Python. Check for the (.venv) marker at the start of the line.
- A build error during installation. Update pip and try again. If that doesn't help, check that your Python version is 3.11 or higher.
Step 2: Write a minimal MCP server with your first tool
Goal of this stage: write a working MCP server with one tool that downloads a page by URL and returns its HTML. This is the foundation we'll build features on top of.
Server code
Open server.py and paste the following code in full:
import sys
import httpx
from mcp.server.fastmcp import FastMCP
mcp = FastMCP('web-collector')
HEADERS = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0 Safari/537.36',
'Accept-Language': 'ru-RU,ru;q=0.9,en;q=0.8',
}
def log(message: str) -> None:
print(message, file=sys.stderr)
@mcp.tool()
def fetch_page(url: str, max_chars: int = 20000) -> str:
'''Скачивает страницу по указанному URL и возвращает её HTML-код.
Используй, когда нужно посмотреть исходную разметку страницы.
Параметр max_chars ограничивает длину ответа, чтобы не переполнять контекст.'''
log(f'fetch_page: {url}')
with httpx.Client(headers=HEADERS, timeout=20.0, follow_redirects=True) as client:
response = client.get(url)
response.raise_for_status()
return response.text[:max_chars]
if __name__ == '__main__':
mcp.run()Line-by-line code breakdown
- FastMCP('web-collector') creates the server object with the name web-collector. That's the name the client will show in the list of connected servers.
- HEADERS are the headers we send to websites. Many sites return incomplete content or an error if the request comes without a familiar browser User-Agent. The Accept-Language header hints that we want the Russian-language version of the page.
- The log function writes messages to stderr. That's how it's done, not with a regular print, because stdout is taken up by the protocol. You'll see these messages in the client's logs and in MCP Inspector.
- @mcp.tool() is a decorator that turns a regular function into an MCP tool. The SDK automatically reads the function name, parameter types, and docstring and builds a description for the model. The default max_chars = 20000 means the parameter is optional.
- The docstring in triple quotes is what the AI will read. Here we explain what the tool does and when to use it. Write these descriptions in detail and in the language you use to talk to the agent.
- httpx.Client with follow_redirects=True automatically follows redirects, and timeout=20.0 keeps the request from hanging indefinitely.
- raise_for_status() throws an error if the site returns a 4xx or 5xx code. The SDK catches it and returns a clear error message to the client instead of silence.
- mcp.run() starts the server with stdio transport by default. It will wait for commands from the client.
First check via MCP Inspector
Running server.py directly is pointless: it'll wait for messages from the client and show nothing. To test it, we'll use MCP Inspector — a web interface that mimics a client and lets you call tools by hand.
- Make sure the virtual environment is activated and you're in the project folder.
- Run the command mcp dev server.py. This command is part of the installed mcp package with the cli extension. On first run, it downloads Inspector via npx, which takes about a minute.
- A URL like http://localhost:6274 will appear in the terminal, along with an access token in newer versions. Open the address in your browser (it often opens by itself).
- In the left panel of Inspector, check that the STDIO transport is selected, the command is python, and the arguments are server.py. Click the Connect button.
- The status indicator turns green and reads Connected. Go to the Tools tab in the top menu and click List Tools.
- The list will show the fetch_page tool with its docstring description and two parameters. Click on it.
- In the url field, enter https://example.com, leave the max_chars field empty or enter 5000. Click Run Tool.
- A result appears on the right: the page's HTML code starting with a doctype tag. Below, in the server logs tab, you'll see the line fetch_page: https://example.com.
✅ Check: Inspector shows the Connected status, the Tools list contains fetch_page, and calling it with example.com returns HTML without errors. If all that's true, your first MCP server works.
Possible problems
- mcp dev says npx isn't found. Node.js isn't installed. Install it and restart the terminal.
- Inspector opens but Connect throws an error. Check that the command field specifies the python from the activated environment. You can enter the full path to python.exe inside .venv.
- SyntaxError when connecting. The code was copied with lost indentation. In Python, indentation is mandatory: function bodies are shifted by four spaces. Check the file in the editor.
- The tool returns a 403 error. The site rejected the request. This won't happen with example.com, but for real sites we'll come back to this in the proxy step.
Step 3: Connect the MCP server to your AI client
Goal of this stage: register the server in the AI client's settings so the agent sees your tool and can call it from a regular chat. We'll cover connecting to Claude Desktop as the most common option and briefly show alternatives.
Connecting to Claude Desktop
- Open Claude Desktop. Go to settings: on Windows via the menu in the top left, Settings; on macOS via the Claude menu, Settings.
- Go to the Developer tab and click Edit Config. The folder with the claude_desktop_config.json file will open. If the file doesn't exist, the client creates it.
- Make a backup of this file by copying it to your desktop.
- Open the file in VS Code or another editor. If the file is empty, paste the contents in full. If it already has other servers, add your block inside the mcpServers object, separated by a comma.
{
"mcpServers": {
"web-collector": {
"command": "C:/mcp-collector/.venv/Scripts/python.exe",
"args": ["C:/mcp-collector/server.py"]
}
}
}On macOS and Linux, replace the paths with your own, for example /Users/ivan/mcp-collector/.venv/bin/python and /Users/ivan/mcp-collector/server.py. Note: even on Windows, paths are written with forward slashes. That's easier because backslashes in JSON need to be doubled, and Windows understands forward slashes just fine.
- Save the file. Make sure there are no extra commas after the last element and all brackets are closed. A single extra comma makes the JSON invalid, and the client silently ignores the config.
- Completely close Claude Desktop and launch it again. On Windows, closing the window isn't enough: right-click the icon in the system tray and choose Quit. The client only reads the config at startup.
- After launch, open a new chat. Under the input field, find the tools icon (a slider or plug icon). Click it: the list should show the web-collector server with one tool, fetch_page.
- Type in the chat: "Download the page https://example.com using fetch_page and tell me what the page's heading is." The client will ask permission to call the tool. Click Allow or Allow for this chat.
- In a couple of seconds, the agent replies that the page's heading is Example Domain. It made a real request through your server.
Connecting to Cursor and VS Code
In Cursor, open Settings, find the MCP section, and click Add new global MCP server. The mcp.json file opens with exactly the same structure as Claude Desktop's. Paste the same block and save. In VS Code with Copilot, create a .vscode/mcp.json file in the root of your workspace, where instead of the mcpServers key you use the servers key, with the same command and args inside. After saving, a Start button appears above the server block. In all clients, the principle is the same: specify the interpreter launch command and the script path.
Tip: In the command field, always specify the python from your virtual environment rather than just the word python. The client launches the process with its own set of environment variables, and the system python command may be a different version without the installed libraries. The full path eliminates this problem once and for all.
✅ Check: The client's interface shows the web-collector server, and the agent calls fetch_page on request and correctly summarizes the contents of example.com. In the client's logs (in Claude Desktop, that's the logs folder next to the config, file mcp-server-web-collector.log), you'll see the line fetch_page: https://example.com.
Possible problems
- The server doesn't appear in the list. Check the JSON for validity: paste the contents into any online JSON validator or open it in VS Code, which will underline errors. Make sure the client was fully restarted.
- A red error indicator next to the server. Open the log file. Most often it says ModuleNotFoundError: the wrong python is specified. Check the path in the command field.
- The agent says it can't access the internet. It didn't see the tool. Make sure tools are enabled in the slider panel, and ask explicitly: "use the fetch_page tool."
- spawn ENOENT error. The path to python or server.py is wrong. Copy the path from the file explorer and replace backslashes with forward slashes.
Step 4: Add data extraction tools
Goal of this stage: teach the server to return useful data instead of raw HTML: clean text, a list of links, and elements found by CSS selector. After this, the agent can collect structured information without wasting context on markup.
Why fetch_page alone isn't enough
A real page's HTML weighs hundreds of kilobytes, and most of it is scripts, styles, and service markup. If you hand the model everything every time, it'll quickly hit the context limit, and you'll pay for extra tokens. The right strategy: the server does the rough cleanup and structuring, and the model works with compact data. That's why we're adding three specialized tools.
Updated code
Replace the contents of server.py with the extended version. The fetch_page function remains, but the shared download logic is moved into a separate _get_html function that all tools use.
import sys
from urllib.parse import urljoin
import httpx
from bs4 import BeautifulSoup
from mcp.server.fastmcp import FastMCP
mcp = FastMCP('web-collector')
HEADERS = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0 Safari/537.36',
'Accept-Language': 'ru-RU,ru;q=0.9,en;q=0.8',
}
def log(message: str) -> None:
print(message, file=sys.stderr)
def _get_html(url: str) -> str:
log(f'GET {url}')
with httpx.Client(headers=HEADERS, timeout=20.0, follow_redirects=True) as client:
response = client.get(url)
response.raise_for_status()
return response.text
def _clean(text: str) -> str:
return ' '.join(text.split())
@mcp.tool()
def fetch_page(url: str, max_chars: int = 20000) -> str:
'''Возвращает сырой HTML страницы. Используй только когда нужна именно разметка,
например чтобы подобрать CSS-селектор. Для чтения содержимого используй extract_text.'''
return _get_html(url)[:max_chars]
@mcp.tool()
def extract_text(url: str, max_chars: int = 15000) -> str:
'''Возвращает чистый текст страницы без скриптов, стилей и разметки.
Лучший выбор, когда нужно прочитать статью, описание товара или отзывы.'''
soup = BeautifulSoup(_get_html(url), 'html.parser')
for tag in soup(['script', 'style', 'noscript', 'svg', 'header', 'footer', 'nav']):
tag.decompose()
title = _clean(soup.title.get_text()) if soup.title else ''
body = _clean(soup.get_text(' '))
return f'Заголовок: {title}. Текст: {body}'[:max_chars]
@mcp.tool()
def extract_links(url: str, limit: int = 100, contains: str = '') -> list[dict]:
'''Возвращает список ссылок со страницы: текст ссылки и полный адр��с.
Параметр contains фильтрует ссылки, в адресе которых есть указанная подстрока,
например /product/ или /catalog/.'''
soup = BeautifulSoup(_get_html(url), 'html.parser')
result = []
seen = set()
for a in soup.find_all('a', href=True):
full = urljoin(url, a['href'])
if full in seen or (contains and contains not in full):
continue
seen.add(full)
result.append({'text': _clean(a.get_text())[:120], 'url': full})
if len(result) >= limit:
break
return result
@mcp.tool()
def select_elements(url: str, css_selector: str, limit: int = 50) -> list[str]:
'''Находит на странице элементы по CSS-селектору и возвращает их текст.
Примеры селекторов: h2, .price, div.product-card, table tr.
Используй, когда нужны конкретные повторяющиеся блоки: цены, названия, строки таблицы.'''
soup = BeautifulSoup(_get_html(url), 'html.parser')
elements = soup.select(css_selector)[:limit]
return [_clean(el.get_text(' ')) for el in elements]
if __name__ == '__main__':
mcp.run()What each tool does
- extract_text removes scripts, styles, header, footer, and menu from the document, and collapses the remaining text into a single line with single spaces. The _clean function uses split and join to strip extra line breaks and tabs. The page title is prepended to the response so the agent immediately knows what it's looking at.
- extract_links collects all a tags, converts relative addresses to absolute ones with urljoin, removes duplicates via the seen set, and lets you filter links by substring. That way the agent gets, for example, all product cards from a catalog in a single call.
- select_elements is the most powerful tool. It takes a CSS selector and returns the text of the matched elements. The agent can first look at a chunk of HTML via fetch_page, figure out that prices live in the price class, and then call select_elements with the .price selector.
Notice the docstrings: we explicitly hint to the model which tool to choose in which situation. This noticeably improves the agent's performance.
How to test
- Run mcp dev server.py and connect in Inspector. The Tools list now has four tools.
- Call extract_links with the url of any news site or catalog and the contains parameter set to part of a section's address. The result is a list of objects with text and url fields.
- Call select_elements with the same address and the h2 selector. You'll get a list of headings.
- Restart Claude Desktop (no config change is needed, only the code changed) and ask: "Collect all h2 headings and links leading to the news section from the homepage of such-and-such site, and format them as a table."
Tip: If you don't know which selector you need, open the page in your browser, press F12, pick the element selector tool (the arrow icon in the top left of the panel), and click the block you want. You'll see its class in the code. A selector with a dot and the class name, like .product-title, usually works. Even better, you can simply ask the agent: "download the HTML and figure out the selector for the prices yourself."
✅ Check: All four tools are visible in Inspector and the client, extract_text returns readable text without tags, extract_links returns a list with absolute addresses, and select_elements with the h2 selector returns headings.
Possible problems
- select_elements returns an empty list. Either the selector is wrong, or the content is loaded by JavaScript after the page loads. Check via fetch_page: if the HTML doesn't contain the data, the site renders it client-side. Such sites need a browser engine — that's a topic for a separate article.
- extract_text outputs garbled text. The site returns a non-standard encoding. Add the line response.encoding = response.charset_encoding or 'utf-8' after response.raise_for_status().
- The response gets truncated. Increase max_chars in the call or ask the agent to request the page in chunks via several selectors.
Step 5: Connect mobile proxies and IP rotation
Goal of this stage: route all MCP server requests through a mobile proxy, add a tool to change the IP, and check the current address. After this, the agent will work as a mobile carrier, not from your home or office IP.
Why a data collector needs a mobile proxy
When you scrape from a single IP address, sites see dozens of identical requests in a row and start serving captchas, truncated content, or a 429 "too many requests" error. A mobile proxy solves several problems at once. First, the address belongs to a real mobile carrier, and such addresses are shared by thousands of subscribers, so sites treat them more leniently. Second, you can change the IP via a link or on a timer, spreading the load. Third, you separate the agent's working activity from your personal sessions. For a marketer, it's also a way to see a site the way a mobile user in a specific region sees it.
⚠️ Warning: A proxy is a tool for stable, correct operation of a scraper, not for breaking the rules. Collect only publicly available data, respect sites' terms of use and robots.txt, don't create excessive load, and don't collect personal data without legal grounds. You're responsible for how you use the tool.
Step-by-step instructions
- Open your mobile proxy provider's dashboard and find the connection details: host, port, username, password. They're usually assembled into a single string like login:password@host:port. Copy the IP rotation link there too, if available.
- In server.py, add import os at the top, after the other imports. Then, below the HEADERS block, add the settings:
PROXY_URL = os.environ.get('MOBILE_PROXY_URL', '')
ROTATE_URL = os.environ.get('PROXY_ROTATE_URL', '')
def _client() -> httpx.Client:
kwargs = {'headers': HEADERS, 'timeout': 30.0, 'follow_redirects': True}
if PROXY_URL:
kwargs['proxy'] = PROXY_URL
return httpx.Client(**kwargs)- In the _get_html function, replace the httpx.Client line with a call to _client(). Now it looks like this: with _client() as client. All tools automatically go through the proxy.
- Add two new tools before the if __name__ line:
@mcp.tool()
def current_ip() -> str:
'''Показывает IP-адрес, с которого сервер сейчас выходит в интернет.
Используй, чтобы убедиться, что прокси подключён, или после смены IP.'''
with _client() as client:
return client.get('https://api.ipify.org').text.strip()
@mcp.tool()
def rotate_ip() -> str:
'''Запрашивает смену IP-адреса мобильного прокси через ссылку из личного кабинета.
Вызывай, если сайт начал отдавать ошибки 429 или капчу. После вызова подожди 5-10 секунд.'''
if not ROTATE_URL:
return 'Ссылка смены IP не настроена в переменной PROXY_ROTATE_URL'
response = httpx.get(ROTATE_URL, timeout=15.0)
log(f'rotate_ip: status {response.status_code}')
return f'Запрос смены IP отправлен, ответ прокси-сервиса: {response.status_code}'- Pass the proxy details via environment variables in the client config. We deliberately don't hard-code the username and password to avoid accidentally sending them somewhere along with the file. Open claude_desktop_config.json and extend the server block with an env section:
{
"mcpServers": {
"web-collector": {
"command": "C:/mcp-collector/.venv/Scripts/python.exe",
"args": ["C:/mcp-collector/server.py"],
"env": {
"MOBILE_PROXY_URL": "http://login:password@proxy-host:port",
"PROXY_ROTATE_URL": "https://ссылка-смены-ip-из-кабинета"
}
}
}
}- Substitute real values for login, password, proxy-host, and port. If your provider gives you a SOCKS5 proxy, replace http:// with socks5:// and install an extra package with pip install httpx[socks].
- Save the config and fully restart the client.
- Ask the agent: "Call current_ip and tell me our address. Then call rotate_ip, wait ten seconds, and check the IP again." The addresses should differ.
Testing via Inspector with a proxy
Inspector can also pass environment variables. In the left panel, expand the Environment Variables section, add MOBILE_PROXY_URL and PROXY_ROTATE_URL with your values, connect, and call current_ip. The response should match the IP shown in your provider's dashboard.
Tip: Don't call rotate_ip before every request. With most providers, changing the IP takes a few seconds, and too-frequent requests may hit a rotation limit. A reasonable strategy: change the address every 30-100 requests, or only when you get 429 and 403 errors. You can bake this logic directly into _get_html, which we'll do in the next step.
✅ Check: The current_ip tool returns the proxy address, not your home one. After rotate_ip and a pause, the address changes. The extract_text and extract_links tools keep working, and the logs show GET lines with page addresses.
Possible problems
- 407 Proxy Authentication Required error. Wrong username or password, or they contain special characters. Characters like @ or : in the password need encoding: @ becomes %40, : becomes %3A.
- ConnectTimeout error. Wrong host or port, or your IP isn't on the provider's allowlist if the plan has that restriction.
- current_ip shows your own address. The environment variable didn't reach the server. Check the spelling of MOBILE_PROXY_URL in the config and make sure the client was restarted.
- rotate_ip returns status 429 or a limit message. You're rotating the IP more often than the plan allows. Increase the interval.
Step 6: Make the server reliable: retries, delays, cache, and limits
Goal of this stage: turn a tutorial example into a tool that doesn't crash on the first network error, doesn't hammer sites with requests, and doesn't overflow the model's context. This is the last mandatory step before real-world use.
What we're adding and why
- Automatic retries. Network errors happen. Instead of immediately returning an error to the agent, we'll try the request two more times with a pause.
- Auto IP rotation on blocks. If the site responds with 429 or 403 and a rotation link is configured, the server changes the address and retries the request itself.
- Delay between requests. A polite scraper doesn't send dozens of requests per second. A one- to two-second pause reduces the load on the site and the risk of being blocked.
- Cache. The agent often requests the same page multiple times using different tools. An in-memory cache for a few minutes saves repeated downloads.
- Size limit. We won't download pages larger than a few megabytes.
Code
Add import time at the top of the file, and replace the _get_html function with this:
CACHE: dict[str, tuple[float, str]] = {}
CACHE_TTL = 300
REQUEST_DELAY = 1.5
MAX_BYTES = 3_000_000
_last_request = 0.0
def _get_html(url: str) -> str:
global _last_request
now = time.time()
cached = CACHE.get(url)
if cached and now - cached[0] < CACHE_TTL:
log(f'cache hit: {url}')
return cached[1]
last_error = None
for attempt in range(3):
wait = REQUEST_DELAY - (time.time() - _last_request)
if wait > 0:
time.sleep(wait)
try:
with _client() as client:
_last_request = time.time()
response = client.get(url)
if response.status_code in (403, 429) and ROTATE_URL:
log(f'status {response.status_code}, rotating ip')
httpx.get(ROTATE_URL, timeout=15.0)
time.sleep(8)
continue
response.raise_for_status()
if len(response.content) > MAX_BYTES:
raise ValueError(f'Страница слишком большая: {len(response.content)} байт')
html = response.text
CACHE[url] = (time.time(), html)
return html
except httpx.HTTPError as error:
last_error = error
log(f'attempt {attempt + 1} failed: {error}')
time.sleep(2 * (attempt + 1))
raise RuntimeError(f'Не удалось загрузить {url} после 3 попыток: {last_error}')How it works
- The CACHE dictionary stores the timestamp and HTML for each address. If a page was requested less than five minutes ago, we return the saved copy without making a request.
- Before each request, we compute how much time has passed since the last one and top up the pause to REQUEST_DELAY seconds if needed.
- A three-attempt loop. On a 403 or 429 with rotation configured, the server changes the IP, waits eight seconds, and tries again. On network errors, it waits two, four, six seconds between attempts.
- If the page is larger than three megabytes, we treat it as an error: such documents won't fit in the context anyway.
- After three failures, we raise a clear error with the address and reason. The agent gets it as text and can tell you or try another path.
We also recommend adding a cache-clearing tool so the agent can force-reload a page:
@mcp.tool()
def clear_cache() -> str:
'''Очищает кэш загруженных страниц. Вызывай, если нужно получить свежую версию страницы.'''
count = len(CACHE)
CACHE.clear()
return f'Кэш очищен, удалено записей: {count}'Tip: It's worth moving REQUEST_DELAY and CACHE_TTL into environment variables, following the proxy pattern, so you can change them without editing code. For price monitoring, a two- to three-second delay and a one-minute cache work well; for article scraping, a one-second delay and a one-hour cache.
✅ Check: Call extract_text for the same page twice in a row. The second time, the logs show a cache hit line, and the response comes back instantly. Enter a non-existent domain — after a few seconds, the agent gets a "Failed to load ... after 3 attempts" message instead of hanging.
Possible problems
- NameError: ROTATE_URL is not defined. The _get_html function is declared above the proxy settings block. Move the PROXY_URL and ROTATE_URL settings higher in the file.
- The agent complains about slow performance. That's normal: delays and IP rotation take time. If you're in a hurry, lower REQUEST_DELAY to 0.5, but keep the risk of blocks in mind.
- Memory grows. The cache stores every page for the session. For long sessions, add cleanup of entries older than TTL on each call or cap the dictionary size.
Checking the result: a checklist for your finished MCP server
Go through the checklist and tick off each item. If all of them hold, your MCP server for web scraping is ready for real work.
What should work
- The mcp dev server.py command runs without errors, Inspector connects, and the status reads Connected.
- The tool list includes fetch_page, extract_text, extract_links, select_elements, current_ip, rotate_ip, and clear_cache.
- The web-collector server shows up in the AI client's tools panel without an error indicator.
- The agent picks the right tool on its own from a free-form request and calls it.
- current_ip shows the mobile proxy address, and after rotate_ip the address changes.
- A repeated request for the same page is served from cache.
- A bad address produces a clear error message instead of a hang.
End-to-end test
- Pick a public site with a catalog or article feed whose data you're allowed to use.
- Ask the agent: "Open the site's homepage, find links to the catalog section, go into the first five cards, collect the name and price, and format them as a table with columns Name, Price, Link."
- Watch the chain of calls: the agent should call extract_links with a filter, then select_elements or extract_text a few times, and build the table at the end.
- Verify a few rows by hand by opening the cards in your browser. The data should match.
Success metrics
Collecting five cards takes no more than 30-40 seconds including delays. The client logs contain no traceback-level errors. The agent doesn't ask which tool to use but acts on its own. If all that's true, congratulations: you've built your own MCP server and connected an AI agent to the web.
Common mistakes when building an MCP server and how to fix them
Here are the problems almost everyone runs into on the first pass. Format: problem, cause, solution.
1. The server connects in Inspector but doesn't work in the client
Cause: the client config points to the system python without the installed libraries, or the file path is wrong. Solution: specify the full path to the python inside .venv and the full path to server.py, use forward slashes, and fully restart the client.
2. The client drops the connection right after launch
Cause: a regular print without file=sys.stderr is still in the code, and the protocol's stdout stream got polluted. Solution: replace all print calls with the log function. Also check that the libraries don't write to stdout: some progress bars do this by default.
3. The agent doesn't call tools and answers from its own knowledge
Cause: the tool descriptions are too short or vague, and the model doesn't understand when to apply them. Solution: expand the docstrings, add "use when..." phrases and examples. In the first requests, name the tool explicitly.
4. 403 error when loading real sites
Cause: the site doesn't accept requests without browser headers or from a suspicious IP. Solution: check that HEADERS are being sent, update the User-Agent to the current browser version, connect a mobile proxy, and make sure rotation works.
5. Empty select_elements result with a correct selector
Cause: the data is loaded by JavaScript after the page loads, and it's not in the source HTML. Solution: check via fetch_page. If the data isn't there, try finding the site's internal API in the browser's Network tab: cards often arrive as JSON from a separate address you can request directly with the same extract_text.
6. 407 or ConnectTimeout error when working through a proxy
Cause: wrong credentials, unencoded special characters in the password, or the wrong port. Solution: copy the connection string from the dashboard again, encode the special characters, and check the http or socks5 protocol.
7. The JSON config isn't applied
Cause: an extra comma, a missing quote, or backslashes in the paths. Solution: validate the file, replace backslashes with forward slashes, and make sure there's no comma after the last element.
8. The server works, but the data comes back in the wrong encoding
Cause: the site doesn't specify the encoding in its headers. Solution: set response.encoding explicitly, or use the response.content attribute with manual decoding via decode('utf-8', errors='ignore').
Extra features: a section for advanced users
The basic server is ready. If you write Python confidently and want more, here are directions for growth, each of which can be implemented in an evening.
Remote server via Streamable HTTP
To make the server run on a separate machine or let multiple clients connect to it, replace the last line with mcp.run(transport='streamable-http'). By default, the server comes up on port 8000, and the connection address is http://machine-address:8000/mcp. In the client config, instead of command and args, use the url key with that address. In this mode, you can write to stdout, but it's better to keep the habit of logging to stderr. Be sure to close the port to the outside world and add header token validation if the server is reachable beyond the local network.
Resources and prompts
A resource with the current settings helps the agent understand the working context:
@mcp.resource('collector://settings')
def settings() -> str:
'''Текущие настройки сборщика.'''
return f'proxy: {"on" if PROXY_URL else "off"}, delay: {REQUEST_DELAY}, cache ttl: {CACHE_TTL}'A prompt provides a ready-made scenario the user invokes with a single command:
@mcp.prompt()
def price_monitor(url: str) -> str:
'''Сценарий мониторинга цен в каталоге.'''
return f'Открой {url}, собери ссылки на карточки товаров, зайди в каждую, вытащи название и цену и составь таблицу. Если увидишь ошибку 429, вызови rotate_ip и продолжи.'Saving results to a file
Add a save_csv tool that takes a list of dictionaries and a file path and writes the data with the csv module. The agent can then not just collect data but also put results into a table you can open in Excel. Restrict the save path to a single folder so the agent can't write anywhere it wants on the disk.
Async and parallel collection
FastMCP supports async functions: declare a tool with async def and use httpx.AsyncClient. Then a fetch_many tool can load ten pages at once via asyncio.gather. Don't forget a semaphore to limit the number of concurrent requests, and note that the delay between requests needs to be calculated differently with parallelism.
Multiple proxies and smart rotation
If you have several mobile proxies for different regions, store them in an environment variable as a comma-separated list and add a region parameter to the tool. The server will pick a proxy by region, and the agent can compare prices the site shows to users in different cities. This is one of the most in-demand tasks for marketers and affiliates.
Packaging in Docker
To run on a server, build an image on python:3.12-slim, copy server.py and the dependencies file, install the packages, and set the entry point with HTTP transport. Pass the proxy variables at container startup rather than baking them into the image.
⚠️ Warning: Never publish code with usernames, passwords, and rotation links in public repositories. Keep them only in environment variables or in a .env file added to .gitignore. A leaked IP rotation link lets outsiders control your proxy.
FAQ: common questions about building an MCP server
Can I build an MCP server in a language other than Python?
Yes. Official SDKs exist for TypeScript, Java, Kotlin, C#, and other languages. The principles are the same: declare tools with descriptions and launch the transport. Python was chosen in this guide for its simplicity and rich set of HTML-processing libraries.
Do I need a paid AI client plan to work with MCP?
Claude Desktop supports local MCP servers on the free plan too, but with message limits. Cursor and VS Code also let you connect servers. Check the current terms for your specific client.
Is a proxy mandatory?
No, the server works directly too. A proxy is needed when the request volume is noticeable, sites are sensitive to request frequency, or you need to see content from a specific region and from a mobile IP.
How do I know requests are actually going through the proxy?
Call the current_ip tool and compare the address with what your provider's dashboard shows. You can also ask the agent to load an IP-detection service page via extract_text.
How many tools can I add to one server?
Technically there are almost no limits, but every description takes up space in the model's context. Practice shows that 5-15 well-described tools work better than 50 tiny ones. Group related functions with parameters.
How do I update the server without restarting the client?
With stdio transport, the client launches the process at startup, so code changes only take effect after the client is restarted. During development, it's easier to test edits via Inspector and restart the client when you're done.
What if a site serves data only after JavaScript runs?
Our server works with the source HTML and won't see such data. Options: find the site's internal API in the browser's Network tab, or connect a browser engine. The second path is covered in other blog articles; we deliberately avoid it here.
How do I restrict the agent so it doesn't visit unwanted sites?
Add a domain check against a whitelist or blacklist from an environment variable in _get_html and return a clear error for forbidden addresses. That's more reliable than relying on chat instructions.
Can I use one MCP server from multiple clients at once?
With stdio, each client launches its own copy of the process, and that's fine: they don't interfere with each other, but their caches are separate. For a shared cache and a single proxy, switch to the HTTP transport from the advanced section.
Conclusion: what you built and where to go next
Let's sum up. You set up a Python environment and installed the official protocol SDK. You wrote an MCP server from scratch and figured out how the model understands tools through their descriptions. You connected the server to an AI client and saw the agent load pages on its own. You added tools for extracting text, links, and elements by selector. You routed traffic through a mobile proxy with IP rotation. Finally, you made the server resilient: retries, delays, cache, and limits. This is no longer a tutorial example but a working tool for everyday tasks.
What's next? Start using the server in real scenarios: monitoring competitor prices, collecting reviews, checking landing pages, analyzing content in your niche. Along the way, you'll realize which tools you're missing and add them following the existing pattern. Every new tool is a function with a clear description — nothing harder than that.
The next level is the advanced section: a remote server over HTTP, parallel collection, working with multiple proxies by region, and saving results to tables. And when you hit sites with dynamic content, check out the related blog articles on browser automation. The main thing is done: your AI agent is out on the web through your own MCP server, and you fully control how it does it.