refactor: split into google-mcp and duckduckgo-mcp
Some checks are pending
CI / lint (push) Waiting to run
CI / test (push) Waiting to run
CI / docker-build (push) Waiting to run
CI / security (push) Waiting to run
CI / build-result (push) Blocked by required conditions

Move the DuckDuckGo MCP server into its own repository. This repo
keeps the Google MCP server, shared search library, and tests.
Compose trimmed to the Google service.
This commit is contained in:
Jarian Cottingham 2026-08-21 18:37:53 +00:00
parent fdc67d5887
commit 4bae588d05
5 changed files with 39 additions and 443 deletions

118
README.md
View File

@ -1,22 +1,23 @@
# MCP Search Servers # Google MCP Server
Model Context Protocol (MCP) servers that expose **Google** and **DuckDuckGo** search Model Context Protocol (MCP) server that exposes **Google** search as a tool for LLM clients (Claude Desktop, Cursor, VS Code, ...), backed by a self-hosted [SearXNG](https://github.com/searxng/searxng) metasearch instance.
as tools for LLM clients (Claude Desktop, Cursor, VS Code, ...), backed by a
self-hosted [SearXNG](https://github.com/searxng/searxng) metasearch instance.
SearXNG aggregates results from Google, DuckDuckGo, Brave, Wikipedia, and more with SearXNG aggregates results from Google, DuckDuckGo, Brave, Wikipedia, and more with its own request handling — no browser fingerprinting, no API keys, fully self-hosted.
its own request handling, so a single SearXNG container serves both MCP servers —
no browser fingerprinting, no API keys, fully self-hosted.
## Services Part of the MCP Search Servers project:
| Repo | What it is |
|------|------------|
| [duckduckgo-mcp](https://git.jarianc.com/jarianc/duckduckgo-mcp) | Same design for DuckDuckGo search |
## Service
| Service | Port | MCP tool | Description | | Service | Port | MCP tool | Description |
|-------------------|------|--------------------|--------------------------------------| |-------------|------|----------------|--------------------------------|
| `google-mcp`| 3001 | `google_search`| Google results via SearXNG | | `google-mcp`| 3001 | `google_search`| Google results via SearXNG |
| `duckduckgo-mcp` | 3002 | `duckduckgo_search`| DuckDuckGo results via SearXNG |
| `searxng` | — | — | Internal metasearch engine (not published) | | `searxng` | — | — | Internal metasearch engine (not published) |
Each server exposes: The server exposes:
- `/sse` + `/messages/` — MCP SSE transport - `/sse` + `/messages/` — MCP SSE transport
- `/search?q=...&num=N` — plain HTTP JSON endpoint - `/search?q=...&num=N` — plain HTTP JSON endpoint
@ -42,9 +43,7 @@ curl http://localhost:3001/health
curl "http://localhost:3001/search?q=python&num=3" curl "http://localhost:3001/search?q=python&num=3"
``` ```
Then point your LLM client at `http://localhost:3001/sse` (Google) and/or Then point your LLM client at `http://localhost:3001/sse`. Client configuration examples for Claude Desktop, Cursor/Windsurf, and VS Code are in [USAGE.md](USAGE.md).
`http://localhost:3002/sse` (DuckDuckGo). Client configuration examples for
Claude Desktop, Cursor/Windsurf, and VS Code are in [USAGE.md](USAGE.md).
## Tool Contract ## Tool Contract
@ -55,80 +54,35 @@ Claude Desktop, Cursor/Windsurf, and VS Code are in [USAGE.md](USAGE.md).
} }
``` ```
Returns numbered results with title, URL, and snippet (max 20 results, Returns numbered results with title, URL, and snippet (max 20 results, queries capped at 500 characters).
queries capped at 500 characters).
## Configuration ## Architecture
| Variable | Default | Description |
|----------------------|----------------------|------------------------------------------|
| `SEARXNG_SECRET` | — (required) | SearXNG session/CSRF secret |
| `RATE_LIMIT_SECONDS` | 5 / 3 | Minimum seconds between search requests |
| `MCP_PORT` | 3001 / 3002 | Port the MCP server listens on |
| `SEARXNG_URL` | `http://searxng:8080`| URL of the SearXNG instance |
## Security Model
- **Internal network** — SearXNG is on a `internal: true` Docker network; only the
two MCP containers can reach it. It is not published to the host.
- **Per-IP rate limiting** — HTTP endpoints allow 30 requests/minute per IP
(`/health` exempt); excess requests get `429`.
- **Request pacing** — each server also enforces a minimum interval between
upstream SearXNG queries to avoid throttling.
- **No auth on `/search` and `/sse`** — the servers assume they are bound to a
trusted network. Expose them publicly only behind an auth proxy.
- **Input validation** — query length capped, result counts clamped to 1..20,
malformed parameters return `400` instead of crashing.
## Optional: Playwright Scraping Fallback
`lib/google_search.py` and `lib/duckduckgo_search.py` contain a legacy
Playwright-based scraping path (single shared browser, stealth user agent,
HTML parsing with multiple fallback strategies). It is **not** used by the
production servers — Google aggressively blocks headless browsers and SearXNG
is the working path. If you want to run the scraper directly:
```bash
pip install playwright
playwright install chromium
``` ```
LLM client (Claude Desktop, Cursor, ...)
## Development | MCP / SSE
v
```bash +---------------+ +----------------+
python3 -m venv .venv && source .venv/bin/activate | google-mcp | ---- | SearXNG | --> Google, Brave, Wikipedia, ...
pip install -e ".[dev]" | (:3001) | | (internal) |
+---------------+ +----------------+
ruff check .
pytest tests/ -v
# End-to-end test (requires running containers)
python3 integration_test.py
``` ```
## Project Structure ## Project Structure
``` ```
├── docker-compose.yml # All 3 services .
├── searxng-settings.yml # SearXNG engine config ├── google-mcp/ # MCP server (SSE transport + tools)
├── google-mcp/ │ ├── server.py
│ ├── Dockerfile │ └── Dockerfile
│ └── server.py # Google MCP server (SearXNG-backed) ├── lib/ # Shared search engine implementations
├── duckduckgo-mcp/ │ ├── google_search.py
│ ├── Dockerfile │ ├── duckduckgo_search.py # also used by the duckduckgo-mcp sibling repo
│ └── server.py # DuckDuckGo MCP server (SearXNG-backed) │ ├── playwright_manager.py
├── lib/ │ └── rate_limiter.py
│ ├── playwright_manager.py # Single-instance browser manager (optional path) ├── tests/ # Unit tests (rate limiter, result parsing)
│ ├── rate_limiter.py # Async minimum-interval rate limiter ├── integration_test.py
│ ├── google_search.py # Playwright Google scraper (optional path) ├── test_client.py
│ └── duckduckgo_search.py # Playwright DuckDuckGo scraper (optional path) ├── searxng-settings.yml
├── tests/ └── docker-compose.yml
│ ├── test_rate_limiter.py
│ └── test_search_parsing.py
├── integration_test.py # End-to-end MCP SSE test
└── test_client.py # Manual stdio/SSE smoke client
``` ```
## License
[MIT](LICENSE)

View File

@ -49,33 +49,6 @@ services:
retries: 3 retries: 3
start_period: 10s start_period: 10s
duckduckgo-search:
build:
context: .
dockerfile: duckduckgo-mcp/Dockerfile
container_name: duckduckgo-mcp
ports:
- "3002:3002"
environment:
- RATE_LIMIT_SECONDS=3
- MCP_PORT=3002
- SEARXNG_URL=http://searxng:8080
- PYTHONUNBUFFERED=1
depends_on:
searxng:
condition: service_healthy
restart: unless-stopped
stop_grace_period: 5s
networks:
- mcp-internal
- mcp-external
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:3002/health')"]
interval: 30s
timeout: 10s
retries: 3
start_period: 10s
networks: networks:
mcp-internal: mcp-internal:
internal: true internal: true

View File

@ -1,18 +0,0 @@
FROM python:3.12-slim
ENV DEBIAN_FRONTEND=noninteractive
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY lib/ ./lib/
COPY duckduckgo-mcp/ ./duckduckgo-mcp/
ENV PYTHONPATH=/app
ENV RATE_LIMIT_SECONDS=3
ENV SEARXNG_URL=http://searxng:8080
ENV MCP_PORT=3002
CMD ["python", "-m", "duckduckgo-mcp.server"]

View File

@ -1,313 +0,0 @@
import asyncio
import logging
import os
import signal
import sys
from collections import defaultdict
from contextlib import asynccontextmanager
from datetime import datetime, timezone
import httpx
from mcp.server import Server
from mcp.server.sse import SseServerTransport
from mcp.types import TextContent, Tool
from starlette.applications import Starlette
from starlette.middleware import Middleware
from starlette.middleware.base import BaseHTTPMiddleware
from starlette.requests import Request
from starlette.responses import JSONResponse, Response
from starlette.routing import Mount, Route
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s [%(levelname)s] %(message)s",
stream=sys.stderr,
)
logger = logging.getLogger("duckduckgo-mcp")
app = Server("duckduckgo-search")
PORT = int(os.environ.get("MCP_PORT", "3002"))
SEARXNG_URL = os.environ.get("SEARXNG_URL", "http://searxng:8080")
RATE_LIMIT_SECONDS = float(os.environ.get("RATE_LIMIT_SECONDS", "3"))
DEFAULT_RESULTS = 10
MAX_RESULTS = 20
MAX_QUERY_LEN = 500
def parse_num_results(value) -> int:
"""Parse a result-count value, clamped to 1..MAX_RESULTS."""
try:
num = int(value)
except (TypeError, ValueError):
return DEFAULT_RESULTS
return max(1, min(num, MAX_RESULTS))
# --- HTTPX singleton (#8) ---
_httpx_client: httpx.AsyncClient | None = None
async def get_httpx_client() -> httpx.AsyncClient:
global _httpx_client
if _httpx_client is None:
_httpx_client = httpx.AsyncClient(
timeout=httpx.Timeout(15.0),
limits=httpx.Limits(max_connections=32, max_keepalive_connections=16),
)
return _httpx_client
async def close_httpx_client():
global _httpx_client
if _httpx_client is not None:
await _httpx_client.aclose()
_httpx_client = None
@asynccontextmanager
async def lifespan(app):
logger.info("HTTPX client initialized")
try:
yield
finally:
await close_httpx_client()
logger.info("HTTPX client closed")
# --- Rate limiter ---
class RateLimiter:
def __init__(self, min_interval: float):
self.min_interval = min_interval
self._last_request = None
self._lock = asyncio.Lock()
async def acquire(self):
async with self._lock:
now = asyncio.get_running_loop().time()
if self._last_request:
elapsed = now - self._last_request
wait = self.min_interval - elapsed
if wait > 0:
logger.info(f"Rate limit: waiting {wait:.1f}s")
await asyncio.sleep(wait)
self._last_request = asyncio.get_running_loop().time()
rate_limiter = RateLimiter(RATE_LIMIT_SECONDS)
# --- Per-IP rate limiting middleware (#3) ---
class RateLimitMiddleware(BaseHTTPMiddleware):
"""Per-IP rate limiting for HTTP endpoints."""
def __init__(self, app, max_requests: int = 30, window_seconds: int = 60):
super().__init__(app)
self.max_requests = max_requests
self.window_seconds = window_seconds
self._requests: dict[str, list[float]] = defaultdict(list)
self._lock = asyncio.Lock()
async def dispatch(self, request: Request, call_next):
if request.url.path == "/health":
return await call_next(request)
client_ip = request.client.host if request.client else "unknown"
now = datetime.now(timezone.utc).timestamp()
cutoff = now - self.window_seconds
async with self._lock:
if len(self._requests) > 1000:
for ip in [
ip
for ip, ts in self._requests.items()
if not ts or ts[-1] <= cutoff
]:
del self._requests[ip]
timestamps = self._requests[client_ip]
self._requests[client_ip] = [t for t in timestamps if t > cutoff]
if len(self._requests[client_ip]) >= self.max_requests:
return JSONResponse(
{"error": "Rate limit exceeded. Try again later."},
status_code=429,
)
self._requests[client_ip].append(now)
response = await call_next(request)
return response
# --- Search logic ---
async def do_search(query: str, num_results: int, engine: str):
"""Shared search logic."""
await rate_limiter.acquire()
client = await get_httpx_client()
resp = await client.get(
f"{SEARXNG_URL}/search",
params={
"q": query,
"format": "json",
"engines": engine,
"categories": "general",
"language": "en",
},
)
resp.raise_for_status()
data = resp.json()
results = data.get("results", [])[:num_results]
return [
{
"title": r.get("title", ""),
"url": r.get("url", ""),
"snippet": r.get("content", "")[:200],
}
for r in results
]
# --- HTTP endpoint ---
async def search_http(request: Request):
query = request.query_params.get("q", "").strip()
if not query:
return JSONResponse({"error": "'q' parameter required"}, status_code=400)
if len(query) > MAX_QUERY_LEN:
return JSONResponse(
{"error": f"'q' must be {MAX_QUERY_LEN} characters or fewer"},
status_code=400,
)
num = parse_num_results(request.query_params.get("num"))
try:
results = await do_search(query, num, "duckduckgo")
return JSONResponse({"query": query, "results": results})
except Exception:
logger.exception("Search failed")
return JSONResponse({"error": "Search service unavailable"}, status_code=503)
# --- Health check (#6) ---
async def health_check(request: Request):
return JSONResponse({"status": "ok", "service": "duckduckgo-mcp"})
# --- MCP tool handlers ---
@app.list_tools()
async def list_tools() -> list[Tool]:
return [
Tool(
name="duckduckgo_search",
description="Search DuckDuckGo and return results with titles, URLs, and snippets.",
inputSchema={
"type": "object",
"properties": {
"query": {"type": "string", "description": "The search query."},
"num_results": {
"type": "integer",
"description": "Maximum number of results (default 10).",
"default": 10,
},
},
"required": ["query"],
},
)
]
@app.call_tool()
async def call_tool(name: str, arguments: dict) -> list[TextContent]:
if name != "duckduckgo_search":
raise ValueError(f"Unknown tool: {name}")
query = arguments.get("query", "")
if not isinstance(query, str) or not query:
return [TextContent(type="text", text="Error: 'query' is required.")]
if len(query) > MAX_QUERY_LEN:
return [
TextContent(
type="text",
text=f"Error: query must be {MAX_QUERY_LEN} characters or fewer.",
)
]
num_results = parse_num_results(arguments.get("num_results"))
try:
results = await do_search(query, num_results, "duckduckgo")
except Exception:
logger.exception("Search failed")
return [TextContent(type="text", text="Search failed: service unavailable")]
if not results:
return [TextContent(type="text", text=f"No results found for: {query}")]
lines = [f"Search results for: {query}\n"]
for i, r in enumerate(results, 1):
lines.append(f"{i}. {r['title']}")
lines.append(f" URL: {r['url']}")
if r["snippet"]:
lines.append(f" {r['snippet']}")
lines.append("")
return [TextContent(type="text", text="\n".join(lines))]
# --- SSE MCP transport ---
sse = SseServerTransport("/messages/")
async def handle_sse(request):
async with sse.connect_sse(
request.scope, request.receive, request._send
) as (read_stream, write_stream):
await app.run(
read_stream,
write_stream,
app.create_initialization_options(),
)
return Response()
starlette_app = Starlette(
debug=False,
lifespan=lifespan,
routes=[
Route("/health", endpoint=health_check),
Route("/search", endpoint=search_http),
Route("/sse", endpoint=handle_sse),
Mount("/messages/", app=sse.handle_post_message),
],
middleware=[Middleware(RateLimitMiddleware)],
)
async def main():
logger.info(f"DuckDuckGo MCP server ready on port {PORT}")
import uvicorn
config = uvicorn.Config(starlette_app, host="0.0.0.0", port=PORT, log_level="info")
server = uvicorn.Server(config)
loop = asyncio.get_running_loop()
def handle_signal():
server.should_exit = True
for sig in (signal.SIGINT, signal.SIGTERM):
loop.add_signal_handler(sig, handle_signal)
await server.serve()
if __name__ == "__main__":
asyncio.run(main())