- mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool - ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server) - run-ocr-server.sh: llama-server launcher with tuned sampling flags - skill/ocr-image: opencode skill teaching agents when/how to call it - README with full install instructions
197 lines
8.4 KiB
Markdown
197 lines
8.4 KiB
Markdown
# Unlimited-OCR MCP
|
|
|
|
Give opencode's **text-only** models the ability to read the images you post.
|
|
|
|
`unlimited-ocr-mcp` runs [baidu/Unlimited-OCR](https://huggingface.co/baidu/Unlimited-OCR) — a 3B vision-language model — locally through [llama.cpp](https://github.com/ggml-org/llama.cpp) and exposes it to opencode as an MCP tool. When you paste a screenshot into a chat that runs on a model without vision, opencode strips the image from the model's context but stores it in its local database. The `ocr_image` tool recovers that image and returns clean, transcribed text — so the conversation can continue as if the model had eyes.
|
|
|
|
```
|
|
you paste a screenshot into opencode
|
|
│
|
|
▼
|
|
opencode (text-only model) ── image stripped from context
|
|
│ │
|
|
│ ▼
|
|
│ opencode.db (image stored, base64)
|
|
│ ▲
|
|
│ │ reads latest image part
|
|
▼ │
|
|
agent sees placeholder ──► ocr_image (MCP tool, stdio)
|
|
│
|
|
▼
|
|
llama-server :11436 (Unlimited-OCR, GPU)
|
|
│
|
|
▼
|
|
transcribed text back into the chat
|
|
```
|
|
|
|
## Components
|
|
|
|
| path | what it is |
|
|
|---|---|
|
|
| `mcp/server.py` | MCP stdio server. Zero dependencies (Python stdlib only). Exposes the `ocr_image` tool. |
|
|
| `ocr-posted.py` | Same pipeline as a plain CLI — useful before the MCP server is registered, or from cron/scripts. |
|
|
| `run-ocr-server.sh` | Launches the Unlimited-OCR llama-server with the recommended sampling flags. All paths overridable via env. |
|
|
| `skill/ocr-image/SKILL.md` | An [opencode skill](https://opencode.ai/docs/skills/) that teaches the agent when and how to call the tool. |
|
|
|
|
## Requirements
|
|
|
|
- **llama.cpp with Unlimited-OCR support** — a build including the DeepSeek-OCR / Unlimited-OCR mtmd support (PRs `ggml-org/llama.cpp#17400`, `#24969`, `#25614`). Builds from ~August 2026 onward include it. The `llama-server` binary must support `--mmproj`.
|
|
- **A GPU with ~5 GB free VRAM** for the Q4_K_M quant (more for Q6/Q8/BF16). CPU-only works but is slow.
|
|
- **opencode** with an active session (the tool reads `~/.local/share/opencode/opencode.db`).
|
|
- Python 3.8+ (stdlib only — no `pip install` needed).
|
|
|
|
## Model download
|
|
|
|
Pre-quantized GGUFs (imatrix) from [`sahilchachra/Unlimited-OCR-GGUF`](https://huggingface.co/sahilchachra/Unlimited-OCR-GGUF). You need **both** files — the language model and the vision projector:
|
|
|
|
```bash
|
|
# New Hugging Face CLI syntax ("hf download", not "huggingface-cli download")
|
|
hf download sahilchachra/Unlimited-OCR-GGUF \
|
|
--include "Unlimited-OCR-Q4_K_M.gguf" \
|
|
--include "mmproj-Unlimited-OCR-F16.gguf" \
|
|
--local-dir ~/models
|
|
```
|
|
|
|
| quant | size | notes |
|
|
|---|---|---|
|
|
| `Unlimited-OCR-Q4_K_M.gguf` | 1.8 GiB | default, good quality/size |
|
|
| `Unlimited-OCR-Q8_0.gguf` | 2.9 GiB | near-lossless, fewer degenerate repetitions |
|
|
| `Unlimited-OCR-BF16.gguf` | ~6 GiB | full precision, overkill for most OCR |
|
|
|
|
The `mmproj-Unlimited-OCR-F16.gguf` (vision projector, 0.77 GiB) is **required** with any quant.
|
|
|
|
## 1. Start the OCR server
|
|
|
|
```bash
|
|
# Point the env vars at your llama-server binary and model files
|
|
export LLAMA_SERVER_BIN=/path/to/llama-server
|
|
export OCR_MODEL=~/models/Unlimited-OCR-Q4_K_M.gguf
|
|
export OCR_MMPROJ=~/models/mmproj-Unlimited-OCR-F16.gguf
|
|
export CUDA_VISIBLE_DEVICES=1 # the GPU with free VRAM
|
|
|
|
./run-ocr-server.sh
|
|
```
|
|
|
|
Or launch `llama-server` directly:
|
|
|
|
```bash
|
|
llama-server \
|
|
--model ~/models/Unlimited-OCR-Q4_K_M.gguf \
|
|
--mmproj ~/models/mmproj-Unlimited-OCR-F16.gguf \
|
|
--alias Unlimited-OCR \
|
|
--host 127.0.0.1 --port 11436 \
|
|
--temp 0 --repeat-penalty 1.2 \
|
|
--ctx-size 32768 --n-gpu-layers 99 \
|
|
--flash-attn on --cont-batching --parallel 1 \
|
|
--cache-type-k q4_0 --cache-type-v q4_0
|
|
```
|
|
|
|
Verify:
|
|
|
|
```bash
|
|
curl -s http://127.0.0.1:11436/health # → {"status":"ok"}
|
|
curl -s http://127.0.0.1:11436/v1/models # → capabilities include "multimodal"
|
|
```
|
|
|
|
> **Sampling matters.** `--temp 0 --repeat-penalty 1.2` is the recipe this project is tuned around.
|
|
> `1.1` tends to loop infinitely; `1.3` starts skipping legitimately repeated rows.
|
|
|
|
## 2. Register the MCP server in opencode
|
|
|
|
Add to your opencode config (`~/.config/opencode/opencode.json`, global — or a project `opencode.json`):
|
|
|
|
```json
|
|
{
|
|
"mcp": {
|
|
"ocr": {
|
|
"type": "local",
|
|
"command": ["python3", "/absolute/path/to/unlimited-ocr-mcp/mcp/server.py"],
|
|
"enabled": true
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
Then **restart opencode** — MCP servers are loaded at startup. The tool appears as `ocr_image`.
|
|
|
|
Environment overrides (optional): `OPENCODE_DB` (path to `opencode.db`, default `~/.local/share/opencode/opencode.db`) and `OCR_URL` (chat completions endpoint, default `http://127.0.0.1:11436/v1/chat/completions`).
|
|
|
|
## 3. Install the skill
|
|
|
|
Copy the skill into opencode's global skills directory so every project knows to use the tool:
|
|
|
|
```bash
|
|
mkdir -p ~/.config/opencode/skills
|
|
cp -r skill/ocr-image ~/.config/opencode/skills/
|
|
```
|
|
|
|
The skill triggers when an image is posted to a text-only model (the
|
|
`Cannot read <file> (this model does not support image input)` placeholder) or when you ask
|
|
"read the image I posted". It also documents the CLI fallback (`ocr-posted.py`) for sessions
|
|
where the MCP server is not loaded yet.
|
|
|
|
## Usage
|
|
|
|
Post any image into opencode and ask about it:
|
|
|
|
> *what does this screenshot say?*
|
|
> *OCR the invoice I sent, I need the total*
|
|
> *describe the chart in this image*
|
|
|
|
The agent calls `ocr_image` with a prompt matched to your intent:
|
|
|
|
| task | prompt sent to the model |
|
|
|---|---|
|
|
| transcribe document/screenshot (default) | `Free OCR.` |
|
|
| structured layout with coordinates | `<\|grounding\|>Convert the document to markdown.` |
|
|
| visual description | `Describe this image in detail.` |
|
|
| locate an element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` |
|
|
|
|
`ocr_image` arguments:
|
|
|
|
| arg | default | meaning |
|
|
|---|---|---|
|
|
| `target` | `latest` | `latest` image part in `opencode.db`, a part id (`prt_...`), `session:<id>`, or a local image file path |
|
|
| `prompt` | `Free OCR.` | instruction sent to the OCR model |
|
|
| `save` | — | optional path to persist the decoded image |
|
|
|
|
Raw output is layout-prefixed (`header [x1,y1,x2,y2]text`, `table [...]`, `<img>` for image
|
|
regions); the skill instructs the agent to strip the layout tokens and present clean text.
|
|
|
|
CLI fallback (identical pipeline, no MCP needed):
|
|
|
|
```bash
|
|
python3 ocr-posted.py latest "Free OCR."
|
|
python3 ocr-posted.py prt_0123abcd... "Describe this image in detail." /tmp/shot.png
|
|
```
|
|
|
|
## How it works
|
|
|
|
1. opencode stores every posted file part in its SQLite database (`~/.local/share/opencode/opencode.db`, table `part`). Image parts keep the full payload as a `data:` URL — even when the active model cannot see them.
|
|
2. `ocr_image` queries that table for the newest image part (or a specific part/session), decodes it, and POSTs it to the llama-server `/v1/chat/completions` endpoint with `temperature 0` and `repeat_penalty 1.2`.
|
|
3. The transcription returns through MCP into the agent's context.
|
|
|
|
## Troubleshooting
|
|
|
|
- **`no image part found`** — no image was posted yet, or you're pointing at a different opencode data dir. Set `OPENCODE_DB` accordingly.
|
|
- **Connection refused on :11436** — the OCR server isn't running. Start it (section 1). Check GPU VRAM with `nvidia-smi`; a model sharing the card may need to be stopped first.
|
|
- **Garbled tail / repeated last row** — known degenerate behavior of the 3B quant on dense tables. Re-run the same call; sampling varies per run. Q8_0 reduces it.
|
|
- **`multimodal` missing from `/v1/models` capabilities** — your llama.cpp build predates Unlimited-OCR support, or `--mmproj` wasn't passed.
|
|
- **Tool not visible after config change** — opencode loads MCP servers at startup; restart it.
|
|
|
|
## Repository layout
|
|
|
|
```
|
|
unlimited-ocr-mcp/
|
|
├── README.md
|
|
├── LICENSE
|
|
├── mcp/server.py # MCP stdio server (ocr_image tool)
|
|
├── ocr-posted.py # CLI equivalent
|
|
├── run-ocr-server.sh # llama-server launcher
|
|
└── skill/ocr-image/SKILL.md # opencode skill
|
|
```
|
|
|
|
## License
|
|
|
|
[MIT](LICENSE) — see also the model card for `baidu/Unlimited-OCR` for the upstream model license.
|