unlimited-ocr-mcp/README.md
Jarian Cottingham c0fa9148b7 Initial commit: Unlimited-OCR MCP for opencode
- mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool
- ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server)
- run-ocr-server.sh: llama-server launcher with tuned sampling flags
- skill/ocr-image: opencode skill teaching agents when/how to call it
- README with full install instructions
2026-08-22 06:24:57 +00:00

197 lines
8.4 KiB
Markdown

# Unlimited-OCR MCP
Give opencode's **text-only** models the ability to read the images you post.
`unlimited-ocr-mcp` runs [baidu/Unlimited-OCR](https://huggingface.co/baidu/Unlimited-OCR) — a 3B vision-language model — locally through [llama.cpp](https://github.com/ggml-org/llama.cpp) and exposes it to opencode as an MCP tool. When you paste a screenshot into a chat that runs on a model without vision, opencode strips the image from the model's context but stores it in its local database. The `ocr_image` tool recovers that image and returns clean, transcribed text — so the conversation can continue as if the model had eyes.
```
you paste a screenshot into opencode
│
▼
opencode (text-only model) ── image stripped from context
│ │
│ ▼
│ opencode.db (image stored, base64)
│ ▲
│ │ reads latest image part
▼ │
agent sees placeholder ──► ocr_image (MCP tool, stdio)
│
▼
llama-server :11436 (Unlimited-OCR, GPU)
│
▼
transcribed text back into the chat
```
## Components
| path | what it is |
|---|---|
| `mcp/server.py` | MCP stdio server. Zero dependencies (Python stdlib only). Exposes the `ocr_image` tool. |
| `ocr-posted.py` | Same pipeline as a plain CLI — useful before the MCP server is registered, or from cron/scripts. |
| `run-ocr-server.sh` | Launches the Unlimited-OCR llama-server with the recommended sampling flags. All paths overridable via env. |
| `skill/ocr-image/SKILL.md` | An [opencode skill](https://opencode.ai/docs/skills/) that teaches the agent when and how to call the tool. |
## Requirements
- **llama.cpp with Unlimited-OCR support** — a build including the DeepSeek-OCR / Unlimited-OCR mtmd support (PRs `ggml-org/llama.cpp#17400`, `#24969`, `#25614`). Builds from ~August 2026 onward include it. The `llama-server` binary must support `--mmproj`.
- **A GPU with ~5 GB free VRAM** for the Q4_K_M quant (more for Q6/Q8/BF16). CPU-only works but is slow.
- **opencode** with an active session (the tool reads `~/.local/share/opencode/opencode.db`).
- Python 3.8+ (stdlib only — no `pip install` needed).
## Model download
Pre-quantized GGUFs (imatrix) from [`sahilchachra/Unlimited-OCR-GGUF`](https://huggingface.co/sahilchachra/Unlimited-OCR-GGUF). You need **both** files — the language model and the vision projector:
```bash
# New Hugging Face CLI syntax ("hf download", not "huggingface-cli download")
hf download sahilchachra/Unlimited-OCR-GGUF \
--include "Unlimited-OCR-Q4_K_M.gguf" \
--include "mmproj-Unlimited-OCR-F16.gguf" \
--local-dir ~/models
```
| quant | size | notes |
|---|---|---|
| `Unlimited-OCR-Q4_K_M.gguf` | 1.8 GiB | default, good quality/size |
| `Unlimited-OCR-Q8_0.gguf` | 2.9 GiB | near-lossless, fewer degenerate repetitions |
| `Unlimited-OCR-BF16.gguf` | ~6 GiB | full precision, overkill for most OCR |
The `mmproj-Unlimited-OCR-F16.gguf` (vision projector, 0.77 GiB) is **required** with any quant.
## 1. Start the OCR server
```bash
# Point the env vars at your llama-server binary and model files
export LLAMA_SERVER_BIN=/path/to/llama-server
export OCR_MODEL=~/models/Unlimited-OCR-Q4_K_M.gguf
export OCR_MMPROJ=~/models/mmproj-Unlimited-OCR-F16.gguf
export CUDA_VISIBLE_DEVICES=1 # the GPU with free VRAM
./run-ocr-server.sh
```
Or launch `llama-server` directly:
```bash
llama-server \
--model ~/models/Unlimited-OCR-Q4_K_M.gguf \
--mmproj ~/models/mmproj-Unlimited-OCR-F16.gguf \
--alias Unlimited-OCR \
--host 127.0.0.1 --port 11436 \
--temp 0 --repeat-penalty 1.2 \
--ctx-size 32768 --n-gpu-layers 99 \
--flash-attn on --cont-batching --parallel 1 \
--cache-type-k q4_0 --cache-type-v q4_0
```
Verify:
```bash
curl -s http://127.0.0.1:11436/health # → {"status":"ok"}
curl -s http://127.0.0.1:11436/v1/models # → capabilities include "multimodal"
```
> **Sampling matters.** `--temp 0 --repeat-penalty 1.2` is the recipe this project is tuned around.
> `1.1` tends to loop infinitely; `1.3` starts skipping legitimately repeated rows.
## 2. Register the MCP server in opencode
Add to your opencode config (`~/.config/opencode/opencode.json`, global — or a project `opencode.json`):
```json
{
"mcp": {
"ocr": {
"type": "local",
"command": ["python3", "/absolute/path/to/unlimited-ocr-mcp/mcp/server.py"],
"enabled": true
}
}
}
```
Then **restart opencode** — MCP servers are loaded at startup. The tool appears as `ocr_image`.
Environment overrides (optional): `OPENCODE_DB` (path to `opencode.db`, default `~/.local/share/opencode/opencode.db`) and `OCR_URL` (chat completions endpoint, default `http://127.0.0.1:11436/v1/chat/completions`).
## 3. Install the skill
Copy the skill into opencode's global skills directory so every project knows to use the tool:
```bash
mkdir -p ~/.config/opencode/skills
cp -r skill/ocr-image ~/.config/opencode/skills/
```
The skill triggers when an image is posted to a text-only model (the
`Cannot read <file> (this model does not support image input)` placeholder) or when you ask
"read the image I posted". It also documents the CLI fallback (`ocr-posted.py`) for sessions
where the MCP server is not loaded yet.
## Usage
Post any image into opencode and ask about it:
> *what does this screenshot say?*
> *OCR the invoice I sent, I need the total*
> *describe the chart in this image*
The agent calls `ocr_image` with a prompt matched to your intent:
| task | prompt sent to the model |
|---|---|
| transcribe document/screenshot (default) | `Free OCR.` |
| structured layout with coordinates | `<\|grounding\|>Convert the document to markdown.` |
| visual description | `Describe this image in detail.` |
| locate an element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` |
`ocr_image` arguments:
| arg | default | meaning |
|---|---|---|
| `target` | `latest` | `latest` image part in `opencode.db`, a part id (`prt_...`), `session:<id>`, or a local image file path |
| `prompt` | `Free OCR.` | instruction sent to the OCR model |
| `save` | — | optional path to persist the decoded image |
Raw output is layout-prefixed (`header [x1,y1,x2,y2]text`, `table [...]`, `<img>` for image
regions); the skill instructs the agent to strip the layout tokens and present clean text.
CLI fallback (identical pipeline, no MCP needed):
```bash
python3 ocr-posted.py latest "Free OCR."
python3 ocr-posted.py prt_0123abcd... "Describe this image in detail." /tmp/shot.png
```
## How it works
1. opencode stores every posted file part in its SQLite database (`~/.local/share/opencode/opencode.db`, table `part`). Image parts keep the full payload as a `data:` URL — even when the active model cannot see them.
2. `ocr_image` queries that table for the newest image part (or a specific part/session), decodes it, and POSTs it to the llama-server `/v1/chat/completions` endpoint with `temperature 0` and `repeat_penalty 1.2`.
3. The transcription returns through MCP into the agent's context.
## Troubleshooting
- **`no image part found`** — no image was posted yet, or you're pointing at a different opencode data dir. Set `OPENCODE_DB` accordingly.
- **Connection refused on :11436** — the OCR server isn't running. Start it (section 1). Check GPU VRAM with `nvidia-smi`; a model sharing the card may need to be stopped first.
- **Garbled tail / repeated last row** — known degenerate behavior of the 3B quant on dense tables. Re-run the same call; sampling varies per run. Q8_0 reduces it.
- **`multimodal` missing from `/v1/models` capabilities** — your llama.cpp build predates Unlimited-OCR support, or `--mmproj` wasn't passed.
- **Tool not visible after config change** — opencode loads MCP servers at startup; restart it.
## Repository layout
```
unlimited-ocr-mcp/
├── README.md
├── LICENSE
├── mcp/server.py # MCP stdio server (ocr_image tool)
├── ocr-posted.py # CLI equivalent
├── run-ocr-server.sh # llama-server launcher
└── skill/ocr-image/SKILL.md # opencode skill
```
## License
[MIT](LICENSE) — see also the model card for `baidu/Unlimited-OCR` for the upstream model license.