# Unlimited-OCR MCP Give opencode's **text-only** models the ability to read the images you post. `unlimited-ocr-mcp` runs [baidu/Unlimited-OCR](https://huggingface.co/baidu/Unlimited-OCR) — a 3B vision-language model — locally through [llama.cpp](https://github.com/ggml-org/llama.cpp) and exposes it to opencode as an MCP tool. When you paste a screenshot into a chat that runs on a model without vision, opencode strips the image from the model's context but stores it in its local database. The `ocr_image` tool recovers that image and returns clean, transcribed text — so the conversation can continue as if the model had eyes. ``` you paste a screenshot into opencode │ ▼ opencode (text-only model) ── image stripped from context │ │ │ ▼ │ opencode.db (image stored, base64) │ ▲ │ │ reads latest image part ▼ │ agent sees placeholder ──► ocr_image (MCP tool, stdio) │ ▼ llama-server :11436 (Unlimited-OCR, GPU) │ ▼ transcribed text back into the chat ``` ## Components | path | what it is | |---|---| | `mcp/server.py` | MCP stdio server. Zero dependencies (Python stdlib only). Exposes the `ocr_image` tool. | | `ocr-posted.py` | Same pipeline as a plain CLI — useful before the MCP server is registered, or from cron/scripts. | | `run-ocr-server.sh` | Launches the Unlimited-OCR llama-server with the recommended sampling flags. All paths overridable via env. | | `skill/ocr-image/SKILL.md` | An [opencode skill](https://opencode.ai/docs/skills/) that teaches the agent when and how to call the tool. | ## Requirements - **llama.cpp with Unlimited-OCR support** — a build including the DeepSeek-OCR / Unlimited-OCR mtmd support (PRs `ggml-org/llama.cpp#17400`, `#24969`, `#25614`). Builds from ~August 2026 onward include it. The `llama-server` binary must support `--mmproj`. - **A GPU with ~5 GB free VRAM** for the Q4_K_M quant (more for Q6/Q8/BF16). CPU-only works but is slow. - **opencode** with an active session (the tool reads `~/.local/share/opencode/opencode.db`). - Python 3.8+ (stdlib only — no `pip install` needed). ## Model download Pre-quantized GGUFs (imatrix) from [`sahilchachra/Unlimited-OCR-GGUF`](https://huggingface.co/sahilchachra/Unlimited-OCR-GGUF). You need **both** files — the language model and the vision projector: ```bash # New Hugging Face CLI syntax ("hf download", not "huggingface-cli download") hf download sahilchachra/Unlimited-OCR-GGUF \ --include "Unlimited-OCR-Q4_K_M.gguf" \ --include "mmproj-Unlimited-OCR-F16.gguf" \ --local-dir ~/models ``` | quant | size | notes | |---|---|---| | `Unlimited-OCR-Q4_K_M.gguf` | 1.8 GiB | default, good quality/size | | `Unlimited-OCR-Q8_0.gguf` | 2.9 GiB | near-lossless, fewer degenerate repetitions | | `Unlimited-OCR-BF16.gguf` | ~6 GiB | full precision, overkill for most OCR | The `mmproj-Unlimited-OCR-F16.gguf` (vision projector, 0.77 GiB) is **required** with any quant. ## 1. Start the OCR server ```bash # Point the env vars at your llama-server binary and model files export LLAMA_SERVER_BIN=/path/to/llama-server export OCR_MODEL=~/models/Unlimited-OCR-Q4_K_M.gguf export OCR_MMPROJ=~/models/mmproj-Unlimited-OCR-F16.gguf export CUDA_VISIBLE_DEVICES=1 # the GPU with free VRAM ./run-ocr-server.sh ``` Or launch `llama-server` directly: ```bash llama-server \ --model ~/models/Unlimited-OCR-Q4_K_M.gguf \ --mmproj ~/models/mmproj-Unlimited-OCR-F16.gguf \ --alias Unlimited-OCR \ --host 127.0.0.1 --port 11436 \ --temp 0 --repeat-penalty 1.2 \ --ctx-size 32768 --n-gpu-layers 99 \ --flash-attn on --cont-batching --parallel 1 \ --cache-type-k q4_0 --cache-type-v q4_0 ``` Verify: ```bash curl -s http://127.0.0.1:11436/health # → {"status":"ok"} curl -s http://127.0.0.1:11436/v1/models # → capabilities include "multimodal" ``` > **Sampling matters.** `--temp 0 --repeat-penalty 1.2` is the recipe this project is tuned around. > `1.1` tends to loop infinitely; `1.3` starts skipping legitimately repeated rows. ## 2. Register the MCP server in opencode Add to your opencode config (`~/.config/opencode/opencode.json`, global — or a project `opencode.json`): ```json { "mcp": { "ocr": { "type": "local", "command": ["python3", "/absolute/path/to/unlimited-ocr-mcp/mcp/server.py"], "enabled": true } } } ``` Then **restart opencode** — MCP servers are loaded at startup. The tool appears as `ocr_image`. Environment overrides (optional): `OPENCODE_DB` (path to `opencode.db`, default `~/.local/share/opencode/opencode.db`) and `OCR_URL` (chat completions endpoint, default `http://127.0.0.1:11436/v1/chat/completions`). ## 3. Install the skill Copy the skill into opencode's global skills directory so every project knows to use the tool: ```bash mkdir -p ~/.config/opencode/skills cp -r skill/ocr-image ~/.config/opencode/skills/ ``` The skill triggers when an image is posted to a text-only model (the `Cannot read (this model does not support image input)` placeholder) or when you ask "read the image I posted". It also documents the CLI fallback (`ocr-posted.py`) for sessions where the MCP server is not loaded yet. ## Usage Post any image into opencode and ask about it: > *what does this screenshot say?* > *OCR the invoice I sent, I need the total* > *describe the chart in this image* The agent calls `ocr_image` with a prompt matched to your intent: | task | prompt sent to the model | |---|---| | transcribe document/screenshot (default) | `Free OCR.` | | structured layout with coordinates | `<\|grounding\|>Convert the document to markdown.` | | visual description | `Describe this image in detail.` | | locate an element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` | `ocr_image` arguments: | arg | default | meaning | |---|---|---| | `target` | `latest` | `latest` image part in `opencode.db`, a part id (`prt_...`), `session:`, or a local image file path | | `prompt` | `Free OCR.` | instruction sent to the OCR model | | `save` | — | optional path to persist the decoded image | Raw output is layout-prefixed (`header [x1,y1,x2,y2]text`, `table [...]`, `` for image regions); the skill instructs the agent to strip the layout tokens and present clean text. CLI fallback (identical pipeline, no MCP needed): ```bash python3 ocr-posted.py latest "Free OCR." python3 ocr-posted.py prt_0123abcd... "Describe this image in detail." /tmp/shot.png ``` ## How it works 1. opencode stores every posted file part in its SQLite database (`~/.local/share/opencode/opencode.db`, table `part`). Image parts keep the full payload as a `data:` URL — even when the active model cannot see them. 2. `ocr_image` queries that table for the newest image part (or a specific part/session), decodes it, and POSTs it to the llama-server `/v1/chat/completions` endpoint with `temperature 0` and `repeat_penalty 1.2`. 3. The transcription returns through MCP into the agent's context. ## Troubleshooting - **`no image part found`** — no image was posted yet, or you're pointing at a different opencode data dir. Set `OPENCODE_DB` accordingly. - **Connection refused on :11436** — the OCR server isn't running. Start it (section 1). Check GPU VRAM with `nvidia-smi`; a model sharing the card may need to be stopped first. - **Garbled tail / repeated last row** — known degenerate behavior of the 3B quant on dense tables. Re-run the same call; sampling varies per run. Q8_0 reduces it. - **`multimodal` missing from `/v1/models` capabilities** — your llama.cpp build predates Unlimited-OCR support, or `--mmproj` wasn't passed. - **Tool not visible after config change** — opencode loads MCP servers at startup; restart it. ## Repository layout ``` unlimited-ocr-mcp/ ├── README.md ├── LICENSE ├── mcp/server.py # MCP stdio server (ocr_image tool) ├── ocr-posted.py # CLI equivalent ├── run-ocr-server.sh # llama-server launcher └── skill/ocr-image/SKILL.md # opencode skill ``` ## License [MIT](LICENSE) — see also the model card for `baidu/Unlimited-OCR` for the upstream model license.