Jarian Cottingham c0fa9148b7 Initial commit: Unlimited-OCR MCP for opencode
- mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool
- ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server)
- run-ocr-server.sh: llama-server launcher with tuned sampling flags
- skill/ocr-image: opencode skill teaching agents when/how to call it
- README with full install instructions
2026-08-22 06:24:57 +00:00

Unlimited-OCR MCP

Give opencode's text-only models the ability to read the images you post.

unlimited-ocr-mcp runs baidu/Unlimited-OCR — a 3B vision-language model — locally through llama.cpp and exposes it to opencode as an MCP tool. When you paste a screenshot into a chat that runs on a model without vision, opencode strips the image from the model's context but stores it in its local database. The ocr_image tool recovers that image and returns clean, transcribed text — so the conversation can continue as if the model had eyes.

you paste a screenshot into opencode
        │
        ▼
 opencode (text-only model) ── image stripped from context
        │                          │
        │                          ▼
        │                 opencode.db  (image stored, base64)
        │                          ▲
        │                          │ reads latest image part
        ▼                          │
 agent sees placeholder ──► ocr_image (MCP tool, stdio)
                                   │
                                   ▼
                     llama-server :11436  (Unlimited-OCR, GPU)
                                   │
                                   ▼
                    transcribed text back into the chat

Components

path what it is
mcp/server.py MCP stdio server. Zero dependencies (Python stdlib only). Exposes the ocr_image tool.
ocr-posted.py Same pipeline as a plain CLI — useful before the MCP server is registered, or from cron/scripts.
run-ocr-server.sh Launches the Unlimited-OCR llama-server with the recommended sampling flags. All paths overridable via env.
skill/ocr-image/SKILL.md An opencode skill that teaches the agent when and how to call the tool.

Requirements

  • llama.cpp with Unlimited-OCR support — a build including the DeepSeek-OCR / Unlimited-OCR mtmd support (PRs ggml-org/llama.cpp#17400, #24969, #25614). Builds from ~August 2026 onward include it. The llama-server binary must support --mmproj.
  • A GPU with ~5 GB free VRAM for the Q4_K_M quant (more for Q6/Q8/BF16). CPU-only works but is slow.
  • opencode with an active session (the tool reads ~/.local/share/opencode/opencode.db).
  • Python 3.8+ (stdlib only — no pip install needed).

Model download

Pre-quantized GGUFs (imatrix) from sahilchachra/Unlimited-OCR-GGUF. You need both files — the language model and the vision projector:

# New Hugging Face CLI syntax ("hf download", not "huggingface-cli download")
hf download sahilchachra/Unlimited-OCR-GGUF \
  --include "Unlimited-OCR-Q4_K_M.gguf" \
  --include "mmproj-Unlimited-OCR-F16.gguf" \
  --local-dir ~/models
quant size notes
Unlimited-OCR-Q4_K_M.gguf 1.8 GiB default, good quality/size
Unlimited-OCR-Q8_0.gguf 2.9 GiB near-lossless, fewer degenerate repetitions
Unlimited-OCR-BF16.gguf ~6 GiB full precision, overkill for most OCR

The mmproj-Unlimited-OCR-F16.gguf (vision projector, 0.77 GiB) is required with any quant.

1. Start the OCR server

# Point the env vars at your llama-server binary and model files
export LLAMA_SERVER_BIN=/path/to/llama-server
export OCR_MODEL=~/models/Unlimited-OCR-Q4_K_M.gguf
export OCR_MMPROJ=~/models/mmproj-Unlimited-OCR-F16.gguf
export CUDA_VISIBLE_DEVICES=1        # the GPU with free VRAM

./run-ocr-server.sh

Or launch llama-server directly:

llama-server \
  --model ~/models/Unlimited-OCR-Q4_K_M.gguf \
  --mmproj ~/models/mmproj-Unlimited-OCR-F16.gguf \
  --alias Unlimited-OCR \
  --host 127.0.0.1 --port 11436 \
  --temp 0 --repeat-penalty 1.2 \
  --ctx-size 32768 --n-gpu-layers 99 \
  --flash-attn on --cont-batching --parallel 1 \
  --cache-type-k q4_0 --cache-type-v q4_0

Verify:

curl -s http://127.0.0.1:11436/health          # → {"status":"ok"}
curl -s http://127.0.0.1:11436/v1/models       # → capabilities include "multimodal"

Sampling matters. --temp 0 --repeat-penalty 1.2 is the recipe this project is tuned around. 1.1 tends to loop infinitely; 1.3 starts skipping legitimately repeated rows.

2. Register the MCP server in opencode

Add to your opencode config (~/.config/opencode/opencode.json, global — or a project opencode.json):

{
  "mcp": {
    "ocr": {
      "type": "local",
      "command": ["python3", "/absolute/path/to/unlimited-ocr-mcp/mcp/server.py"],
      "enabled": true
    }
  }
}

Then restart opencode — MCP servers are loaded at startup. The tool appears as ocr_image.

Environment overrides (optional): OPENCODE_DB (path to opencode.db, default ~/.local/share/opencode/opencode.db) and OCR_URL (chat completions endpoint, default http://127.0.0.1:11436/v1/chat/completions).

3. Install the skill

Copy the skill into opencode's global skills directory so every project knows to use the tool:

mkdir -p ~/.config/opencode/skills
cp -r skill/ocr-image ~/.config/opencode/skills/

The skill triggers when an image is posted to a text-only model (the Cannot read <file> (this model does not support image input) placeholder) or when you ask "read the image I posted". It also documents the CLI fallback (ocr-posted.py) for sessions where the MCP server is not loaded yet.

Usage

Post any image into opencode and ask about it:

what does this screenshot say? OCR the invoice I sent, I need the total describe the chart in this image

The agent calls ocr_image with a prompt matched to your intent:

task prompt sent to the model
transcribe document/screenshot (default) Free OCR.
structured layout with coordinates <|grounding|>Convert the document to markdown.
visual description Describe this image in detail.
locate an element <|grounding|>Locate <|ref|>TEXT<|/ref|> in the image.

ocr_image arguments:

arg default meaning
target latest latest image part in opencode.db, a part id (prt_...), session:<id>, or a local image file path
prompt Free OCR. instruction sent to the OCR model
save — optional path to persist the decoded image

Raw output is layout-prefixed (header [x1,y1,x2,y2]text, table [...], <img> for image regions); the skill instructs the agent to strip the layout tokens and present clean text.

CLI fallback (identical pipeline, no MCP needed):

python3 ocr-posted.py latest "Free OCR."
python3 ocr-posted.py prt_0123abcd... "Describe this image in detail." /tmp/shot.png

How it works

  1. opencode stores every posted file part in its SQLite database (~/.local/share/opencode/opencode.db, table part). Image parts keep the full payload as a data: URL — even when the active model cannot see them.
  2. ocr_image queries that table for the newest image part (or a specific part/session), decodes it, and POSTs it to the llama-server /v1/chat/completions endpoint with temperature 0 and repeat_penalty 1.2.
  3. The transcription returns through MCP into the agent's context.

Troubleshooting

  • no image part found — no image was posted yet, or you're pointing at a different opencode data dir. Set OPENCODE_DB accordingly.
  • Connection refused on :11436 — the OCR server isn't running. Start it (section 1). Check GPU VRAM with nvidia-smi; a model sharing the card may need to be stopped first.
  • Garbled tail / repeated last row — known degenerate behavior of the 3B quant on dense tables. Re-run the same call; sampling varies per run. Q8_0 reduces it.
  • multimodal missing from /v1/models capabilities — your llama.cpp build predates Unlimited-OCR support, or --mmproj wasn't passed.
  • Tool not visible after config change — opencode loads MCP servers at startup; restart it.

Repository layout

unlimited-ocr-mcp/
├── README.md
├── LICENSE
├── mcp/server.py            # MCP stdio server (ocr_image tool)
├── ocr-posted.py            # CLI equivalent
├── run-ocr-server.sh        # llama-server launcher
└── skill/ocr-image/SKILL.md # opencode skill

License

MIT — see also the model card for baidu/Unlimited-OCR for the upstream model license.

Description
Run baidu/Unlimited-OCR locally via llama.cpp and give opencode text-only models the ability to read the images you post. MCP server, CLI, launcher, opencode skill.
Readme 34 KiB
Languages
Python 90.1%
Shell 9.9%