Initial commit: Unlimited-OCR MCP for opencode

- mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool
- ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server)
- run-ocr-server.sh: llama-server launcher with tuned sampling flags
- skill/ocr-image: opencode skill teaching agents when/how to call it
- README with full install instructions
This commit is contained in:
Jarian Cottingham 2026-08-22 06:24:57 +00:00
commit c0fa9148b7
7 changed files with 590 additions and 0 deletions

3
.gitignore vendored Normal file
View File

@ -0,0 +1,3 @@
__pycache__/
*.pyc
.env

21
LICENSE Normal file
View File

@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Jarian Cottingham
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

196
README.md Normal file
View File

@ -0,0 +1,196 @@
# Unlimited-OCR MCP
Give opencode's **text-only** models the ability to read the images you post.
`unlimited-ocr-mcp` runs [baidu/Unlimited-OCR](https://huggingface.co/baidu/Unlimited-OCR) — a 3B vision-language model — locally through [llama.cpp](https://github.com/ggml-org/llama.cpp) and exposes it to opencode as an MCP tool. When you paste a screenshot into a chat that runs on a model without vision, opencode strips the image from the model's context but stores it in its local database. The `ocr_image` tool recovers that image and returns clean, transcribed text — so the conversation can continue as if the model had eyes.
```
you paste a screenshot into opencode
│
▼
opencode (text-only model) ── image stripped from context
│ │
│ ▼
│ opencode.db (image stored, base64)
│ ▲
│ │ reads latest image part
▼ │
agent sees placeholder ──► ocr_image (MCP tool, stdio)
│
▼
llama-server :11436 (Unlimited-OCR, GPU)
│
▼
transcribed text back into the chat
```
## Components
| path | what it is |
|---|---|
| `mcp/server.py` | MCP stdio server. Zero dependencies (Python stdlib only). Exposes the `ocr_image` tool. |
| `ocr-posted.py` | Same pipeline as a plain CLI — useful before the MCP server is registered, or from cron/scripts. |
| `run-ocr-server.sh` | Launches the Unlimited-OCR llama-server with the recommended sampling flags. All paths overridable via env. |
| `skill/ocr-image/SKILL.md` | An [opencode skill](https://opencode.ai/docs/skills/) that teaches the agent when and how to call the tool. |
## Requirements
- **llama.cpp with Unlimited-OCR support** — a build including the DeepSeek-OCR / Unlimited-OCR mtmd support (PRs `ggml-org/llama.cpp#17400`, `#24969`, `#25614`). Builds from ~August 2026 onward include it. The `llama-server` binary must support `--mmproj`.
- **A GPU with ~5 GB free VRAM** for the Q4_K_M quant (more for Q6/Q8/BF16). CPU-only works but is slow.
- **opencode** with an active session (the tool reads `~/.local/share/opencode/opencode.db`).
- Python 3.8+ (stdlib only — no `pip install` needed).
## Model download
Pre-quantized GGUFs (imatrix) from [`sahilchachra/Unlimited-OCR-GGUF`](https://huggingface.co/sahilchachra/Unlimited-OCR-GGUF). You need **both** files — the language model and the vision projector:
```bash
# New Hugging Face CLI syntax ("hf download", not "huggingface-cli download")
hf download sahilchachra/Unlimited-OCR-GGUF \
--include "Unlimited-OCR-Q4_K_M.gguf" \
--include "mmproj-Unlimited-OCR-F16.gguf" \
--local-dir ~/models
```
| quant | size | notes |
|---|---|---|
| `Unlimited-OCR-Q4_K_M.gguf` | 1.8 GiB | default, good quality/size |
| `Unlimited-OCR-Q8_0.gguf` | 2.9 GiB | near-lossless, fewer degenerate repetitions |
| `Unlimited-OCR-BF16.gguf` | ~6 GiB | full precision, overkill for most OCR |
The `mmproj-Unlimited-OCR-F16.gguf` (vision projector, 0.77 GiB) is **required** with any quant.
## 1. Start the OCR server
```bash
# Point the env vars at your llama-server binary and model files
export LLAMA_SERVER_BIN=/path/to/llama-server
export OCR_MODEL=~/models/Unlimited-OCR-Q4_K_M.gguf
export OCR_MMPROJ=~/models/mmproj-Unlimited-OCR-F16.gguf
export CUDA_VISIBLE_DEVICES=1 # the GPU with free VRAM
./run-ocr-server.sh
```
Or launch `llama-server` directly:
```bash
llama-server \
--model ~/models/Unlimited-OCR-Q4_K_M.gguf \
--mmproj ~/models/mmproj-Unlimited-OCR-F16.gguf \
--alias Unlimited-OCR \
--host 127.0.0.1 --port 11436 \
--temp 0 --repeat-penalty 1.2 \
--ctx-size 32768 --n-gpu-layers 99 \
--flash-attn on --cont-batching --parallel 1 \
--cache-type-k q4_0 --cache-type-v q4_0
```
Verify:
```bash
curl -s http://127.0.0.1:11436/health # → {"status":"ok"}
curl -s http://127.0.0.1:11436/v1/models # → capabilities include "multimodal"
```
> **Sampling matters.** `--temp 0 --repeat-penalty 1.2` is the recipe this project is tuned around.
> `1.1` tends to loop infinitely; `1.3` starts skipping legitimately repeated rows.
## 2. Register the MCP server in opencode
Add to your opencode config (`~/.config/opencode/opencode.json`, global — or a project `opencode.json`):
```json
{
"mcp": {
"ocr": {
"type": "local",
"command": ["python3", "/absolute/path/to/unlimited-ocr-mcp/mcp/server.py"],
"enabled": true
}
}
}
```
Then **restart opencode** — MCP servers are loaded at startup. The tool appears as `ocr_image`.
Environment overrides (optional): `OPENCODE_DB` (path to `opencode.db`, default `~/.local/share/opencode/opencode.db`) and `OCR_URL` (chat completions endpoint, default `http://127.0.0.1:11436/v1/chat/completions`).
## 3. Install the skill
Copy the skill into opencode's global skills directory so every project knows to use the tool:
```bash
mkdir -p ~/.config/opencode/skills
cp -r skill/ocr-image ~/.config/opencode/skills/
```
The skill triggers when an image is posted to a text-only model (the
`Cannot read <file> (this model does not support image input)` placeholder) or when you ask
"read the image I posted". It also documents the CLI fallback (`ocr-posted.py`) for sessions
where the MCP server is not loaded yet.
## Usage
Post any image into opencode and ask about it:
> *what does this screenshot say?*
> *OCR the invoice I sent, I need the total*
> *describe the chart in this image*
The agent calls `ocr_image` with a prompt matched to your intent:
| task | prompt sent to the model |
|---|---|
| transcribe document/screenshot (default) | `Free OCR.` |
| structured layout with coordinates | `<\|grounding\|>Convert the document to markdown.` |
| visual description | `Describe this image in detail.` |
| locate an element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` |
`ocr_image` arguments:
| arg | default | meaning |
|---|---|---|
| `target` | `latest` | `latest` image part in `opencode.db`, a part id (`prt_...`), `session:<id>`, or a local image file path |
| `prompt` | `Free OCR.` | instruction sent to the OCR model |
| `save` | — | optional path to persist the decoded image |
Raw output is layout-prefixed (`header [x1,y1,x2,y2]text`, `table [...]`, `<img>` for image
regions); the skill instructs the agent to strip the layout tokens and present clean text.
CLI fallback (identical pipeline, no MCP needed):
```bash
python3 ocr-posted.py latest "Free OCR."
python3 ocr-posted.py prt_0123abcd... "Describe this image in detail." /tmp/shot.png
```
## How it works
1. opencode stores every posted file part in its SQLite database (`~/.local/share/opencode/opencode.db`, table `part`). Image parts keep the full payload as a `data:` URL — even when the active model cannot see them.
2. `ocr_image` queries that table for the newest image part (or a specific part/session), decodes it, and POSTs it to the llama-server `/v1/chat/completions` endpoint with `temperature 0` and `repeat_penalty 1.2`.
3. The transcription returns through MCP into the agent's context.
## Troubleshooting
- **`no image part found`** — no image was posted yet, or you're pointing at a different opencode data dir. Set `OPENCODE_DB` accordingly.
- **Connection refused on :11436** — the OCR server isn't running. Start it (section 1). Check GPU VRAM with `nvidia-smi`; a model sharing the card may need to be stopped first.
- **Garbled tail / repeated last row** — known degenerate behavior of the 3B quant on dense tables. Re-run the same call; sampling varies per run. Q8_0 reduces it.
- **`multimodal` missing from `/v1/models` capabilities** — your llama.cpp build predates Unlimited-OCR support, or `--mmproj` wasn't passed.
- **Tool not visible after config change** — opencode loads MCP servers at startup; restart it.
## Repository layout
```
unlimited-ocr-mcp/
├── README.md
├── LICENSE
├── mcp/server.py # MCP stdio server (ocr_image tool)
├── ocr-posted.py # CLI equivalent
├── run-ocr-server.sh # llama-server launcher
└── skill/ocr-image/SKILL.md # opencode skill
```
## License
[MIT](LICENSE) — see also the model card for `baidu/Unlimited-OCR` for the upstream model license.

190
mcp/server.py Executable file
View File

@ -0,0 +1,190 @@
#!/usr/bin/env python3
"""MCP stdio server: OCR images posted into opencode via a local Unlimited-OCR backend.
Transport: newline-delimited JSON-RPC over stdin/stdout (MCP stdio).
Stdlib only. Env overrides: OPENCODE_DB, OCR_URL.
"""
import base64
import json
import os
import sqlite3
import sys
import urllib.request
DB = os.environ.get(
'OPENCODE_DB',
os.path.expanduser('~/.local/share/opencode/opencode.db'))
OCR_URL = os.environ.get('OCR_URL', 'http://127.0.0.1:11436/v1/chat/completions')
SERVER_INFO = {'name': 'ocr-posted', 'version': '1.0.0'}
TOOLS = [
{
'name': 'ocr_image',
'description': (
'OCR an image the user posted into opencode (or any local image file) using the '
'local Unlimited-OCR VLM server. Use this when the user attaches or pastes an image '
'that you cannot see directly, e.g. when the message shows '
"'Cannot read <file> (this model does not support image input)'. "
'Returns transcribed text (layout-prefixed lines with prompts that ask for grounding).'
),
'inputSchema': {
'type': 'object',
'properties': {
'target': {
'type': 'string',
'description': (
"'latest' (default): most recent image part in opencode.db. "
"Or a part id (prt_...), 'session:<session_id>', or a local image file path."
),
},
'prompt': {
'type': 'string',
'description': (
"Instruction sent to the OCR model. Default 'Free OCR.' (plain transcription). "
"Use '<|grounding|>Convert the document to markdown.' for structured layout, "
"'Describe this image in detail.' for visual QA."
),
},
'save': {
'type': 'string',
'description': 'Optional local path to save the decoded image to.',
},
},
'required': [],
},
}
]
def resolve_data_url(target):
"""Return (data_url, meta) for the requested target."""
if target and target != 'latest' and not target.startswith('session:'):
if os.path.isfile(target):
ext = os.path.splitext(target)[1].lower()
mime = {'.png': 'image/png', '.jpg': 'image/jpeg', '.jpeg': 'image/jpeg',
'.gif': 'image/gif', '.webp': 'image/webp'}.get(ext, 'image/png')
b64 = base64.b64encode(open(target, 'rb').read()).decode()
return 'data:%s;base64,%s' % (mime, b64), {'source': 'file', 'filename': target}
db = sqlite3.connect(DB, timeout=10)
row = db.execute('SELECT id, data FROM part WHERE id=?', (target,)).fetchone()
if not row:
raise ValueError('part id not found: %s' % target)
pid, data = row
else:
db = sqlite3.connect(DB, timeout=10)
if target and target.startswith('session:'):
row = db.execute(
"SELECT id, data FROM part WHERE session_id=? "
"AND data LIKE '%\"type\":\"file\"%' AND data LIKE '%data:image%' "
"ORDER BY time_created DESC LIMIT 1", (target[8:],)).fetchone()
else:
row = db.execute(
"SELECT id, data FROM part WHERE data LIKE '%\"type\":\"file\"%' "
"AND data LIKE '%data:image%' ORDER BY time_created DESC LIMIT 1").fetchone()
if not row:
scope = ' in session %s' % target[8:] if target and target.startswith('session:') else ''
raise ValueError('no image part found%s' % scope)
pid, data = row
d = json.loads(data)
url = d.get('url', '')
if not url.startswith('data:image'):
raise ValueError('part %s has no inline image data' % pid)
return url, {'source': 'opencode.db', 'part': pid, 'filename': d.get('filename')}
def run_ocr(data_url, prompt):
body = json.dumps({
'model': 'Unlimited-OCR',
'temperature': 0,
'repeat_penalty': 1.2,
'max_tokens': 4096,
'messages': [{'role': 'user', 'content': [
{'type': 'image_url', 'image_url': {'url': data_url}},
{'type': 'text', 'text': prompt},
]}],
}).encode()
req = urllib.request.Request(OCR_URL, data=body,
headers={'Content-Type': 'application/json'})
with urllib.request.urlopen(req, timeout=1800) as r:
out = json.load(r)
return out['choices'][0]['message']['content']
def tool_ocr_image(args):
target = args.get('target', 'latest')
prompt = args.get('prompt', 'Free OCR.')
save = args.get('save')
data_url, meta = resolve_data_url(target)
if save:
open(save, 'wb').write(base64.b64decode(data_url.split(',', 1)[1]))
meta['saved_to'] = save
text = run_ocr(data_url, prompt)
header = 'source=%s part=%s file=%s' % (meta.get('source'), meta.get('part', '-'), meta.get('filename', '-'))
if meta.get('saved_to'):
header += ' saved_to=%s' % meta['saved_to']
return header + '\n' + text
def dispatch(msg):
method = msg.get('method')
mid = msg.get('id')
if method is None:
return None
if method == 'initialize':
params = msg.get('params', {})
return {
'protocolVersion': params.get('protocolVersion', '2025-06-18'),
'capabilities': {'tools': {}},
'serverInfo': SERVER_INFO,
}
if method == 'notifications/initialized':
return None
if method == 'ping':
return {}
if method == 'tools/list':
return {'tools': TOOLS}
if method == 'tools/call':
params = msg.get('params', {})
name = params.get('name')
args = params.get('arguments', {}) or {}
if name != 'ocr_image':
return None, {'code': -32602, 'message': 'unknown tool: %s' % name}
try:
text = tool_ocr_image(args)
return {'content': [{'type': 'text', 'text': text}], 'isError': False}, None
except Exception as e:
return {'content': [{'type': 'text', 'text': 'ocr_image failed: %s: %s' % (type(e).__name__, e)}],
'isError': True}, None
if method in ('resources/list', 'prompts/list'):
return {'resources': []} if method == 'resources/list' else {'prompts': []}
if mid is None:
return None
return None, {'code': -32601, 'message': 'method not supported: %s' % method}
def main():
for line in sys.stdin:
line = line.strip()
if not line:
continue
try:
msg = json.loads(line)
except json.JSONDecodeError:
continue
result = dispatch(msg)
if result is None:
continue
if isinstance(result, tuple):
payload, error = result
if error is not None:
out = {'jsonrpc': '2.0', 'id': msg.get('id'), 'error': error}
else:
out = {'jsonrpc': '2.0', 'id': msg.get('id'), 'result': payload}
else:
out = {'jsonrpc': '2.0', 'id': msg.get('id'), 'result': result}
sys.stdout.write(json.dumps(out) + '\n')
sys.stdout.flush()
if __name__ == '__main__':
main()

76
ocr-posted.py Executable file
View File

@ -0,0 +1,76 @@
#!/usr/bin/env python3
"""Pull an image posted into opencode from opencode.db and OCR it via local Unlimited-OCR.
Usage:
ocr-posted.py [latest | <part_id> | session:<session_id>] [prompt] [save_path]
Defaults: latest image part in DB, prompt "Free OCR."
Env: OPENCODE_DB, OCR_URL
"""
import base64, json, os, sqlite3, sys, urllib.request
DB = os.environ.get(
'OPENCODE_DB',
os.path.expanduser('~/.local/share/opencode/opencode.db'))
OCR_URL = os.environ.get('OCR_URL', 'http://127.0.0.1:11436/v1/chat/completions')
def pick(db, arg):
cur = db.cursor()
if arg and arg != 'latest':
if arg.startswith('session:'):
cur.execute(
"SELECT id, data FROM part WHERE session_id=? "
"AND data LIKE '%\"type\":\"file\"%' AND data LIKE '%data:image%' "
"ORDER BY time_created DESC LIMIT 1",
(arg[8:],))
else:
cur.execute("SELECT id, data FROM part WHERE id=?", (arg,))
else:
cur.execute(
"SELECT id, data FROM part WHERE data LIKE '%\"type\":\"file\"%' "
"AND data LIKE '%data:image%' ORDER BY time_created DESC LIMIT 1")
return cur.fetchone()
def ocr(data_url, prompt, max_tokens=4096):
body = json.dumps({
"model": "Unlimited-OCR",
"temperature": 0,
"repeat_penalty": 1.2,
"max_tokens": max_tokens,
"messages": [{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": data_url}},
{"type": "text", "text": prompt},
]}],
}).encode()
req = urllib.request.Request(
OCR_URL, data=body, headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=1800) as r:
return json.load(r)['choices'][0]['message']['content']
def main():
arg = sys.argv[1] if len(sys.argv) > 1 else 'latest'
prompt = sys.argv[2] if len(sys.argv) > 2 else 'Free OCR.'
save = sys.argv[3] if len(sys.argv) > 3 else None
db = sqlite3.connect(DB, timeout=10)
row = pick(db, arg)
if not row:
print('no image part found', file=sys.stderr)
sys.exit(1)
pid, data = row
d = json.loads(data)
url = d.get('url', '')
if not url.startswith('data:image'):
print('part %s has no inline image data' % pid, file=sys.stderr)
sys.exit(1)
if save:
open(save, 'wb').write(base64.b64decode(url.split(',', 1)[1]))
print('saved image to %s' % save, file=sys.stderr)
print('part: %s file: %s' % (pid, d.get('filename')), file=sys.stderr)
print(ocr(url, prompt))
if __name__ == '__main__':
main()

36
run-ocr-server.sh Executable file
View File

@ -0,0 +1,36 @@
#!/usr/bin/env bash
# Launch baidu/Unlimited-OCR via llama-server (OpenAI-compatible, multimodal).
# Foreground — Ctrl+C to stop. All paths overridable via env.
set -euo pipefail
export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-1}"
BIN="${LLAMA_SERVER_BIN:-/usr/local/bin/llama-server}"
MODEL="${OCR_MODEL:-$HOME/models/Unlimited-OCR-Q4_K_M.gguf}"
MMPROJ="${OCR_MMPROJ:-$HOME/models/mmproj-Unlimited-OCR-F16.gguf}"
HOST="${OCR_HOST:-127.0.0.1}"
PORT="${OCR_PORT:-11436}"
CTX="${OCR_CTX:-32768}"
TEMP="${OCR_TEMP:-0}"
REPEAT_PENALTY="${OCR_REPEAT_PENALTY:-1.2}"
[[ -x $BIN ]] || { echo "MISSING (or not executable): $BIN" >&2; exit 1; }
[[ -f $MODEL ]] || { echo "MISSING: $MODEL" >&2; exit 1; }
[[ -f $MMPROJ ]] || { echo "MISSING: $MMPROJ" >&2; exit 1; }
exec "$BIN" \
--model "$MODEL" \
--mmproj "$MMPROJ" \
--alias Unlimited-OCR \
--host "$HOST" \
--port "$PORT" \
--temp "$TEMP" \
--repeat-penalty "$REPEAT_PENALTY" \
--ctx-size "$CTX" \
--n-gpu-layers 99 \
--flash-attn on \
--cont-batching \
--parallel 1 \
--cache-type-k q4_0 \
--cache-type-v q4_0 \
--metrics

68
skill/ocr-image/SKILL.md Normal file
View File

@ -0,0 +1,68 @@
---
name: ocr-image
description: Read images the user posts/attaches in opencode when the current model cannot see them. Use when an image arrives with a placeholder like "Cannot read <file> (this model does not support image input)", or when the user asks to read, OCR, transcribe, or summarize a posted image, screenshot, photo, or document. Calls the local Unlimited-OCR VLM via the ocr_image MCP tool.
license: MIT
---
# OCR Posted Images
A local Unlimited-OCR VLM (llama-server, OpenAI-compatible) reads images the text-only model cannot see.
## When to use
- User posts/attaches an image and you cannot see it — the message shows a placeholder such as
`Cannot read image.png (this model does not support image input)`. The image bytes are safe in
`opencode.db`; only your context lost them.
- User says "read/OCR/transcribe/summarize the image I posted" or similar.
- If you CAN see the image directly (vision model), do not OCR it.
## Tool
Preferred: MCP tool `ocr_image` (server `ocr`):
| arg | meaning |
|---|---|
| `target` | `latest` (default, most recent image part in opencode.db) · part id `prt_...` · `session:<id>` · local file path |
| `prompt` | instruction to the OCR model (see below) |
| `save` | optional path to persist the decoded image |
Fallback (MCP not loaded yet, e.g. before opencode restart): run the CLI from this repo:
```bash
python3 /path/to/unlimited-ocr-mcp/ocr-posted.py [latest|part_id|session:<id>] [prompt] [save_path]
```
## Prompt choice
| task | prompt |
|---|---|
| transcribe document/screenshot | `Free OCR.` (default) |
| structured layout + coords | `<\|grounding\|>Convert the document to markdown.` |
| visual description | `Describe this image in detail.` |
| locate element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` |
Pick the prompt from the user's intent, not from habit.
## Output format
Lines look like `type [x1,y1,x2,y2]content` (types: header, text, table, footer, image,
plus a `source=... part=... file=...` header line). `<img>` marks an image region.
Strip layout prefixes and the source header before presenting — give the user clean text.
## Caveats
- Dense tables: the model can degenerate into repeating the last row (garbled repetition).
If the tail looks looped, re-run the same call — sampling varies per run.
- Animated images (gif): first frame only.
- Sampling is fixed by the server: temp 0, repeat_penalty 1.2. Do not override.
- Backend down? `curl -s http://127.0.0.1:11436/health` → if unreachable, start it with
`run-ocr-server.sh` from this repo (set `CUDA_VISIBLE_DEVICES`, `OCR_MODEL`, `OCR_MMPROJ`).
If the GPU is shared with other models and VRAM is tight, ask the user before stopping anything.
## Example
User: [posts screenshot] "what does this say?"
→ `ocr_image {target: "latest", prompt: "Free OCR."}` → strip layout tokens → answer with the text.
User: "ocr the invoice I sent, I need the total"
→ `ocr_image {target: "latest"}` → extract the total from the transcription.