Initial commit: Unlimited-OCR MCP for opencode
- mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool - ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server) - run-ocr-server.sh: llama-server launcher with tuned sampling flags - skill/ocr-image: opencode skill teaching agents when/how to call it - README with full install instructions
This commit is contained in:
commit
c0fa9148b7
3
.gitignore
vendored
Normal file
3
.gitignore
vendored
Normal file
@ -0,0 +1,3 @@
|
|||||||
|
__pycache__/
|
||||||
|
*.pyc
|
||||||
|
.env
|
||||||
21
LICENSE
Normal file
21
LICENSE
Normal file
@ -0,0 +1,21 @@
|
|||||||
|
MIT License
|
||||||
|
|
||||||
|
Copyright (c) 2026 Jarian Cottingham
|
||||||
|
|
||||||
|
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||||
|
of this software and associated documentation files (the "Software"), to deal
|
||||||
|
in the Software without restriction, including without limitation the rights
|
||||||
|
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||||
|
copies of the Software, and to permit persons to whom the Software is
|
||||||
|
furnished to do so, subject to the following conditions:
|
||||||
|
|
||||||
|
The above copyright notice and this permission notice shall be included in all
|
||||||
|
copies or substantial portions of the Software.
|
||||||
|
|
||||||
|
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||||
|
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||||
|
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||||
|
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||||
|
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||||
|
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||||
|
SOFTWARE.
|
||||||
196
README.md
Normal file
196
README.md
Normal file
@ -0,0 +1,196 @@
|
|||||||
|
# Unlimited-OCR MCP
|
||||||
|
|
||||||
|
Give opencode's **text-only** models the ability to read the images you post.
|
||||||
|
|
||||||
|
`unlimited-ocr-mcp` runs [baidu/Unlimited-OCR](https://huggingface.co/baidu/Unlimited-OCR) — a 3B vision-language model — locally through [llama.cpp](https://github.com/ggml-org/llama.cpp) and exposes it to opencode as an MCP tool. When you paste a screenshot into a chat that runs on a model without vision, opencode strips the image from the model's context but stores it in its local database. The `ocr_image` tool recovers that image and returns clean, transcribed text — so the conversation can continue as if the model had eyes.
|
||||||
|
|
||||||
|
```
|
||||||
|
you paste a screenshot into opencode
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
opencode (text-only model) ── image stripped from context
|
||||||
|
│ │
|
||||||
|
│ ▼
|
||||||
|
│ opencode.db (image stored, base64)
|
||||||
|
│ ▲
|
||||||
|
│ │ reads latest image part
|
||||||
|
▼ │
|
||||||
|
agent sees placeholder ──► ocr_image (MCP tool, stdio)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
llama-server :11436 (Unlimited-OCR, GPU)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
transcribed text back into the chat
|
||||||
|
```
|
||||||
|
|
||||||
|
## Components
|
||||||
|
|
||||||
|
| path | what it is |
|
||||||
|
|---|---|
|
||||||
|
| `mcp/server.py` | MCP stdio server. Zero dependencies (Python stdlib only). Exposes the `ocr_image` tool. |
|
||||||
|
| `ocr-posted.py` | Same pipeline as a plain CLI — useful before the MCP server is registered, or from cron/scripts. |
|
||||||
|
| `run-ocr-server.sh` | Launches the Unlimited-OCR llama-server with the recommended sampling flags. All paths overridable via env. |
|
||||||
|
| `skill/ocr-image/SKILL.md` | An [opencode skill](https://opencode.ai/docs/skills/) that teaches the agent when and how to call the tool. |
|
||||||
|
|
||||||
|
## Requirements
|
||||||
|
|
||||||
|
- **llama.cpp with Unlimited-OCR support** — a build including the DeepSeek-OCR / Unlimited-OCR mtmd support (PRs `ggml-org/llama.cpp#17400`, `#24969`, `#25614`). Builds from ~August 2026 onward include it. The `llama-server` binary must support `--mmproj`.
|
||||||
|
- **A GPU with ~5 GB free VRAM** for the Q4_K_M quant (more for Q6/Q8/BF16). CPU-only works but is slow.
|
||||||
|
- **opencode** with an active session (the tool reads `~/.local/share/opencode/opencode.db`).
|
||||||
|
- Python 3.8+ (stdlib only — no `pip install` needed).
|
||||||
|
|
||||||
|
## Model download
|
||||||
|
|
||||||
|
Pre-quantized GGUFs (imatrix) from [`sahilchachra/Unlimited-OCR-GGUF`](https://huggingface.co/sahilchachra/Unlimited-OCR-GGUF). You need **both** files — the language model and the vision projector:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# New Hugging Face CLI syntax ("hf download", not "huggingface-cli download")
|
||||||
|
hf download sahilchachra/Unlimited-OCR-GGUF \
|
||||||
|
--include "Unlimited-OCR-Q4_K_M.gguf" \
|
||||||
|
--include "mmproj-Unlimited-OCR-F16.gguf" \
|
||||||
|
--local-dir ~/models
|
||||||
|
```
|
||||||
|
|
||||||
|
| quant | size | notes |
|
||||||
|
|---|---|---|
|
||||||
|
| `Unlimited-OCR-Q4_K_M.gguf` | 1.8 GiB | default, good quality/size |
|
||||||
|
| `Unlimited-OCR-Q8_0.gguf` | 2.9 GiB | near-lossless, fewer degenerate repetitions |
|
||||||
|
| `Unlimited-OCR-BF16.gguf` | ~6 GiB | full precision, overkill for most OCR |
|
||||||
|
|
||||||
|
The `mmproj-Unlimited-OCR-F16.gguf` (vision projector, 0.77 GiB) is **required** with any quant.
|
||||||
|
|
||||||
|
## 1. Start the OCR server
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Point the env vars at your llama-server binary and model files
|
||||||
|
export LLAMA_SERVER_BIN=/path/to/llama-server
|
||||||
|
export OCR_MODEL=~/models/Unlimited-OCR-Q4_K_M.gguf
|
||||||
|
export OCR_MMPROJ=~/models/mmproj-Unlimited-OCR-F16.gguf
|
||||||
|
export CUDA_VISIBLE_DEVICES=1 # the GPU with free VRAM
|
||||||
|
|
||||||
|
./run-ocr-server.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
Or launch `llama-server` directly:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
llama-server \
|
||||||
|
--model ~/models/Unlimited-OCR-Q4_K_M.gguf \
|
||||||
|
--mmproj ~/models/mmproj-Unlimited-OCR-F16.gguf \
|
||||||
|
--alias Unlimited-OCR \
|
||||||
|
--host 127.0.0.1 --port 11436 \
|
||||||
|
--temp 0 --repeat-penalty 1.2 \
|
||||||
|
--ctx-size 32768 --n-gpu-layers 99 \
|
||||||
|
--flash-attn on --cont-batching --parallel 1 \
|
||||||
|
--cache-type-k q4_0 --cache-type-v q4_0
|
||||||
|
```
|
||||||
|
|
||||||
|
Verify:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -s http://127.0.0.1:11436/health # → {"status":"ok"}
|
||||||
|
curl -s http://127.0.0.1:11436/v1/models # → capabilities include "multimodal"
|
||||||
|
```
|
||||||
|
|
||||||
|
> **Sampling matters.** `--temp 0 --repeat-penalty 1.2` is the recipe this project is tuned around.
|
||||||
|
> `1.1` tends to loop infinitely; `1.3` starts skipping legitimately repeated rows.
|
||||||
|
|
||||||
|
## 2. Register the MCP server in opencode
|
||||||
|
|
||||||
|
Add to your opencode config (`~/.config/opencode/opencode.json`, global — or a project `opencode.json`):
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"mcp": {
|
||||||
|
"ocr": {
|
||||||
|
"type": "local",
|
||||||
|
"command": ["python3", "/absolute/path/to/unlimited-ocr-mcp/mcp/server.py"],
|
||||||
|
"enabled": true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Then **restart opencode** — MCP servers are loaded at startup. The tool appears as `ocr_image`.
|
||||||
|
|
||||||
|
Environment overrides (optional): `OPENCODE_DB` (path to `opencode.db`, default `~/.local/share/opencode/opencode.db`) and `OCR_URL` (chat completions endpoint, default `http://127.0.0.1:11436/v1/chat/completions`).
|
||||||
|
|
||||||
|
## 3. Install the skill
|
||||||
|
|
||||||
|
Copy the skill into opencode's global skills directory so every project knows to use the tool:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
mkdir -p ~/.config/opencode/skills
|
||||||
|
cp -r skill/ocr-image ~/.config/opencode/skills/
|
||||||
|
```
|
||||||
|
|
||||||
|
The skill triggers when an image is posted to a text-only model (the
|
||||||
|
`Cannot read <file> (this model does not support image input)` placeholder) or when you ask
|
||||||
|
"read the image I posted". It also documents the CLI fallback (`ocr-posted.py`) for sessions
|
||||||
|
where the MCP server is not loaded yet.
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
Post any image into opencode and ask about it:
|
||||||
|
|
||||||
|
> *what does this screenshot say?*
|
||||||
|
> *OCR the invoice I sent, I need the total*
|
||||||
|
> *describe the chart in this image*
|
||||||
|
|
||||||
|
The agent calls `ocr_image` with a prompt matched to your intent:
|
||||||
|
|
||||||
|
| task | prompt sent to the model |
|
||||||
|
|---|---|
|
||||||
|
| transcribe document/screenshot (default) | `Free OCR.` |
|
||||||
|
| structured layout with coordinates | `<\|grounding\|>Convert the document to markdown.` |
|
||||||
|
| visual description | `Describe this image in detail.` |
|
||||||
|
| locate an element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` |
|
||||||
|
|
||||||
|
`ocr_image` arguments:
|
||||||
|
|
||||||
|
| arg | default | meaning |
|
||||||
|
|---|---|---|
|
||||||
|
| `target` | `latest` | `latest` image part in `opencode.db`, a part id (`prt_...`), `session:<id>`, or a local image file path |
|
||||||
|
| `prompt` | `Free OCR.` | instruction sent to the OCR model |
|
||||||
|
| `save` | — | optional path to persist the decoded image |
|
||||||
|
|
||||||
|
Raw output is layout-prefixed (`header [x1,y1,x2,y2]text`, `table [...]`, `<img>` for image
|
||||||
|
regions); the skill instructs the agent to strip the layout tokens and present clean text.
|
||||||
|
|
||||||
|
CLI fallback (identical pipeline, no MCP needed):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 ocr-posted.py latest "Free OCR."
|
||||||
|
python3 ocr-posted.py prt_0123abcd... "Describe this image in detail." /tmp/shot.png
|
||||||
|
```
|
||||||
|
|
||||||
|
## How it works
|
||||||
|
|
||||||
|
1. opencode stores every posted file part in its SQLite database (`~/.local/share/opencode/opencode.db`, table `part`). Image parts keep the full payload as a `data:` URL — even when the active model cannot see them.
|
||||||
|
2. `ocr_image` queries that table for the newest image part (or a specific part/session), decodes it, and POSTs it to the llama-server `/v1/chat/completions` endpoint with `temperature 0` and `repeat_penalty 1.2`.
|
||||||
|
3. The transcription returns through MCP into the agent's context.
|
||||||
|
|
||||||
|
## Troubleshooting
|
||||||
|
|
||||||
|
- **`no image part found`** — no image was posted yet, or you're pointing at a different opencode data dir. Set `OPENCODE_DB` accordingly.
|
||||||
|
- **Connection refused on :11436** — the OCR server isn't running. Start it (section 1). Check GPU VRAM with `nvidia-smi`; a model sharing the card may need to be stopped first.
|
||||||
|
- **Garbled tail / repeated last row** — known degenerate behavior of the 3B quant on dense tables. Re-run the same call; sampling varies per run. Q8_0 reduces it.
|
||||||
|
- **`multimodal` missing from `/v1/models` capabilities** — your llama.cpp build predates Unlimited-OCR support, or `--mmproj` wasn't passed.
|
||||||
|
- **Tool not visible after config change** — opencode loads MCP servers at startup; restart it.
|
||||||
|
|
||||||
|
## Repository layout
|
||||||
|
|
||||||
|
```
|
||||||
|
unlimited-ocr-mcp/
|
||||||
|
├── README.md
|
||||||
|
├── LICENSE
|
||||||
|
├── mcp/server.py # MCP stdio server (ocr_image tool)
|
||||||
|
├── ocr-posted.py # CLI equivalent
|
||||||
|
├── run-ocr-server.sh # llama-server launcher
|
||||||
|
└── skill/ocr-image/SKILL.md # opencode skill
|
||||||
|
```
|
||||||
|
|
||||||
|
## License
|
||||||
|
|
||||||
|
[MIT](LICENSE) — see also the model card for `baidu/Unlimited-OCR` for the upstream model license.
|
||||||
190
mcp/server.py
Executable file
190
mcp/server.py
Executable file
@ -0,0 +1,190 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""MCP stdio server: OCR images posted into opencode via a local Unlimited-OCR backend.
|
||||||
|
|
||||||
|
Transport: newline-delimited JSON-RPC over stdin/stdout (MCP stdio).
|
||||||
|
Stdlib only. Env overrides: OPENCODE_DB, OCR_URL.
|
||||||
|
"""
|
||||||
|
import base64
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import sqlite3
|
||||||
|
import sys
|
||||||
|
import urllib.request
|
||||||
|
|
||||||
|
DB = os.environ.get(
|
||||||
|
'OPENCODE_DB',
|
||||||
|
os.path.expanduser('~/.local/share/opencode/opencode.db'))
|
||||||
|
OCR_URL = os.environ.get('OCR_URL', 'http://127.0.0.1:11436/v1/chat/completions')
|
||||||
|
SERVER_INFO = {'name': 'ocr-posted', 'version': '1.0.0'}
|
||||||
|
|
||||||
|
TOOLS = [
|
||||||
|
{
|
||||||
|
'name': 'ocr_image',
|
||||||
|
'description': (
|
||||||
|
'OCR an image the user posted into opencode (or any local image file) using the '
|
||||||
|
'local Unlimited-OCR VLM server. Use this when the user attaches or pastes an image '
|
||||||
|
'that you cannot see directly, e.g. when the message shows '
|
||||||
|
"'Cannot read <file> (this model does not support image input)'. "
|
||||||
|
'Returns transcribed text (layout-prefixed lines with prompts that ask for grounding).'
|
||||||
|
),
|
||||||
|
'inputSchema': {
|
||||||
|
'type': 'object',
|
||||||
|
'properties': {
|
||||||
|
'target': {
|
||||||
|
'type': 'string',
|
||||||
|
'description': (
|
||||||
|
"'latest' (default): most recent image part in opencode.db. "
|
||||||
|
"Or a part id (prt_...), 'session:<session_id>', or a local image file path."
|
||||||
|
),
|
||||||
|
},
|
||||||
|
'prompt': {
|
||||||
|
'type': 'string',
|
||||||
|
'description': (
|
||||||
|
"Instruction sent to the OCR model. Default 'Free OCR.' (plain transcription). "
|
||||||
|
"Use '<|grounding|>Convert the document to markdown.' for structured layout, "
|
||||||
|
"'Describe this image in detail.' for visual QA."
|
||||||
|
),
|
||||||
|
},
|
||||||
|
'save': {
|
||||||
|
'type': 'string',
|
||||||
|
'description': 'Optional local path to save the decoded image to.',
|
||||||
|
},
|
||||||
|
},
|
||||||
|
'required': [],
|
||||||
|
},
|
||||||
|
}
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_data_url(target):
|
||||||
|
"""Return (data_url, meta) for the requested target."""
|
||||||
|
if target and target != 'latest' and not target.startswith('session:'):
|
||||||
|
if os.path.isfile(target):
|
||||||
|
ext = os.path.splitext(target)[1].lower()
|
||||||
|
mime = {'.png': 'image/png', '.jpg': 'image/jpeg', '.jpeg': 'image/jpeg',
|
||||||
|
'.gif': 'image/gif', '.webp': 'image/webp'}.get(ext, 'image/png')
|
||||||
|
b64 = base64.b64encode(open(target, 'rb').read()).decode()
|
||||||
|
return 'data:%s;base64,%s' % (mime, b64), {'source': 'file', 'filename': target}
|
||||||
|
db = sqlite3.connect(DB, timeout=10)
|
||||||
|
row = db.execute('SELECT id, data FROM part WHERE id=?', (target,)).fetchone()
|
||||||
|
if not row:
|
||||||
|
raise ValueError('part id not found: %s' % target)
|
||||||
|
pid, data = row
|
||||||
|
else:
|
||||||
|
db = sqlite3.connect(DB, timeout=10)
|
||||||
|
if target and target.startswith('session:'):
|
||||||
|
row = db.execute(
|
||||||
|
"SELECT id, data FROM part WHERE session_id=? "
|
||||||
|
"AND data LIKE '%\"type\":\"file\"%' AND data LIKE '%data:image%' "
|
||||||
|
"ORDER BY time_created DESC LIMIT 1", (target[8:],)).fetchone()
|
||||||
|
else:
|
||||||
|
row = db.execute(
|
||||||
|
"SELECT id, data FROM part WHERE data LIKE '%\"type\":\"file\"%' "
|
||||||
|
"AND data LIKE '%data:image%' ORDER BY time_created DESC LIMIT 1").fetchone()
|
||||||
|
if not row:
|
||||||
|
scope = ' in session %s' % target[8:] if target and target.startswith('session:') else ''
|
||||||
|
raise ValueError('no image part found%s' % scope)
|
||||||
|
pid, data = row
|
||||||
|
d = json.loads(data)
|
||||||
|
url = d.get('url', '')
|
||||||
|
if not url.startswith('data:image'):
|
||||||
|
raise ValueError('part %s has no inline image data' % pid)
|
||||||
|
return url, {'source': 'opencode.db', 'part': pid, 'filename': d.get('filename')}
|
||||||
|
|
||||||
|
|
||||||
|
def run_ocr(data_url, prompt):
|
||||||
|
body = json.dumps({
|
||||||
|
'model': 'Unlimited-OCR',
|
||||||
|
'temperature': 0,
|
||||||
|
'repeat_penalty': 1.2,
|
||||||
|
'max_tokens': 4096,
|
||||||
|
'messages': [{'role': 'user', 'content': [
|
||||||
|
{'type': 'image_url', 'image_url': {'url': data_url}},
|
||||||
|
{'type': 'text', 'text': prompt},
|
||||||
|
]}],
|
||||||
|
}).encode()
|
||||||
|
req = urllib.request.Request(OCR_URL, data=body,
|
||||||
|
headers={'Content-Type': 'application/json'})
|
||||||
|
with urllib.request.urlopen(req, timeout=1800) as r:
|
||||||
|
out = json.load(r)
|
||||||
|
return out['choices'][0]['message']['content']
|
||||||
|
|
||||||
|
|
||||||
|
def tool_ocr_image(args):
|
||||||
|
target = args.get('target', 'latest')
|
||||||
|
prompt = args.get('prompt', 'Free OCR.')
|
||||||
|
save = args.get('save')
|
||||||
|
data_url, meta = resolve_data_url(target)
|
||||||
|
if save:
|
||||||
|
open(save, 'wb').write(base64.b64decode(data_url.split(',', 1)[1]))
|
||||||
|
meta['saved_to'] = save
|
||||||
|
text = run_ocr(data_url, prompt)
|
||||||
|
header = 'source=%s part=%s file=%s' % (meta.get('source'), meta.get('part', '-'), meta.get('filename', '-'))
|
||||||
|
if meta.get('saved_to'):
|
||||||
|
header += ' saved_to=%s' % meta['saved_to']
|
||||||
|
return header + '\n' + text
|
||||||
|
|
||||||
|
|
||||||
|
def dispatch(msg):
|
||||||
|
method = msg.get('method')
|
||||||
|
mid = msg.get('id')
|
||||||
|
if method is None:
|
||||||
|
return None
|
||||||
|
if method == 'initialize':
|
||||||
|
params = msg.get('params', {})
|
||||||
|
return {
|
||||||
|
'protocolVersion': params.get('protocolVersion', '2025-06-18'),
|
||||||
|
'capabilities': {'tools': {}},
|
||||||
|
'serverInfo': SERVER_INFO,
|
||||||
|
}
|
||||||
|
if method == 'notifications/initialized':
|
||||||
|
return None
|
||||||
|
if method == 'ping':
|
||||||
|
return {}
|
||||||
|
if method == 'tools/list':
|
||||||
|
return {'tools': TOOLS}
|
||||||
|
if method == 'tools/call':
|
||||||
|
params = msg.get('params', {})
|
||||||
|
name = params.get('name')
|
||||||
|
args = params.get('arguments', {}) or {}
|
||||||
|
if name != 'ocr_image':
|
||||||
|
return None, {'code': -32602, 'message': 'unknown tool: %s' % name}
|
||||||
|
try:
|
||||||
|
text = tool_ocr_image(args)
|
||||||
|
return {'content': [{'type': 'text', 'text': text}], 'isError': False}, None
|
||||||
|
except Exception as e:
|
||||||
|
return {'content': [{'type': 'text', 'text': 'ocr_image failed: %s: %s' % (type(e).__name__, e)}],
|
||||||
|
'isError': True}, None
|
||||||
|
if method in ('resources/list', 'prompts/list'):
|
||||||
|
return {'resources': []} if method == 'resources/list' else {'prompts': []}
|
||||||
|
if mid is None:
|
||||||
|
return None
|
||||||
|
return None, {'code': -32601, 'message': 'method not supported: %s' % method}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
for line in sys.stdin:
|
||||||
|
line = line.strip()
|
||||||
|
if not line:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
msg = json.loads(line)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
continue
|
||||||
|
result = dispatch(msg)
|
||||||
|
if result is None:
|
||||||
|
continue
|
||||||
|
if isinstance(result, tuple):
|
||||||
|
payload, error = result
|
||||||
|
if error is not None:
|
||||||
|
out = {'jsonrpc': '2.0', 'id': msg.get('id'), 'error': error}
|
||||||
|
else:
|
||||||
|
out = {'jsonrpc': '2.0', 'id': msg.get('id'), 'result': payload}
|
||||||
|
else:
|
||||||
|
out = {'jsonrpc': '2.0', 'id': msg.get('id'), 'result': result}
|
||||||
|
sys.stdout.write(json.dumps(out) + '\n')
|
||||||
|
sys.stdout.flush()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
76
ocr-posted.py
Executable file
76
ocr-posted.py
Executable file
@ -0,0 +1,76 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Pull an image posted into opencode from opencode.db and OCR it via local Unlimited-OCR.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
ocr-posted.py [latest | <part_id> | session:<session_id>] [prompt] [save_path]
|
||||||
|
|
||||||
|
Defaults: latest image part in DB, prompt "Free OCR."
|
||||||
|
Env: OPENCODE_DB, OCR_URL
|
||||||
|
"""
|
||||||
|
import base64, json, os, sqlite3, sys, urllib.request
|
||||||
|
|
||||||
|
DB = os.environ.get(
|
||||||
|
'OPENCODE_DB',
|
||||||
|
os.path.expanduser('~/.local/share/opencode/opencode.db'))
|
||||||
|
OCR_URL = os.environ.get('OCR_URL', 'http://127.0.0.1:11436/v1/chat/completions')
|
||||||
|
|
||||||
|
|
||||||
|
def pick(db, arg):
|
||||||
|
cur = db.cursor()
|
||||||
|
if arg and arg != 'latest':
|
||||||
|
if arg.startswith('session:'):
|
||||||
|
cur.execute(
|
||||||
|
"SELECT id, data FROM part WHERE session_id=? "
|
||||||
|
"AND data LIKE '%\"type\":\"file\"%' AND data LIKE '%data:image%' "
|
||||||
|
"ORDER BY time_created DESC LIMIT 1",
|
||||||
|
(arg[8:],))
|
||||||
|
else:
|
||||||
|
cur.execute("SELECT id, data FROM part WHERE id=?", (arg,))
|
||||||
|
else:
|
||||||
|
cur.execute(
|
||||||
|
"SELECT id, data FROM part WHERE data LIKE '%\"type\":\"file\"%' "
|
||||||
|
"AND data LIKE '%data:image%' ORDER BY time_created DESC LIMIT 1")
|
||||||
|
return cur.fetchone()
|
||||||
|
|
||||||
|
|
||||||
|
def ocr(data_url, prompt, max_tokens=4096):
|
||||||
|
body = json.dumps({
|
||||||
|
"model": "Unlimited-OCR",
|
||||||
|
"temperature": 0,
|
||||||
|
"repeat_penalty": 1.2,
|
||||||
|
"max_tokens": max_tokens,
|
||||||
|
"messages": [{"role": "user", "content": [
|
||||||
|
{"type": "image_url", "image_url": {"url": data_url}},
|
||||||
|
{"type": "text", "text": prompt},
|
||||||
|
]}],
|
||||||
|
}).encode()
|
||||||
|
req = urllib.request.Request(
|
||||||
|
OCR_URL, data=body, headers={"Content-Type": "application/json"})
|
||||||
|
with urllib.request.urlopen(req, timeout=1800) as r:
|
||||||
|
return json.load(r)['choices'][0]['message']['content']
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
arg = sys.argv[1] if len(sys.argv) > 1 else 'latest'
|
||||||
|
prompt = sys.argv[2] if len(sys.argv) > 2 else 'Free OCR.'
|
||||||
|
save = sys.argv[3] if len(sys.argv) > 3 else None
|
||||||
|
db = sqlite3.connect(DB, timeout=10)
|
||||||
|
row = pick(db, arg)
|
||||||
|
if not row:
|
||||||
|
print('no image part found', file=sys.stderr)
|
||||||
|
sys.exit(1)
|
||||||
|
pid, data = row
|
||||||
|
d = json.loads(data)
|
||||||
|
url = d.get('url', '')
|
||||||
|
if not url.startswith('data:image'):
|
||||||
|
print('part %s has no inline image data' % pid, file=sys.stderr)
|
||||||
|
sys.exit(1)
|
||||||
|
if save:
|
||||||
|
open(save, 'wb').write(base64.b64decode(url.split(',', 1)[1]))
|
||||||
|
print('saved image to %s' % save, file=sys.stderr)
|
||||||
|
print('part: %s file: %s' % (pid, d.get('filename')), file=sys.stderr)
|
||||||
|
print(ocr(url, prompt))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
36
run-ocr-server.sh
Executable file
36
run-ocr-server.sh
Executable file
@ -0,0 +1,36 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Launch baidu/Unlimited-OCR via llama-server (OpenAI-compatible, multimodal).
|
||||||
|
# Foreground — Ctrl+C to stop. All paths overridable via env.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-1}"
|
||||||
|
|
||||||
|
BIN="${LLAMA_SERVER_BIN:-/usr/local/bin/llama-server}"
|
||||||
|
MODEL="${OCR_MODEL:-$HOME/models/Unlimited-OCR-Q4_K_M.gguf}"
|
||||||
|
MMPROJ="${OCR_MMPROJ:-$HOME/models/mmproj-Unlimited-OCR-F16.gguf}"
|
||||||
|
HOST="${OCR_HOST:-127.0.0.1}"
|
||||||
|
PORT="${OCR_PORT:-11436}"
|
||||||
|
CTX="${OCR_CTX:-32768}"
|
||||||
|
TEMP="${OCR_TEMP:-0}"
|
||||||
|
REPEAT_PENALTY="${OCR_REPEAT_PENALTY:-1.2}"
|
||||||
|
|
||||||
|
[[ -x $BIN ]] || { echo "MISSING (or not executable): $BIN" >&2; exit 1; }
|
||||||
|
[[ -f $MODEL ]] || { echo "MISSING: $MODEL" >&2; exit 1; }
|
||||||
|
[[ -f $MMPROJ ]] || { echo "MISSING: $MMPROJ" >&2; exit 1; }
|
||||||
|
|
||||||
|
exec "$BIN" \
|
||||||
|
--model "$MODEL" \
|
||||||
|
--mmproj "$MMPROJ" \
|
||||||
|
--alias Unlimited-OCR \
|
||||||
|
--host "$HOST" \
|
||||||
|
--port "$PORT" \
|
||||||
|
--temp "$TEMP" \
|
||||||
|
--repeat-penalty "$REPEAT_PENALTY" \
|
||||||
|
--ctx-size "$CTX" \
|
||||||
|
--n-gpu-layers 99 \
|
||||||
|
--flash-attn on \
|
||||||
|
--cont-batching \
|
||||||
|
--parallel 1 \
|
||||||
|
--cache-type-k q4_0 \
|
||||||
|
--cache-type-v q4_0 \
|
||||||
|
--metrics
|
||||||
68
skill/ocr-image/SKILL.md
Normal file
68
skill/ocr-image/SKILL.md
Normal file
@ -0,0 +1,68 @@
|
|||||||
|
---
|
||||||
|
name: ocr-image
|
||||||
|
description: Read images the user posts/attaches in opencode when the current model cannot see them. Use when an image arrives with a placeholder like "Cannot read <file> (this model does not support image input)", or when the user asks to read, OCR, transcribe, or summarize a posted image, screenshot, photo, or document. Calls the local Unlimited-OCR VLM via the ocr_image MCP tool.
|
||||||
|
license: MIT
|
||||||
|
---
|
||||||
|
|
||||||
|
# OCR Posted Images
|
||||||
|
|
||||||
|
A local Unlimited-OCR VLM (llama-server, OpenAI-compatible) reads images the text-only model cannot see.
|
||||||
|
|
||||||
|
## When to use
|
||||||
|
|
||||||
|
- User posts/attaches an image and you cannot see it — the message shows a placeholder such as
|
||||||
|
`Cannot read image.png (this model does not support image input)`. The image bytes are safe in
|
||||||
|
`opencode.db`; only your context lost them.
|
||||||
|
- User says "read/OCR/transcribe/summarize the image I posted" or similar.
|
||||||
|
- If you CAN see the image directly (vision model), do not OCR it.
|
||||||
|
|
||||||
|
## Tool
|
||||||
|
|
||||||
|
Preferred: MCP tool `ocr_image` (server `ocr`):
|
||||||
|
|
||||||
|
| arg | meaning |
|
||||||
|
|---|---|
|
||||||
|
| `target` | `latest` (default, most recent image part in opencode.db) · part id `prt_...` · `session:<id>` · local file path |
|
||||||
|
| `prompt` | instruction to the OCR model (see below) |
|
||||||
|
| `save` | optional path to persist the decoded image |
|
||||||
|
|
||||||
|
Fallback (MCP not loaded yet, e.g. before opencode restart): run the CLI from this repo:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 /path/to/unlimited-ocr-mcp/ocr-posted.py [latest|part_id|session:<id>] [prompt] [save_path]
|
||||||
|
```
|
||||||
|
|
||||||
|
## Prompt choice
|
||||||
|
|
||||||
|
| task | prompt |
|
||||||
|
|---|---|
|
||||||
|
| transcribe document/screenshot | `Free OCR.` (default) |
|
||||||
|
| structured layout + coords | `<\|grounding\|>Convert the document to markdown.` |
|
||||||
|
| visual description | `Describe this image in detail.` |
|
||||||
|
| locate element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` |
|
||||||
|
|
||||||
|
Pick the prompt from the user's intent, not from habit.
|
||||||
|
|
||||||
|
## Output format
|
||||||
|
|
||||||
|
Lines look like `type [x1,y1,x2,y2]content` (types: header, text, table, footer, image,
|
||||||
|
plus a `source=... part=... file=...` header line). `<img>` marks an image region.
|
||||||
|
Strip layout prefixes and the source header before presenting — give the user clean text.
|
||||||
|
|
||||||
|
## Caveats
|
||||||
|
|
||||||
|
- Dense tables: the model can degenerate into repeating the last row (garbled repetition).
|
||||||
|
If the tail looks looped, re-run the same call — sampling varies per run.
|
||||||
|
- Animated images (gif): first frame only.
|
||||||
|
- Sampling is fixed by the server: temp 0, repeat_penalty 1.2. Do not override.
|
||||||
|
- Backend down? `curl -s http://127.0.0.1:11436/health` → if unreachable, start it with
|
||||||
|
`run-ocr-server.sh` from this repo (set `CUDA_VISIBLE_DEVICES`, `OCR_MODEL`, `OCR_MMPROJ`).
|
||||||
|
If the GPU is shared with other models and VRAM is tight, ask the user before stopping anything.
|
||||||
|
|
||||||
|
## Example
|
||||||
|
|
||||||
|
User: [posts screenshot] "what does this say?"
|
||||||
|
→ `ocr_image {target: "latest", prompt: "Free OCR."}` → strip layout tokens → answer with the text.
|
||||||
|
|
||||||
|
User: "ocr the invoice I sent, I need the total"
|
||||||
|
→ `ocr_image {target: "latest"}` → extract the total from the transcription.
|
||||||
Loading…
x
Reference in New Issue
Block a user