commit c0fa9148b79dfcde9d9e974ee38ceb4e26f436d4 Author: Jarian Cottingham Date: Sat Aug 22 06:24:57 2026 +0000 Initial commit: Unlimited-OCR MCP for opencode - mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool - ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server) - run-ocr-server.sh: llama-server launcher with tuned sampling flags - skill/ocr-image: opencode skill teaching agents when/how to call it - README with full install instructions diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..b6cf5f0 --- /dev/null +++ b/.gitignore @@ -0,0 +1,3 @@ +__pycache__/ +*.pyc +.env diff --git a/LICENSE b/LICENSE new file mode 100644 index 0000000..eba8c1d --- /dev/null +++ b/LICENSE @@ -0,0 +1,21 @@ +MIT License + +Copyright (c) 2026 Jarian Cottingham + +Permission is hereby granted, free of charge, to any person obtaining a copy +of this software and associated documentation files (the "Software"), to deal +in the Software without restriction, including without limitation the rights +to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +copies of the Software, and to permit persons to whom the Software is +furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all +copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +SOFTWARE. diff --git a/README.md b/README.md new file mode 100644 index 0000000..2187349 --- /dev/null +++ b/README.md @@ -0,0 +1,196 @@ +# Unlimited-OCR MCP + +Give opencode's **text-only** models the ability to read the images you post. + +`unlimited-ocr-mcp` runs [baidu/Unlimited-OCR](https://huggingface.co/baidu/Unlimited-OCR) — a 3B vision-language model — locally through [llama.cpp](https://github.com/ggml-org/llama.cpp) and exposes it to opencode as an MCP tool. When you paste a screenshot into a chat that runs on a model without vision, opencode strips the image from the model's context but stores it in its local database. The `ocr_image` tool recovers that image and returns clean, transcribed text — so the conversation can continue as if the model had eyes. + +``` +you paste a screenshot into opencode + │ + ▼ + opencode (text-only model) ── image stripped from context + │ │ + │ ▼ + │ opencode.db (image stored, base64) + │ ▲ + │ │ reads latest image part + ▼ │ + agent sees placeholder ──► ocr_image (MCP tool, stdio) + │ + ▼ + llama-server :11436 (Unlimited-OCR, GPU) + │ + ▼ + transcribed text back into the chat +``` + +## Components + +| path | what it is | +|---|---| +| `mcp/server.py` | MCP stdio server. Zero dependencies (Python stdlib only). Exposes the `ocr_image` tool. | +| `ocr-posted.py` | Same pipeline as a plain CLI — useful before the MCP server is registered, or from cron/scripts. | +| `run-ocr-server.sh` | Launches the Unlimited-OCR llama-server with the recommended sampling flags. All paths overridable via env. | +| `skill/ocr-image/SKILL.md` | An [opencode skill](https://opencode.ai/docs/skills/) that teaches the agent when and how to call the tool. | + +## Requirements + +- **llama.cpp with Unlimited-OCR support** — a build including the DeepSeek-OCR / Unlimited-OCR mtmd support (PRs `ggml-org/llama.cpp#17400`, `#24969`, `#25614`). Builds from ~August 2026 onward include it. The `llama-server` binary must support `--mmproj`. +- **A GPU with ~5 GB free VRAM** for the Q4_K_M quant (more for Q6/Q8/BF16). CPU-only works but is slow. +- **opencode** with an active session (the tool reads `~/.local/share/opencode/opencode.db`). +- Python 3.8+ (stdlib only — no `pip install` needed). + +## Model download + +Pre-quantized GGUFs (imatrix) from [`sahilchachra/Unlimited-OCR-GGUF`](https://huggingface.co/sahilchachra/Unlimited-OCR-GGUF). You need **both** files — the language model and the vision projector: + +```bash +# New Hugging Face CLI syntax ("hf download", not "huggingface-cli download") +hf download sahilchachra/Unlimited-OCR-GGUF \ + --include "Unlimited-OCR-Q4_K_M.gguf" \ + --include "mmproj-Unlimited-OCR-F16.gguf" \ + --local-dir ~/models +``` + +| quant | size | notes | +|---|---|---| +| `Unlimited-OCR-Q4_K_M.gguf` | 1.8 GiB | default, good quality/size | +| `Unlimited-OCR-Q8_0.gguf` | 2.9 GiB | near-lossless, fewer degenerate repetitions | +| `Unlimited-OCR-BF16.gguf` | ~6 GiB | full precision, overkill for most OCR | + +The `mmproj-Unlimited-OCR-F16.gguf` (vision projector, 0.77 GiB) is **required** with any quant. + +## 1. Start the OCR server + +```bash +# Point the env vars at your llama-server binary and model files +export LLAMA_SERVER_BIN=/path/to/llama-server +export OCR_MODEL=~/models/Unlimited-OCR-Q4_K_M.gguf +export OCR_MMPROJ=~/models/mmproj-Unlimited-OCR-F16.gguf +export CUDA_VISIBLE_DEVICES=1 # the GPU with free VRAM + +./run-ocr-server.sh +``` + +Or launch `llama-server` directly: + +```bash +llama-server \ + --model ~/models/Unlimited-OCR-Q4_K_M.gguf \ + --mmproj ~/models/mmproj-Unlimited-OCR-F16.gguf \ + --alias Unlimited-OCR \ + --host 127.0.0.1 --port 11436 \ + --temp 0 --repeat-penalty 1.2 \ + --ctx-size 32768 --n-gpu-layers 99 \ + --flash-attn on --cont-batching --parallel 1 \ + --cache-type-k q4_0 --cache-type-v q4_0 +``` + +Verify: + +```bash +curl -s http://127.0.0.1:11436/health # → {"status":"ok"} +curl -s http://127.0.0.1:11436/v1/models # → capabilities include "multimodal" +``` + +> **Sampling matters.** `--temp 0 --repeat-penalty 1.2` is the recipe this project is tuned around. +> `1.1` tends to loop infinitely; `1.3` starts skipping legitimately repeated rows. + +## 2. Register the MCP server in opencode + +Add to your opencode config (`~/.config/opencode/opencode.json`, global — or a project `opencode.json`): + +```json +{ + "mcp": { + "ocr": { + "type": "local", + "command": ["python3", "/absolute/path/to/unlimited-ocr-mcp/mcp/server.py"], + "enabled": true + } + } +} +``` + +Then **restart opencode** — MCP servers are loaded at startup. The tool appears as `ocr_image`. + +Environment overrides (optional): `OPENCODE_DB` (path to `opencode.db`, default `~/.local/share/opencode/opencode.db`) and `OCR_URL` (chat completions endpoint, default `http://127.0.0.1:11436/v1/chat/completions`). + +## 3. Install the skill + +Copy the skill into opencode's global skills directory so every project knows to use the tool: + +```bash +mkdir -p ~/.config/opencode/skills +cp -r skill/ocr-image ~/.config/opencode/skills/ +``` + +The skill triggers when an image is posted to a text-only model (the +`Cannot read (this model does not support image input)` placeholder) or when you ask +"read the image I posted". It also documents the CLI fallback (`ocr-posted.py`) for sessions +where the MCP server is not loaded yet. + +## Usage + +Post any image into opencode and ask about it: + +> *what does this screenshot say?* +> *OCR the invoice I sent, I need the total* +> *describe the chart in this image* + +The agent calls `ocr_image` with a prompt matched to your intent: + +| task | prompt sent to the model | +|---|---| +| transcribe document/screenshot (default) | `Free OCR.` | +| structured layout with coordinates | `<\|grounding\|>Convert the document to markdown.` | +| visual description | `Describe this image in detail.` | +| locate an element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` | + +`ocr_image` arguments: + +| arg | default | meaning | +|---|---|---| +| `target` | `latest` | `latest` image part in `opencode.db`, a part id (`prt_...`), `session:`, or a local image file path | +| `prompt` | `Free OCR.` | instruction sent to the OCR model | +| `save` | — | optional path to persist the decoded image | + +Raw output is layout-prefixed (`header [x1,y1,x2,y2]text`, `table [...]`, `` for image +regions); the skill instructs the agent to strip the layout tokens and present clean text. + +CLI fallback (identical pipeline, no MCP needed): + +```bash +python3 ocr-posted.py latest "Free OCR." +python3 ocr-posted.py prt_0123abcd... "Describe this image in detail." /tmp/shot.png +``` + +## How it works + +1. opencode stores every posted file part in its SQLite database (`~/.local/share/opencode/opencode.db`, table `part`). Image parts keep the full payload as a `data:` URL — even when the active model cannot see them. +2. `ocr_image` queries that table for the newest image part (or a specific part/session), decodes it, and POSTs it to the llama-server `/v1/chat/completions` endpoint with `temperature 0` and `repeat_penalty 1.2`. +3. The transcription returns through MCP into the agent's context. + +## Troubleshooting + +- **`no image part found`** — no image was posted yet, or you're pointing at a different opencode data dir. Set `OPENCODE_DB` accordingly. +- **Connection refused on :11436** — the OCR server isn't running. Start it (section 1). Check GPU VRAM with `nvidia-smi`; a model sharing the card may need to be stopped first. +- **Garbled tail / repeated last row** — known degenerate behavior of the 3B quant on dense tables. Re-run the same call; sampling varies per run. Q8_0 reduces it. +- **`multimodal` missing from `/v1/models` capabilities** — your llama.cpp build predates Unlimited-OCR support, or `--mmproj` wasn't passed. +- **Tool not visible after config change** — opencode loads MCP servers at startup; restart it. + +## Repository layout + +``` +unlimited-ocr-mcp/ +├── README.md +├── LICENSE +├── mcp/server.py # MCP stdio server (ocr_image tool) +├── ocr-posted.py # CLI equivalent +├── run-ocr-server.sh # llama-server launcher +└── skill/ocr-image/SKILL.md # opencode skill +``` + +## License + +[MIT](LICENSE) — see also the model card for `baidu/Unlimited-OCR` for the upstream model license. diff --git a/mcp/server.py b/mcp/server.py new file mode 100755 index 0000000..7296997 --- /dev/null +++ b/mcp/server.py @@ -0,0 +1,190 @@ +#!/usr/bin/env python3 +"""MCP stdio server: OCR images posted into opencode via a local Unlimited-OCR backend. + +Transport: newline-delimited JSON-RPC over stdin/stdout (MCP stdio). +Stdlib only. Env overrides: OPENCODE_DB, OCR_URL. +""" +import base64 +import json +import os +import sqlite3 +import sys +import urllib.request + +DB = os.environ.get( + 'OPENCODE_DB', + os.path.expanduser('~/.local/share/opencode/opencode.db')) +OCR_URL = os.environ.get('OCR_URL', 'http://127.0.0.1:11436/v1/chat/completions') +SERVER_INFO = {'name': 'ocr-posted', 'version': '1.0.0'} + +TOOLS = [ + { + 'name': 'ocr_image', + 'description': ( + 'OCR an image the user posted into opencode (or any local image file) using the ' + 'local Unlimited-OCR VLM server. Use this when the user attaches or pastes an image ' + 'that you cannot see directly, e.g. when the message shows ' + "'Cannot read (this model does not support image input)'. " + 'Returns transcribed text (layout-prefixed lines with prompts that ask for grounding).' + ), + 'inputSchema': { + 'type': 'object', + 'properties': { + 'target': { + 'type': 'string', + 'description': ( + "'latest' (default): most recent image part in opencode.db. " + "Or a part id (prt_...), 'session:', or a local image file path." + ), + }, + 'prompt': { + 'type': 'string', + 'description': ( + "Instruction sent to the OCR model. Default 'Free OCR.' (plain transcription). " + "Use '<|grounding|>Convert the document to markdown.' for structured layout, " + "'Describe this image in detail.' for visual QA." + ), + }, + 'save': { + 'type': 'string', + 'description': 'Optional local path to save the decoded image to.', + }, + }, + 'required': [], + }, + } +] + + +def resolve_data_url(target): + """Return (data_url, meta) for the requested target.""" + if target and target != 'latest' and not target.startswith('session:'): + if os.path.isfile(target): + ext = os.path.splitext(target)[1].lower() + mime = {'.png': 'image/png', '.jpg': 'image/jpeg', '.jpeg': 'image/jpeg', + '.gif': 'image/gif', '.webp': 'image/webp'}.get(ext, 'image/png') + b64 = base64.b64encode(open(target, 'rb').read()).decode() + return 'data:%s;base64,%s' % (mime, b64), {'source': 'file', 'filename': target} + db = sqlite3.connect(DB, timeout=10) + row = db.execute('SELECT id, data FROM part WHERE id=?', (target,)).fetchone() + if not row: + raise ValueError('part id not found: %s' % target) + pid, data = row + else: + db = sqlite3.connect(DB, timeout=10) + if target and target.startswith('session:'): + row = db.execute( + "SELECT id, data FROM part WHERE session_id=? " + "AND data LIKE '%\"type\":\"file\"%' AND data LIKE '%data:image%' " + "ORDER BY time_created DESC LIMIT 1", (target[8:],)).fetchone() + else: + row = db.execute( + "SELECT id, data FROM part WHERE data LIKE '%\"type\":\"file\"%' " + "AND data LIKE '%data:image%' ORDER BY time_created DESC LIMIT 1").fetchone() + if not row: + scope = ' in session %s' % target[8:] if target and target.startswith('session:') else '' + raise ValueError('no image part found%s' % scope) + pid, data = row + d = json.loads(data) + url = d.get('url', '') + if not url.startswith('data:image'): + raise ValueError('part %s has no inline image data' % pid) + return url, {'source': 'opencode.db', 'part': pid, 'filename': d.get('filename')} + + +def run_ocr(data_url, prompt): + body = json.dumps({ + 'model': 'Unlimited-OCR', + 'temperature': 0, + 'repeat_penalty': 1.2, + 'max_tokens': 4096, + 'messages': [{'role': 'user', 'content': [ + {'type': 'image_url', 'image_url': {'url': data_url}}, + {'type': 'text', 'text': prompt}, + ]}], + }).encode() + req = urllib.request.Request(OCR_URL, data=body, + headers={'Content-Type': 'application/json'}) + with urllib.request.urlopen(req, timeout=1800) as r: + out = json.load(r) + return out['choices'][0]['message']['content'] + + +def tool_ocr_image(args): + target = args.get('target', 'latest') + prompt = args.get('prompt', 'Free OCR.') + save = args.get('save') + data_url, meta = resolve_data_url(target) + if save: + open(save, 'wb').write(base64.b64decode(data_url.split(',', 1)[1])) + meta['saved_to'] = save + text = run_ocr(data_url, prompt) + header = 'source=%s part=%s file=%s' % (meta.get('source'), meta.get('part', '-'), meta.get('filename', '-')) + if meta.get('saved_to'): + header += ' saved_to=%s' % meta['saved_to'] + return header + '\n' + text + + +def dispatch(msg): + method = msg.get('method') + mid = msg.get('id') + if method is None: + return None + if method == 'initialize': + params = msg.get('params', {}) + return { + 'protocolVersion': params.get('protocolVersion', '2025-06-18'), + 'capabilities': {'tools': {}}, + 'serverInfo': SERVER_INFO, + } + if method == 'notifications/initialized': + return None + if method == 'ping': + return {} + if method == 'tools/list': + return {'tools': TOOLS} + if method == 'tools/call': + params = msg.get('params', {}) + name = params.get('name') + args = params.get('arguments', {}) or {} + if name != 'ocr_image': + return None, {'code': -32602, 'message': 'unknown tool: %s' % name} + try: + text = tool_ocr_image(args) + return {'content': [{'type': 'text', 'text': text}], 'isError': False}, None + except Exception as e: + return {'content': [{'type': 'text', 'text': 'ocr_image failed: %s: %s' % (type(e).__name__, e)}], + 'isError': True}, None + if method in ('resources/list', 'prompts/list'): + return {'resources': []} if method == 'resources/list' else {'prompts': []} + if mid is None: + return None + return None, {'code': -32601, 'message': 'method not supported: %s' % method} + + +def main(): + for line in sys.stdin: + line = line.strip() + if not line: + continue + try: + msg = json.loads(line) + except json.JSONDecodeError: + continue + result = dispatch(msg) + if result is None: + continue + if isinstance(result, tuple): + payload, error = result + if error is not None: + out = {'jsonrpc': '2.0', 'id': msg.get('id'), 'error': error} + else: + out = {'jsonrpc': '2.0', 'id': msg.get('id'), 'result': payload} + else: + out = {'jsonrpc': '2.0', 'id': msg.get('id'), 'result': result} + sys.stdout.write(json.dumps(out) + '\n') + sys.stdout.flush() + + +if __name__ == '__main__': + main() diff --git a/ocr-posted.py b/ocr-posted.py new file mode 100755 index 0000000..e8e0998 --- /dev/null +++ b/ocr-posted.py @@ -0,0 +1,76 @@ +#!/usr/bin/env python3 +"""Pull an image posted into opencode from opencode.db and OCR it via local Unlimited-OCR. + +Usage: + ocr-posted.py [latest | | session:] [prompt] [save_path] + +Defaults: latest image part in DB, prompt "Free OCR." +Env: OPENCODE_DB, OCR_URL +""" +import base64, json, os, sqlite3, sys, urllib.request + +DB = os.environ.get( + 'OPENCODE_DB', + os.path.expanduser('~/.local/share/opencode/opencode.db')) +OCR_URL = os.environ.get('OCR_URL', 'http://127.0.0.1:11436/v1/chat/completions') + + +def pick(db, arg): + cur = db.cursor() + if arg and arg != 'latest': + if arg.startswith('session:'): + cur.execute( + "SELECT id, data FROM part WHERE session_id=? " + "AND data LIKE '%\"type\":\"file\"%' AND data LIKE '%data:image%' " + "ORDER BY time_created DESC LIMIT 1", + (arg[8:],)) + else: + cur.execute("SELECT id, data FROM part WHERE id=?", (arg,)) + else: + cur.execute( + "SELECT id, data FROM part WHERE data LIKE '%\"type\":\"file\"%' " + "AND data LIKE '%data:image%' ORDER BY time_created DESC LIMIT 1") + return cur.fetchone() + + +def ocr(data_url, prompt, max_tokens=4096): + body = json.dumps({ + "model": "Unlimited-OCR", + "temperature": 0, + "repeat_penalty": 1.2, + "max_tokens": max_tokens, + "messages": [{"role": "user", "content": [ + {"type": "image_url", "image_url": {"url": data_url}}, + {"type": "text", "text": prompt}, + ]}], + }).encode() + req = urllib.request.Request( + OCR_URL, data=body, headers={"Content-Type": "application/json"}) + with urllib.request.urlopen(req, timeout=1800) as r: + return json.load(r)['choices'][0]['message']['content'] + + +def main(): + arg = sys.argv[1] if len(sys.argv) > 1 else 'latest' + prompt = sys.argv[2] if len(sys.argv) > 2 else 'Free OCR.' + save = sys.argv[3] if len(sys.argv) > 3 else None + db = sqlite3.connect(DB, timeout=10) + row = pick(db, arg) + if not row: + print('no image part found', file=sys.stderr) + sys.exit(1) + pid, data = row + d = json.loads(data) + url = d.get('url', '') + if not url.startswith('data:image'): + print('part %s has no inline image data' % pid, file=sys.stderr) + sys.exit(1) + if save: + open(save, 'wb').write(base64.b64decode(url.split(',', 1)[1])) + print('saved image to %s' % save, file=sys.stderr) + print('part: %s file: %s' % (pid, d.get('filename')), file=sys.stderr) + print(ocr(url, prompt)) + + +if __name__ == '__main__': + main() diff --git a/run-ocr-server.sh b/run-ocr-server.sh new file mode 100755 index 0000000..2f71e3b --- /dev/null +++ b/run-ocr-server.sh @@ -0,0 +1,36 @@ +#!/usr/bin/env bash +# Launch baidu/Unlimited-OCR via llama-server (OpenAI-compatible, multimodal). +# Foreground — Ctrl+C to stop. All paths overridable via env. +set -euo pipefail + +export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-1}" + +BIN="${LLAMA_SERVER_BIN:-/usr/local/bin/llama-server}" +MODEL="${OCR_MODEL:-$HOME/models/Unlimited-OCR-Q4_K_M.gguf}" +MMPROJ="${OCR_MMPROJ:-$HOME/models/mmproj-Unlimited-OCR-F16.gguf}" +HOST="${OCR_HOST:-127.0.0.1}" +PORT="${OCR_PORT:-11436}" +CTX="${OCR_CTX:-32768}" +TEMP="${OCR_TEMP:-0}" +REPEAT_PENALTY="${OCR_REPEAT_PENALTY:-1.2}" + +[[ -x $BIN ]] || { echo "MISSING (or not executable): $BIN" >&2; exit 1; } +[[ -f $MODEL ]] || { echo "MISSING: $MODEL" >&2; exit 1; } +[[ -f $MMPROJ ]] || { echo "MISSING: $MMPROJ" >&2; exit 1; } + +exec "$BIN" \ + --model "$MODEL" \ + --mmproj "$MMPROJ" \ + --alias Unlimited-OCR \ + --host "$HOST" \ + --port "$PORT" \ + --temp "$TEMP" \ + --repeat-penalty "$REPEAT_PENALTY" \ + --ctx-size "$CTX" \ + --n-gpu-layers 99 \ + --flash-attn on \ + --cont-batching \ + --parallel 1 \ + --cache-type-k q4_0 \ + --cache-type-v q4_0 \ + --metrics diff --git a/skill/ocr-image/SKILL.md b/skill/ocr-image/SKILL.md new file mode 100644 index 0000000..09921d6 --- /dev/null +++ b/skill/ocr-image/SKILL.md @@ -0,0 +1,68 @@ +--- +name: ocr-image +description: Read images the user posts/attaches in opencode when the current model cannot see them. Use when an image arrives with a placeholder like "Cannot read (this model does not support image input)", or when the user asks to read, OCR, transcribe, or summarize a posted image, screenshot, photo, or document. Calls the local Unlimited-OCR VLM via the ocr_image MCP tool. +license: MIT +--- + +# OCR Posted Images + +A local Unlimited-OCR VLM (llama-server, OpenAI-compatible) reads images the text-only model cannot see. + +## When to use + +- User posts/attaches an image and you cannot see it — the message shows a placeholder such as + `Cannot read image.png (this model does not support image input)`. The image bytes are safe in + `opencode.db`; only your context lost them. +- User says "read/OCR/transcribe/summarize the image I posted" or similar. +- If you CAN see the image directly (vision model), do not OCR it. + +## Tool + +Preferred: MCP tool `ocr_image` (server `ocr`): + +| arg | meaning | +|---|---| +| `target` | `latest` (default, most recent image part in opencode.db) · part id `prt_...` · `session:` · local file path | +| `prompt` | instruction to the OCR model (see below) | +| `save` | optional path to persist the decoded image | + +Fallback (MCP not loaded yet, e.g. before opencode restart): run the CLI from this repo: + +```bash +python3 /path/to/unlimited-ocr-mcp/ocr-posted.py [latest|part_id|session:] [prompt] [save_path] +``` + +## Prompt choice + +| task | prompt | +|---|---| +| transcribe document/screenshot | `Free OCR.` (default) | +| structured layout + coords | `<\|grounding\|>Convert the document to markdown.` | +| visual description | `Describe this image in detail.` | +| locate element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` | + +Pick the prompt from the user's intent, not from habit. + +## Output format + +Lines look like `type [x1,y1,x2,y2]content` (types: header, text, table, footer, image, +plus a `source=... part=... file=...` header line). `` marks an image region. +Strip layout prefixes and the source header before presenting — give the user clean text. + +## Caveats + +- Dense tables: the model can degenerate into repeating the last row (garbled repetition). + If the tail looks looped, re-run the same call — sampling varies per run. +- Animated images (gif): first frame only. +- Sampling is fixed by the server: temp 0, repeat_penalty 1.2. Do not override. +- Backend down? `curl -s http://127.0.0.1:11436/health` → if unreachable, start it with + `run-ocr-server.sh` from this repo (set `CUDA_VISIBLE_DEVICES`, `OCR_MODEL`, `OCR_MMPROJ`). + If the GPU is shared with other models and VRAM is tight, ask the user before stopping anything. + +## Example + +User: [posts screenshot] "what does this say?" +→ `ocr_image {target: "latest", prompt: "Free OCR."}` → strip layout tokens → answer with the text. + +User: "ocr the invoice I sent, I need the total" +→ `ocr_image {target: "latest"}` → extract the total from the transcription.