Jarian Cottingham c0fa9148b7 Initial commit: Unlimited-OCR MCP for opencode
- mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool
- ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server)
- run-ocr-server.sh: llama-server launcher with tuned sampling flags
- skill/ocr-image: opencode skill teaching agents when/how to call it
- README with full install instructions
2026-08-22 06:24:57 +00:00

69 lines
3.0 KiB
Markdown

---
name: ocr-image
description: Read images the user posts/attaches in opencode when the current model cannot see them. Use when an image arrives with a placeholder like "Cannot read <file> (this model does not support image input)", or when the user asks to read, OCR, transcribe, or summarize a posted image, screenshot, photo, or document. Calls the local Unlimited-OCR VLM via the ocr_image MCP tool.
license: MIT
---
# OCR Posted Images
A local Unlimited-OCR VLM (llama-server, OpenAI-compatible) reads images the text-only model cannot see.
## When to use
- User posts/attaches an image and you cannot see it — the message shows a placeholder such as
`Cannot read image.png (this model does not support image input)`. The image bytes are safe in
`opencode.db`; only your context lost them.
- User says "read/OCR/transcribe/summarize the image I posted" or similar.
- If you CAN see the image directly (vision model), do not OCR it.
## Tool
Preferred: MCP tool `ocr_image` (server `ocr`):
| arg | meaning |
|---|---|
| `target` | `latest` (default, most recent image part in opencode.db) · part id `prt_...` · `session:<id>` · local file path |
| `prompt` | instruction to the OCR model (see below) |
| `save` | optional path to persist the decoded image |
Fallback (MCP not loaded yet, e.g. before opencode restart): run the CLI from this repo:
```bash
python3 /path/to/unlimited-ocr-mcp/ocr-posted.py [latest|part_id|session:<id>] [prompt] [save_path]
```
## Prompt choice
| task | prompt |
|---|---|
| transcribe document/screenshot | `Free OCR.` (default) |
| structured layout + coords | `<\|grounding\|>Convert the document to markdown.` |
| visual description | `Describe this image in detail.` |
| locate element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` |
Pick the prompt from the user's intent, not from habit.
## Output format
Lines look like `type [x1,y1,x2,y2]content` (types: header, text, table, footer, image,
plus a `source=... part=... file=...` header line). `<img>` marks an image region.
Strip layout prefixes and the source header before presenting — give the user clean text.
## Caveats
- Dense tables: the model can degenerate into repeating the last row (garbled repetition).
If the tail looks looped, re-run the same call — sampling varies per run.
- Animated images (gif): first frame only.
- Sampling is fixed by the server: temp 0, repeat_penalty 1.2. Do not override.
- Backend down? `curl -s http://127.0.0.1:11436/health` → if unreachable, start it with
`run-ocr-server.sh` from this repo (set `CUDA_VISIBLE_DEVICES`, `OCR_MODEL`, `OCR_MMPROJ`).
If the GPU is shared with other models and VRAM is tight, ask the user before stopping anything.
## Example
User: [posts screenshot] "what does this say?"
→ `ocr_image {target: "latest", prompt: "Free OCR."}` → strip layout tokens → answer with the text.
User: "ocr the invoice I sent, I need the total"
→ `ocr_image {target: "latest"}` → extract the total from the transcription.