- mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool - ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server) - run-ocr-server.sh: llama-server launcher with tuned sampling flags - skill/ocr-image: opencode skill teaching agents when/how to call it - README with full install instructions
69 lines
3.0 KiB
Markdown
69 lines
3.0 KiB
Markdown
---
|
|
name: ocr-image
|
|
description: Read images the user posts/attaches in opencode when the current model cannot see them. Use when an image arrives with a placeholder like "Cannot read <file> (this model does not support image input)", or when the user asks to read, OCR, transcribe, or summarize a posted image, screenshot, photo, or document. Calls the local Unlimited-OCR VLM via the ocr_image MCP tool.
|
|
license: MIT
|
|
---
|
|
|
|
# OCR Posted Images
|
|
|
|
A local Unlimited-OCR VLM (llama-server, OpenAI-compatible) reads images the text-only model cannot see.
|
|
|
|
## When to use
|
|
|
|
- User posts/attaches an image and you cannot see it — the message shows a placeholder such as
|
|
`Cannot read image.png (this model does not support image input)`. The image bytes are safe in
|
|
`opencode.db`; only your context lost them.
|
|
- User says "read/OCR/transcribe/summarize the image I posted" or similar.
|
|
- If you CAN see the image directly (vision model), do not OCR it.
|
|
|
|
## Tool
|
|
|
|
Preferred: MCP tool `ocr_image` (server `ocr`):
|
|
|
|
| arg | meaning |
|
|
|---|---|
|
|
| `target` | `latest` (default, most recent image part in opencode.db) · part id `prt_...` · `session:<id>` · local file path |
|
|
| `prompt` | instruction to the OCR model (see below) |
|
|
| `save` | optional path to persist the decoded image |
|
|
|
|
Fallback (MCP not loaded yet, e.g. before opencode restart): run the CLI from this repo:
|
|
|
|
```bash
|
|
python3 /path/to/unlimited-ocr-mcp/ocr-posted.py [latest|part_id|session:<id>] [prompt] [save_path]
|
|
```
|
|
|
|
## Prompt choice
|
|
|
|
| task | prompt |
|
|
|---|---|
|
|
| transcribe document/screenshot | `Free OCR.` (default) |
|
|
| structured layout + coords | `<\|grounding\|>Convert the document to markdown.` |
|
|
| visual description | `Describe this image in detail.` |
|
|
| locate element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` |
|
|
|
|
Pick the prompt from the user's intent, not from habit.
|
|
|
|
## Output format
|
|
|
|
Lines look like `type [x1,y1,x2,y2]content` (types: header, text, table, footer, image,
|
|
plus a `source=... part=... file=...` header line). `<img>` marks an image region.
|
|
Strip layout prefixes and the source header before presenting — give the user clean text.
|
|
|
|
## Caveats
|
|
|
|
- Dense tables: the model can degenerate into repeating the last row (garbled repetition).
|
|
If the tail looks looped, re-run the same call — sampling varies per run.
|
|
- Animated images (gif): first frame only.
|
|
- Sampling is fixed by the server: temp 0, repeat_penalty 1.2. Do not override.
|
|
- Backend down? `curl -s http://127.0.0.1:11436/health` → if unreachable, start it with
|
|
`run-ocr-server.sh` from this repo (set `CUDA_VISIBLE_DEVICES`, `OCR_MODEL`, `OCR_MMPROJ`).
|
|
If the GPU is shared with other models and VRAM is tight, ask the user before stopping anything.
|
|
|
|
## Example
|
|
|
|
User: [posts screenshot] "what does this say?"
|
|
→ `ocr_image {target: "latest", prompt: "Free OCR."}` → strip layout tokens → answer with the text.
|
|
|
|
User: "ocr the invoice I sent, I need the total"
|
|
→ `ocr_image {target: "latest"}` → extract the total from the transcription.
|