- mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool - ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server) - run-ocr-server.sh: llama-server launcher with tuned sampling flags - skill/ocr-image: opencode skill teaching agents when/how to call it - README with full install instructions
3.0 KiB
| name | description | license |
|---|---|---|
| ocr-image | Read images the user posts/attaches in opencode when the current model cannot see them. Use when an image arrives with a placeholder like "Cannot read <file> (this model does not support image input)", or when the user asks to read, OCR, transcribe, or summarize a posted image, screenshot, photo, or document. Calls the local Unlimited-OCR VLM via the ocr_image MCP tool. | MIT |
OCR Posted Images
A local Unlimited-OCR VLM (llama-server, OpenAI-compatible) reads images the text-only model cannot see.
When to use
- User posts/attaches an image and you cannot see it — the message shows a placeholder such as
Cannot read image.png (this model does not support image input). The image bytes are safe inopencode.db; only your context lost them. - User says "read/OCR/transcribe/summarize the image I posted" or similar.
- If you CAN see the image directly (vision model), do not OCR it.
Tool
Preferred: MCP tool ocr_image (server ocr):
| arg | meaning |
|---|---|
target |
latest (default, most recent image part in opencode.db) · part id prt_... · session:<id> · local file path |
prompt |
instruction to the OCR model (see below) |
save |
optional path to persist the decoded image |
Fallback (MCP not loaded yet, e.g. before opencode restart): run the CLI from this repo:
python3 /path/to/unlimited-ocr-mcp/ocr-posted.py [latest|part_id|session:<id>] [prompt] [save_path]
Prompt choice
| task | prompt |
|---|---|
| transcribe document/screenshot | Free OCR. (default) |
| structured layout + coords | <|grounding|>Convert the document to markdown. |
| visual description | Describe this image in detail. |
| locate element | <|grounding|>Locate <|ref|>TEXT<|/ref|> in the image. |
Pick the prompt from the user's intent, not from habit.
Output format
Lines look like type [x1,y1,x2,y2]content (types: header, text, table, footer, image,
plus a source=... part=... file=... header line). <img> marks an image region.
Strip layout prefixes and the source header before presenting — give the user clean text.
Caveats
- Dense tables: the model can degenerate into repeating the last row (garbled repetition). If the tail looks looped, re-run the same call — sampling varies per run.
- Animated images (gif): first frame only.
- Sampling is fixed by the server: temp 0, repeat_penalty 1.2. Do not override.
- Backend down?
curl -s http://127.0.0.1:11436/health→ if unreachable, start it withrun-ocr-server.shfrom this repo (setCUDA_VISIBLE_DEVICES,OCR_MODEL,OCR_MMPROJ). If the GPU is shared with other models and VRAM is tight, ask the user before stopping anything.
Example
User: [posts screenshot] "what does this say?"
→ ocr_image {target: "latest", prompt: "Free OCR."} → strip layout tokens → answer with the text.
User: "ocr the invoice I sent, I need the total"
→ ocr_image {target: "latest"} → extract the total from the transcription.