Jarian Cottingham c0fa9148b7 Initial commit: Unlimited-OCR MCP for opencode
- mcp/server.py: stdlib-only MCP stdio server exposing ocr_image tool
- ocr-posted.py: CLI equivalent (DB lookup + OCR via llama-server)
- run-ocr-server.sh: llama-server launcher with tuned sampling flags
- skill/ocr-image: opencode skill teaching agents when/how to call it
- README with full install instructions
2026-08-22 06:24:57 +00:00

3.0 KiB

name description license
ocr-image Read images the user posts/attaches in opencode when the current model cannot see them. Use when an image arrives with a placeholder like "Cannot read <file> (this model does not support image input)", or when the user asks to read, OCR, transcribe, or summarize a posted image, screenshot, photo, or document. Calls the local Unlimited-OCR VLM via the ocr_image MCP tool. MIT

OCR Posted Images

A local Unlimited-OCR VLM (llama-server, OpenAI-compatible) reads images the text-only model cannot see.

When to use

  • User posts/attaches an image and you cannot see it — the message shows a placeholder such as Cannot read image.png (this model does not support image input). The image bytes are safe in opencode.db; only your context lost them.
  • User says "read/OCR/transcribe/summarize the image I posted" or similar.
  • If you CAN see the image directly (vision model), do not OCR it.

Tool

Preferred: MCP tool ocr_image (server ocr):

arg meaning
target latest (default, most recent image part in opencode.db) · part id prt_... · session:<id> · local file path
prompt instruction to the OCR model (see below)
save optional path to persist the decoded image

Fallback (MCP not loaded yet, e.g. before opencode restart): run the CLI from this repo:

python3 /path/to/unlimited-ocr-mcp/ocr-posted.py [latest|part_id|session:<id>] [prompt] [save_path]

Prompt choice

task prompt
transcribe document/screenshot Free OCR. (default)
structured layout + coords <|grounding|>Convert the document to markdown.
visual description Describe this image in detail.
locate element <|grounding|>Locate <|ref|>TEXT<|/ref|> in the image.

Pick the prompt from the user's intent, not from habit.

Output format

Lines look like type [x1,y1,x2,y2]content (types: header, text, table, footer, image, plus a source=... part=... file=... header line). <img> marks an image region. Strip layout prefixes and the source header before presenting — give the user clean text.

Caveats

  • Dense tables: the model can degenerate into repeating the last row (garbled repetition). If the tail looks looped, re-run the same call — sampling varies per run.
  • Animated images (gif): first frame only.
  • Sampling is fixed by the server: temp 0, repeat_penalty 1.2. Do not override.
  • Backend down? curl -s http://127.0.0.1:11436/health → if unreachable, start it with run-ocr-server.sh from this repo (set CUDA_VISIBLE_DEVICES, OCR_MODEL, OCR_MMPROJ). If the GPU is shared with other models and VRAM is tight, ask the user before stopping anything.

Example

User: [posts screenshot] "what does this say?" → ocr_image {target: "latest", prompt: "Free OCR."} → strip layout tokens → answer with the text.

User: "ocr the invoice I sent, I need the total" → ocr_image {target: "latest"} → extract the total from the transcription.