--- name: ocr-image description: Read images the user posts/attaches in opencode when the current model cannot see them. Use when an image arrives with a placeholder like "Cannot read (this model does not support image input)", or when the user asks to read, OCR, transcribe, or summarize a posted image, screenshot, photo, or document. Calls the local Unlimited-OCR VLM via the ocr_image MCP tool. license: MIT --- # OCR Posted Images A local Unlimited-OCR VLM (llama-server, OpenAI-compatible) reads images the text-only model cannot see. ## When to use - User posts/attaches an image and you cannot see it — the message shows a placeholder such as `Cannot read image.png (this model does not support image input)`. The image bytes are safe in `opencode.db`; only your context lost them. - User says "read/OCR/transcribe/summarize the image I posted" or similar. - If you CAN see the image directly (vision model), do not OCR it. ## Tool Preferred: MCP tool `ocr_image` (server `ocr`): | arg | meaning | |---|---| | `target` | `latest` (default, most recent image part in opencode.db) · part id `prt_...` · `session:` · local file path | | `prompt` | instruction to the OCR model (see below) | | `save` | optional path to persist the decoded image | Fallback (MCP not loaded yet, e.g. before opencode restart): run the CLI from this repo: ```bash python3 /path/to/unlimited-ocr-mcp/ocr-posted.py [latest|part_id|session:] [prompt] [save_path] ``` ## Prompt choice | task | prompt | |---|---| | transcribe document/screenshot | `Free OCR.` (default) | | structured layout + coords | `<\|grounding\|>Convert the document to markdown.` | | visual description | `Describe this image in detail.` | | locate element | `<\|grounding\|>Locate <\|ref\|>TEXT<\|/ref\|> in the image.` | Pick the prompt from the user's intent, not from habit. ## Output format Lines look like `type [x1,y1,x2,y2]content` (types: header, text, table, footer, image, plus a `source=... part=... file=...` header line). `` marks an image region. Strip layout prefixes and the source header before presenting — give the user clean text. ## Caveats - Dense tables: the model can degenerate into repeating the last row (garbled repetition). If the tail looks looped, re-run the same call — sampling varies per run. - Animated images (gif): first frame only. - Sampling is fixed by the server: temp 0, repeat_penalty 1.2. Do not override. - Backend down? `curl -s http://127.0.0.1:11436/health` → if unreachable, start it with `run-ocr-server.sh` from this repo (set `CUDA_VISIBLE_DEVICES`, `OCR_MODEL`, `OCR_MMPROJ`). If the GPU is shared with other models and VRAM is tight, ask the user before stopping anything. ## Example User: [posts screenshot] "what does this say?" → `ocr_image {target: "latest", prompt: "Free OCR."}` → strip layout tokens → answer with the text. User: "ocr the invoice I sent, I need the total" → `ocr_image {target: "latest"}` → extract the total from the transcription.