Open OCR CLI
Provides provider-neutral multimodal OCR for images, PDFs, and URLs, offering structured data extraction and agentic workflows via various AI models.
Provides provider-neutral multimodal OCR for images, PDFs, and URLs, offering structured data extraction and agentic workflows via various AI models.
Open OCR CLI offers a powerful, agent-first, and provider-neutral solution for extracting information from images, PDFs, and public URLs. It supports various AI models like Gemini, Kimi, Muse, OpenRouter, and any OpenAI-compatible endpoint, allowing users to convert documents into plain text or validated structured data. Beyond simple extraction, it facilitates complex agentic workflows and can serve OCR tools over the Model Context Protocol (MCP) via a stdio transport, making it highly versatile for integration into automated systems and developer workflows.