The standard DeepSeek V4 Flash route is text-only. DeepSeek also offers a separate V4 Flash Vision experimental model, so check which route your harness uses. A text-only route cannot inspect a mockup or compare a screenshot with a design without another vision tool.

What ModLens Does

ModLens is a plugin and skill-based bridge for supported agent harnesses. It sends an image to a configured vision engine and returns structured evidence to the text-only model.

Text-only model
  -> ModLens
    -> Configured vision engine
      -> Structured JSON evidence
        -> Text-only model

The result includes OCR text, layout regions, entities, relationships, and uncertainty that the calling model can inspect.

Install It

For Codex, Claude Code, Pi, or OpenCode, follow the current ModLens installation guide and run its health check before using an image.

For DeepSeek Harness, use the current plugin command in the ModLens README.

Paste an image path into the agent conversation and ask a focused question about its contents. The exact trigger and provider depend on the harness and configuration you chose.

Privacy Boundary

ModLens sends the image to the vision engine you configure. Review that provider's terms, quotas, and data handling before using private screenshots.

Sources