DeepSeek added DeepSeek-V4-Flash-Vision-Exp to its API on August 21, 2026, and announced Harness 0.1.1 support the same day. The useful change is a documented path for image-aware agent workflows. The experimental label remains important when deciding what traffic to send through it.
Image inputs arrive through existing API surfaces
The release announcement identifies the experimental model name and supports mixed text and images through Chat Completions, Messages and Responses. Images can be supplied as base64, URLs or Files API references. The provider documents up to 384 billed tokens per image at V4-Flash pricing.
API-shape compatibility does not guarantee equivalent behavior between interfaces or existing text workflows. Check serialization, streamed responses and tool calls in the actual client. Keep a small set of image fixtures that can distinguish request-shape errors from model interpretation errors.
File reuse saves upload work, not every inference cost
The announcement also introduces a free-to-use Files API for uploading an image once and reusing its identifier. That reduces repeated upload bandwidth. It should not be interpreted as making the model's image processing free.
For sensitive images, decide who may upload, reference and remove a file before exposing reuse in a product. Track file lifecycle alongside request ownership, and avoid passing arbitrary remote image URLs without considering the application's existing fetch and access controls.
Harness is an execution layer to review
The Harness repository describes a plugin-based architecture and MIT licensing. The dated release artifact provides a concrete version record around the integration. The August 21 announcement establishes the model-support milestone, not the project's first-ever public launch.
A plugin system can simplify extension, but its permissions are part of the application. Review installed plugins, execution privileges and network access. Pin versions and retain an explicit inventory rather than letting an agent discover unrestricted capabilities at runtime.
Use visual grounding without overclaiming readiness
Start with verifiable tasks such as reading a chart label or identifying the next step in a screenshot. Include misleading visual context and require the agent to distinguish what it sees from what it assumes. Source screenshots can contain instructions that should be treated as task data rather than authority.
The provider reports improved multimodal benchmark performance, but that does not establish accuracy on another team's documents or interfaces. Keep production rollout bounded, record failures and maintain a fallback path already validated for the task. Experimental access expands evaluation options without proving dependable autonomous operation.