DeepSeek V4.1 Flash and inclusionAI Ling 3.0 Flash-VL add open-weight multimodal options to September’s model landscape. Their official model cards identify MIT licensing and different architectures and input capabilities. They are candidates for controlled deployment, not automatically lightweight replacements for hosted APIs.
The DeepSeek card describes image-and-text input and a context of up to one million tokens. The Ling card describes image and video input with up to 256K context. The DeepSeek repository history records creation and uploads on September 10, and Ling’s history records its initial commit on September 4; those are artifact milestones, not proof of identical public announcement dates.
The models solve different serving problems
DeepSeek emphasizes compressed cache architecture for input-heavy agent work. Ling describes a native multimodal model with video support and provides serving recipes. A team should match the input it needs before comparing benchmark scores.
A scanned document, a sequence of video frames, and a screenshot-driven agent expose different failure modes. Test actual formats, resolution, and context length. The largest supported context is a limit to validate, not a recommendation to send all organizational data in every request.
Activated parameters are not the deployment footprint
The cards describe mixture-of-experts models with a much smaller activated parameter count than total parameters. That can affect compute per token, but the complete weights and serving runtime still have to fit the deployment architecture.
Include model storage, memory, cache growth, concurrency, and interconnect needs in a capacity test. Quantization and runtime support can also change output quality. Do not choose hardware solely by comparing the activated parameter count with a dense model’s size.
MIT licensing does not complete the deployment review
Both primary cards identify MIT licensing. Preserve the license and attribution requirements for the exact downloaded revision. The artifact’s license should be evaluated separately from any hosted provider’s terms or a derivative quantization supplied by another party.
Operational control also requires a patching and recovery owner. Downloadable weights let a team own the serving path, but that team then needs to investigate failed requests, manage upgrades, and protect data access. Open availability does not supply those processes automatically.
How to compare the options fairly
Use the same bounded task set where capabilities overlap. Measure answer correctness, visual grounding, tool behavior, latency, and complete infrastructure cost. Run video-specific tasks separately because DeepSeek’s documented image/text interface should not be credited with a feature it does not describe.
Nerova’s assessment is that the two releases broaden the practical open-model shortlist. Their value comes from fit and controllable deployment, with hardware and verification requirements made explicit. The right choice is the one that completes the intended workload reliably within the organization’s real capacity.