DeepSeek released DeepSeek V4.1 Flash on September 10, 2026. The model represents a significant architectural change from earlier DeepSeek systems. It combines a large total parameter count with relatively small active computation for each token.
DeepSeek V4.1 Flash is a 552 billion parameter Mixture-of-Experts model. Its new Causal Encoder-Decoder architecture activates about 8 billion parameters while processing input and 16 billion during output generation. It also includes native visual understanding rather than depending on a separate experimental vision model.
AI made PMs faster. Multiplayer mode is still broken.

A PM can summarize research, draft a PRD, and mock up a prototype before lunch. The hard part starts when the team has to decide what actually gets built.
Jira Product Discovery gives product teams one place to capture insights, prioritize ideas with consistent frameworks, and build living roadmaps stakeholders can rally around.
And because it’s connected to Jira, the context behind every decision stays with the work—so developers and their agents know not just what to build, but why.
AI helps PMs move faster. Jira Product Discovery helps the whole team build with confidence.
The model supports a one-million-token context window. DeepSeek also provides tool calling, structured JSON output, Responses API compatibility and an Anthropic-compatible API interface. Its reasoning effort can be controlled continuously from 1 to 100. This differs from systems that provide only several fixed reasoning levels.
Coding and agent functions are an important part of V4.1 Flash. DeepSeek reports scores of 90.6 percent on Terminal-Bench 2.1 and 74.2 percent on DeepSWE v1.1. On the same DeepSeek evaluation table, GPT-5.6 Sol scored 88.8 and 73.0 percent respectively. DeepSeek also reported 54.8 percent on AutomationBench and 31.8 percent on Agent's Last Exam. These results were produced using specified agent scaffolds and maximum reasoning effort, so they should not be treated as independent rankings.
API pricing creates another distinction. DeepSeek V4.1 Flash costs $0.15 per million uncached input tokens and $0.60 per million output tokens during off-peak periods. Cached input costs $0.003 per million tokens. Peak pricing doubles these figures to $0.30 for uncached input, $0.006 for cached input and $1.20 for output.
OpenAI's GPT-5.6 Luna occupies a similar high-volume API category. It currently costs $0.20 per million input tokens, $0.02 for cached input and $1.20 for output. Luna provides a 1.05-million-token context window and up to 128,000 output tokens. Its managed tool support includes web search, file search, Code Interpreter, hosted shell, computer use, MCP and image generation. It also supports multiple reasoning-effort levels.
200+ Proven Ways to Make Money With AI in 2026
The next wave of millionaires will be people who figured out how to make AI work for them.
The window to get ahead is still open. But not for long.
Here are 200+ proven ways to make money with AI in 2026.
Sign up for Superhuman AI, the free daily newsletter read by 1M+ professionals, and get instant access to all 200+ ways to profit from AI this year.
Google's Gemini 3.1 Flash-Lite costs $0.25 per million text, image or video input tokens and $1.50 per million output tokens. It provides a one-million-token input context and supports multimodal processing. Google's API environment also provides Search and Maps grounding, code execution, URL context and file-related functions, although some capabilities carry separate charges.
Claude Haiku 4.5 has a different cost and context profile. Anthropic lists standard pricing at $1 per million input tokens and $5 per million output tokens, with cache reads at $0.10 per million. Its context window is 200,000 tokens and maximum output is 64,000 tokens. Haiku supports image and text input, tool use and extended reasoning, but its context capacity is smaller than the other three models considered here.
The comparison shows different design priorities. DeepSeek V4.1 Flash combines a very large sparse model with low active parameter counts, long context, native vision and agent-oriented functions. GPT-5.6 Luna provides a broader managed tool environment at a higher output-token price. Gemini 3.1 Flash-Lite combines long context with wider native media input. Claude Haiku 4.5 has a shorter context window and higher token rates, while remaining part of Anthropic's established tool and reasoning system.
For developers, token price alone does not determine operating cost. Reasoning length, cache reuse, tool calls, output size, retry rates and task completion accuracy can materially change the cost of running the same workload across these models.
Hire Ava, the AI BDR built for enterprise
Ava is the first AI BDR to run outbound end to end, finding leads or ingesting your CRM accounts, sending personalized emails on your reps' behalf, and booking meetings, autonomously or on copilot. She runs outbound for DoorDash and Grammarly. She's SOC 2 Type II audited, SSO and GDPR ready.




