Independent timing finally put a number on the Flash war. On 23 September 2026, MRKT30 ran 320 streamed API calls from Budapest comparing DeepSeek V4.1 Flash and Qwen3.8 Flash through OpenRouter, pinned to each lab’s own endpoint with fallbacks off. Headline: DeepSeek answered 2.2× to 3.2× faster in every like-for-like setting. On default afternoon settings, median complete-answer time was 2.8s for DeepSeek versus 7.9s for Qwen; with thinking disabled, 1.4s versus 3.4s. DeepSeek also streamed roughly four times as many answer tokens per second.
The more operational finding is nastier than the stopwatch. Both models think by default, and hidden reasoning tokens share the same max_tokens budget as the visible answer. With a 400-token cap, DeepSeek returned empty on 19/40 calls and Qwen on 13/40; on a JSON task DeepSeek blanked 8/8. Raising the cap to 2,000 still left about one call in ten empty. Switching thinking off was the only setting that cleared empties across 120 calls each. Reasoning ate ~83% of DeepSeek’s billed output tokens and ~78% of Qwen’s when thinking stayed on. That is not a toy footnote — it is how Flash economics explode in production agent loops that “just call the cheap model.”
Pricing nuance matters for European mornings. DeepSeek halves prices off-peak; peak windows include 06:00–10:00 UTC weekdays — exactly when Central European workdays start — so the “cheap Flash” story can double mid-morning. Qwen’s Flash price does not flip by the hour. Per-answer cost rankings therefore depend on thinking, hour and token verbosity (Qwen wrote longer median answers). Endpoint pinning mattered too: OpenRouter listed dozens of DeepSeek resellers with a wide price spread; an unpinned call can silently hit a third party while you still brand the result “DeepSeek.”
How this fits the ReadIsi China beat. We already covered DeepSeek’s V4.1 Flash routing and Alibaba’s Qwen 4 still-in-training messaging. This test is the missing developer layer: Flash is only “fast” if you configure it like a production engineer, not like a demo chatbot. For GEO prototypes that fire DeepSeek-with-web-search or Qwen Model Studio search tools, empty answers under low token caps will look like “the model hates my site” when the real bug is thinking-mode budget exhaustion.
Practical checklist: disable thinking unless you need chain-of-thought; set max_tokens far above the expected answer; pin the provider; decide whether European business hours should prefer Qwen’s steadier tariff or DeepSeek’s off-peak wins. Quality was out of scope for MRKT30’s stopwatch — run your own citation and factuality panel before you crown a Flash winner for customer-facing search agents.
Opinion: Flash marketing without thinking-mode documentation is borderline malpractice. DeepSeek winning the Budapest stopwatch is real; shipping Flash with defaults that blank one call in ten is also real. Treat configuration as part of the product, and treat any GEO vendor who will not disclose thinking / token settings as selling screenshots, not systems.






1 Comment