DeepSeek's New Model Sets a Template for Powerful LLMs that Run Lean
The model charges a fraction of a cent per million tokens and will replace V4-Pro as DeepSeek intensifies pressure on rivals.
- On Thursday, Sep 10, Chinese startup DeepSeek unveiled its V4.1 Flash model, a cost-optimized platform charging as little as a fraction of a cent per million tokens while claiming to outperform mainstays from Anthropic and Z.AI.
- DeepSeek redesigned its architecture to reduce key-value cache consumption to between 13 percent and 25 percent of prior requirements, while introducing N-gram parameters that function as implicit memory to boost intelligence without proportional memory increases.
- At 763 billion parameters, the model outperforms the V4-Pro in coding tasks, though it trails flagship models from Anthropic and OpenAI. Starting Sep 14, DeepSeek will retire the V4-Pro by automatically rerouting all inference tasks to the V4.1 Flash at cheaper rates.
- Industry competition is intensifying as Alibaba recently revealed its Qwen 3.8-Flash-Next model, which similarly employs N-gram techniques to challenge ChatGPT-developer OpenAI and other US rivals in a broadening price battle.
- By offloading N-gram weights to cheaper system RAM, the model reduces minimum GPU memory requirements from 763 GB to around 567 GB, enabling efficient deployment across high-throughput production environments without sacrificing performance.
23 Articles
23 Articles
DeepSeek's New Hyper-Efficient Model Stokes Fears Over Korea's Memory Makers
DeepSeek's New Hyper-Efficient Model Stokes Fears Over Korea's Memory Makers Samsung Electronics and SK Hynix each fell more than 3% in Seoul on Friday after DeepSeek said its newest AI model needs a fraction of the memory required by its predecessor. Both stocks had been trying to recover from July's selloff and remain more than 25% below their highs. Local retail traders, who helped drive the rally earlier this year, have sold around $10 billi…
China's DeepSeak has unveiled AI Model V4.1 Flash, which reduces costs and memory usage. By improving KV Cache and SSD efficiency, it has lowered the consumption of computational resources; this news caused global semiconductor and AI-related stocks to fall together. Experts predict that intensifying price competition in the AI industry due to technological standardization will threaten corporate survival.
With 763 billion parameters and innovative memory management, DeepSeek V4.1 Flash wants to make future AI models more efficient.
Coverage Details
Bias Distribution
- 80% of the sources lean Right
Factuality
To view factuality data please Upgrade to Premium





![[your]NEWS](/_next/image?url=https%3A%2F%2Fgroundnews.b-cdn.net%2Finterests%2Ffb6dc495f74049f513563c33352175eaa0ecd509.jpg%3Fwidth%3D60&w=128&q=75)















