<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>DeepSeek on Code Now</title><link>https://blog.0xnullpath.cc/en/tags/deepseek/</link><description>Recent content in DeepSeek on Code Now</description><generator>Hugo</generator><language>en</language><lastBuildDate>Thu, 01 Oct 2026 20:30:00 +0800</lastBuildDate><atom:link href="https://blog.0xnullpath.cc/en/tags/deepseek/index.xml" rel="self" type="application/rss+xml"/><item><title>DeepSeek's Attention Evolution - [DSA, CSA, CSA2]</title><link>https://blog.0xnullpath.cc/en/posts/note-snippet-21-deepseeks-attention-evolution-dsa-csa-csa2/</link><pubDate>Thu, 01 Oct 2026 20:30:00 +0800</pubDate><guid>https://blog.0xnullpath.cc/en/posts/note-snippet-21-deepseeks-attention-evolution-dsa-csa-csa2/</guid><description>&lt;div class="markdown-alert markdown-alert-note">
 &lt;p class="markdown-alert-title">
 &lt;span class="markdown-alert-kind">Note&lt;/span>
 &lt;span class="markdown-alert-num">&lt;/span>
 &lt;span class="markdown-alert-sep">|&lt;/span>
 &lt;span class="markdown-alert-text">Preface&lt;/span>
 &lt;/p>
 &lt;div class="markdown-alert-body">
 &lt;p>This is the second post in the DeepSeek Attention evolution series. The previous one, &lt;a href="https://blog.0xnullpath.cc/posts/note-snippet-20-deepseek-%E7%9A%84-attention%E6%BC%94%E8%BF%9B-mla/">DeepSeek&amp;rsquo;s Attention Evolution -【MLA】&lt;/a>, covered the evolution from MHA to MLA; this post picks up with DSA, CSA and CSA2. MLA solved the storage form of the KVCache, but its compute complexity is still $O(N^2)$ and the KVCache is still $O(N)$, so it still struggles with 256K/1M long contexts. This post follows the DeepSeek V3.2-exp → V4 → V4.1-Flash line to see how it pushes the algorithmic and engineering tradeoffs to the limit, step by step, across sparse attention, KV compression and cross-layer sharing.&lt;/p></description></item><item><title>DeepSeek's attention evolution - [MLA]</title><link>https://blog.0xnullpath.cc/en/posts/note-snippet-20-deepseeks-attention-evolution-mla/</link><pubDate>Tue, 15 Sep 2026 23:30:00 +0800</pubDate><guid>https://blog.0xnullpath.cc/en/posts/note-snippet-20-deepseeks-attention-evolution-mla/</guid><description>&lt;blockquote>
 &lt;p>Foreword: After reading the DeepSeekV4.1 paper and seeing that DeepSeek had once again made major changes to the model architecture, I was quite excited. So I looked back at DeepSeek&amp;rsquo;s incremental improvements to the model architecture and wrote this series of articles in the gaps while waiting for agents to finish running and slacking off at work — a summary of what I&amp;rsquo;ve learned about model architectures over the year since I moved into AI Infra.&lt;/p></description></item></channel></rss>