
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Yusheng Zheng</title>
      <link>https://www.yunwei37.com/blog</link>
      <description>Yusheng Zheng is a systems researcher working on GPU runtimes, distributed AI infrastructure, programmable systems, and agent observability.</description>
      <language>en-us</language>
      <managingEditor>yunwei356@gmail.com (Yusheng Zheng)</managingEditor>
      <webMaster>yunwei356@gmail.com (Yusheng Zheng)</webMaster>
      <lastBuildDate>Sat, 19 Sep 2026 15:30:00 GMT</lastBuildDate>
      <atom:link href="https://www.yunwei37.com/tags/profiling/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.yunwei37.com/blog/deepseek-v41-engram-nowait-dgx-spark</guid>
    <title>The Cached-Read Bottleneck: Speeding Up DeepSeek V4.1 Engram on Four DGX Sparks</title>
    <link>https://www.yunwei37.com/blog/deepseek-v41-engram-nowait-dgx-spark</link>
    <description>A production investigation of DeepSeek V4.1 disk-backed Engram on four DGX Sparks: how syscall, page-cache, and component profiling found a Python scheduling bottleneck, how RWF_NOWAIT fixed it, and what the production evidence does and does not prove.</description>
    <pubDate>Sat, 19 Sep 2026 15:30:00 GMT</pubDate>
    <author>yunwei356@gmail.com (Yusheng Zheng)</author>
    <category>llm-inference</category><category>deepseek</category><category>vllm</category><category>linux</category><category>storage</category><category>profiling</category>
  </item>

    </channel>
  </rss>
