<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>WIP on Code Now</title><link>https://blog.0xnullpath.cc/en/tags/wip/</link><description>Recent content in WIP on Code Now</description><generator>Hugo</generator><language>en</language><lastBuildDate>Mon, 20 Apr 2026 04:27:43 +0000</lastBuildDate><atom:link href="https://blog.0xnullpath.cc/en/tags/wip/index.xml" rel="self" type="application/rss+xml"/><item><title>[Model Inference] A Brief Look at CUDA Graph</title><link>https://blog.0xnullpath.cc/en/posts/note-snippet-15-model-inference-a-brief-look-at-cuda-graph/</link><pubDate>Mon, 20 Apr 2026 04:27:43 +0000</pubDate><guid>https://blog.0xnullpath.cc/en/posts/note-snippet-15-model-inference-a-brief-look-at-cuda-graph/</guid><description>&lt;p>CUDA Graph is probably one of the most frequently mentioned features in inference optimization. The principle isn&amp;rsquo;t complicated: package the kernels that were originally launched one at a time into a static graph, hand the whole thing to the GPU in one shot, and skip the scheduling overhead in between.&lt;/p>
&lt;p>But how exactly does this &amp;ldquo;packaging&amp;rdquo; work? How do you use it in PyTorch? And most importantly — how much faster is it actually on an H200?&lt;/p></description></item></channel></rss>