Blog
Research

10 min read

Megakernels

What are megakernels, why we care, what makes them hard, and how Hoid tackles them

Pavle Pađin & Hoid team

The word MEGAKERNELS drawn as tiles in a GPU profiler trace, one lane per SM, with MEGA in orange and the Hoid emblem as the full stop.

Introduction

Most inference optimization is focused on improving the performance of individual kernels in models, such as attention, mixture of experts, matmuls and others. That work is extremely important and is a big part of how frameworks like vLLM, SGLang and TensorRT-LLM get impressive performance for LLM inference.

However, in order to squeeze the most out of your GPUs, it is not enough to simply optimize standalone kernels. In memory-bound regimes, most notably in LLM decode, a big chunk of performance is lost between kernel launches.

The inter-kernel overhead consists of three main parts:

  1. Kernel launch overhead, which is several microseconds before every kernel launch. This overhead can be hidden in compute-intensive workloads, but LLM decode has many short kernels, which can’t be overlapped with the launch.
One decode layer as a timeline: a launch gap before each of norm, qkv, attn, o_proj, norm, gate_up and down. Several launch gaps are wider than the kernel that follows them.
Fig. 1 - Kernel launch overhead. When operations are compute-heavy (which is the case in LLM prefill), this launch overhead is almost negligible. However, in LLM decode, where many kernels are similar in length to the launch overhead itself, we cannot dismiss kernel launch overhead.
  1. Wave quantization, which occurs when the number of tasks is not divisible by the number of SMs. Then, at the epilogue of that operation, some SMs have to stay idle due to lack of work.
Wave quantization on 4 SMs, one row per SM. Top: 6 tasks take 2 waves and SM2 and SM3 sit idle for the second. Bottom: 50 tasks take 13 waves and the same two SMs sit idle only in the last.
Fig. 2 - Wave quantization. In example 1, we see that wave quantization leads to two SMs being idle in the second wave of tasks, which is a significant fraction of this op's runtime. This example is representative of what happens in LLM decode, where the number of tasks is small relative to the number of SMs. In example 2, when the operation is more compute-heavy, we have many tasks relative to the number of SMs. Therefore, wave quantization has a smaller impact on the utilization of the whole operation. This resembles what happens in LLM prefill.
  1. Lack of communication and computation overlap. That kind of overlap is, in general, one of the main ways to hide latency in kernel optimization. Overlapping one kernel’s computation with another’s loads and stores is very beneficial.
Two kernels, each load, compute, store. Example 1: kernel 2 starts loading only after kernel 1 stores. Example 2: kernel 2 loads during kernel 1's compute and is already computing when kernel 1 stores, finishing earlier.
Fig. 3 - Overlapping loads and stores with compute. Traditional kernel boundaries only allow us to start loading inputs for kernel 2 once the whole of kernel 1 has finished executing. However, if kernel 2 has inputs that don't depend on kernel 1's outputs, this is purely a constraint of the standard software programming model. The second image shows what would be the desired outcome, limited only by causal dependencies between inputs and outputs.

These three failure modes of the standard programming model motivate a new, different paradigm for kernel design, called megakernels. The next section will focus on explaining this mysterious and often misunderstood term.

What is a megakernel?

Many people in the AI performance engineering community have heard the term megakernel in the last year or so. Hazy Research Lab coined the term, and it has become one of those overloaded computer science terms (like the term “kernel” itself) where you need to explicitly state what you mean when you say “megakernel”. It has even been memed for how vague the term actually is:

Megakernel Iceberg meme: seven layered definitions of a megakernel, from a large fused kernel at the tip down to “the GPU is a megakernel” at the bottom.
Fig. 4 — Mark Saroufim, “The megakernel literature, summarized”.

In this blog post, we will explain parts of this iceberg, and which part of it Hoid takes.

Megakernel iceberg

Tip of the iceberg says that a large fused kernel is a megakernel. Kernel fusion is the process of joining two or more kernels together to avoid having to store and load the intermediate results back to global memory. This is an important AI compiler optimization, but it doesn’t encapsulate the core of the megakernel. A megakernel’s biggest advantage is removing the scheduling overhead of standard GPU kernels.

Layer 2 of the iceberg says that GPU-side scheduling is a megakernel. This states that the GPU has scheduling logic that decides which operation executes next, and not the host (CPU). While this often ends up being the case in efficient megakernels, it is only a single step toward making megakernels actually efficient.

Layer 3 says that a CUDA graph is a megakernel. CUDA graphs record the exact sequence of kernels that you want to execute and then record all the parameters for launching them. The next time the same graph should be executed, CUDA graphs just replay the sequence of kernels that was prerecorded. That way, they remove the kernel launch overhead before every kernel. They are very useful and widely used in the ecosystem, especially in LLM decode, when launch overhead accounts for a larger portion of the whole execution time.

Unfortunately, there are two main limitations of CUDA graphs:

  1. They reduce the kernel launch overhead, but still ~1 microsecond per kernel is needed for launching the kernel.
  2. They still enforce synchronization between two separate kernels, which means that all blocks of kernel 1 have to finish and before moving on to kernel 2. This results in wave quantization and doesn’t allow overlapping two consecutive kernel executions.

Layer 4 states that device-launched CUDA graphs are megakernels. As the name suggests, this allows the GPU to decide whether to launch a CUDA graph without host intervention. While this adds more flexibility to the CUDA graph execution, its main problems remain.

Layer 5 says that PDL (Programmatic Dependent Launch) is a megakernel. PDL enables overlapping part of the end of one kernel with the loading of static inputs of the next kernel. It is one step beyond CUDA graphs and helps reduce the inter-kernel overhead. It cannot, however, remove the boundaries between two kernels, and the blocks of the first kernel still have to finish before any block of the second kernel starts executing.

Layer 6 states that fine-grained overlap is a megakernel. This is the layer where Hoid’s megakernels stand! This means that you merge multiple kernels into a single megakernel in such a way that you schedule, dispatch, and manage all dependencies on device. There is no global synchronization between two different operations. Each tile can be executed at the exact moment its dependencies finish executing.

Additionally, managing dependencies manually gives you the ability to prefetch any inputs.

Layer 7. This layer is just for the meme!

Hoid’s approach to megakernels

Hoid is adopting the on-GPU interpreter megakernel model to manage fine-grained dependencies between operations. Unlike the traditional kernel programming model, where each kernel poses a synchronization boundary, the on-GPU interpreter treats one block as a single unit of work. That means that each block will only wait for its direct producer blocks to finish before launching. The standard kernel model forces every block of the latter kernel to wait for every block of the former kernel to finish. This is a simple way to avoid race conditions, but it comes at the cost of performance.

a) Traditional kernel design. Kernel 1 and kernel 2 are grids of blocks split by a kernel boundary; on the SM timeline every SM waits idle at a barrier for kernel 1's slowest block before kernel 2 starts. b) The megakernel on-GPU interpreter. Each SM calls the task queue's pop_next_task(), which returns a task such as kernel 1 block 2,3 or kernel 2 block 1,1; on the timeline kernel 2 blocks start as soon as an SM is free, with no barrier.
Fig. 5 - On-GPU interpreter. a) In traditional kernel design, we think of one kernel as one unit of work. That means that each block of kernel 2 has to wait for all blocks of kernel 1 to finish.

b) When we adopt the On-GPU interpreter model, one unit of work is one block. Therefore, a single block only has to wait for its producers to finish, which removes the idle time on SMs.

With this said, it is obvious that such a megakernel design is a big leap in terms of complexity compared to the standard kernel design. That leads to one of the key takeaways about megakernels: it is not hard to make a megakernel, but it is very hard to make a performant megakernel!

Two failure modes that are common in megakernels are scheduling overhead and inefficient block-level implementations of kernels.

Three timelines of the same two-kernel workload on 4 SMs. a) Well scheduled with efficient blocks, the reference. b) Scheduling overhead: every kernel 1 block runs as early as possible, but three kernel 2 blocks wait for all of kernel 1 to finish instead of just their own producer, leaving bubbles until the long k1 · 2,4 ends. c) Inefficient block-level kernels: the same schedule as a), but every block takes about twice as long, so the run ends latest.
Fig. 6 - Two failure modes of megakernels. a) A well-scheduled megakernel with efficient blocks. b) Bad scheduling makes blocks wait when they shouldn't. This means that dependency management for this megakernel is not fine-grained enough. c) Inefficient block-level implementations defeat the purpose of megakernels.

The craft of making performant megakernels boils down to avoiding these two failure modes. Hoid takes an agent–compiler co-design approach for this purpose. AI compiler infrastructure provides scheduling primitives, and agents have the ability to write optimized block-level kernels. We believe this is the only viable way to generate performant megakernels at scale.

Results

The following figure shows a comparison of Hoid’s megakernel with vLLM, SGLang, and TensorRT-LLM. We are beating all the popular inference engines on a variety of batch sizes. We are also making our work open source! The code to run the vLLM server with Hoid’s megakernel and to reproduce the benchmarks is available at: https://github.com/hoid-ai/hoid-megakernel-qwen/.

Grouped bar chart of interactivity in tokens per second per user for Qwen3-4B decode, 8K in and 1K out on one H200, at batch sizes 1, 2, 4 and 8. Hoid megakernel: 317, 283, 237, 176. vLLM: 255, 236, 208, 166. SGLang: 270, 241, 209, 169. TensorRT-LLM: 219, 198, 178, 143. Hoid is highest at every batch size.
Fig. 7 - Hoid megakernels beat vLLM, SGLang, and TensorRT-LLM on several batch sizes. For a fair comparison, the Hoid megakernel has been integrated into vLLM as a plugin, to add the same scheduling overhead that appears in inference engines.

This is just the beginning. We are aiming to build the first AI compiler that can consistently generate performant megakernels.

More posts

All posts

Be there when
the next world opens.

Book a call