metakernel

The measuring instrument for the kernel-design loop, packaged as a kernel and loadable through kernels: trustworthy timing, adversarial correctness comparison, and measured facts about the device. The reference standard is the suite of measurement invariants any correctly functioning device must satisfy, certified on multiple architectures.

An agent iterating on CUDA candidates is good at writing code and bad at trusting its own measurements. It needs three things it cannot reliably improvise: timing that rejects thermal throttling, comparison that distinguishes real bugs from legitimate reduction-order noise, and device numbers that were measured rather than copied from a spec sheet. metakernel ships those as loadable ops that run anywhere get_kernel runs, including headless CI where Nsight and CUPTI are unavailable, which is exactly where agent loops live.

A roofline plot draws itself from a live device probe: the bandwidth diagonal, five measured compute roofs, and a real GEMM plotting onto it

The dossier drawing its own roofline, every line measured live by probe kernels: the 790 GB/s DRAM diagonal, compute roofs from fp64's 1.2 TF to int8's 715 TOP/s, the dependent-load latency ladder from 30 shared-memory cycles to 666 at DRAM, and a real fp16 GEMM plotting itself at its measured 116 TF with throttled samples rejected from the median.

Usage

from kernels import get_kernel
mk = get_kernel("phanerozoic/metakernel", version=1, trust_remote_code=True)

dossier = mk.probe_device()                          # measured device facts
report  = mk.bench(lambda: cand.op(x), iters=200,    # throttle-rejected medians
                   bytes=B, flops=F, dossier=dossier)
verdict = mk.compare(ref_op, cand_op, gen=make_inputs)   # pass/fail with ULP evidence
table   = mk.sweep(factory, {"BLOCK": [64, 128, 256]})   # config search

Every op returns plain dicts and tensors with fixed schemas, never prose, so an agent parses results instead of scraping stdout. version selects the release branch; trust_remote_code is required by kernels for publishers without the trusted-publisher mark.

API

Symbol Purpose
probe_device(quick=False) empirical dossier: DRAM/L2/shared/transfer bandwidth, pointer-chase latencies, FMA and per-dtype MMA rates, atomic contention curve, launch overhead, grid-barrier cost, occupancy tables, clock under load
bench(fn, iters, bytes, flops, dossier, dtype, graph, clock_stride) cudaEvent bracketing; median with CI; rejection of samples whose SM clock sagged; achieved GB/s and TFLOP/s; roofline placement; CUDA-graph capture for microsecond kernels
compare(ref, cand, gen, order_shuffle, atol, rtol, zero_atol) max/mean ULP, log2 ULP histogram, first divergent index, NaN/Inf census; with an order shuffle every element is judged against its own pointwise reduction-order band; carries worst_elements for debugging
order_shuffles.last_dim() / .matmul_k() built-in shuffles for the common contractual-noise axes
fuzz_shapes(op, ref, make, dtypes, isolate) the ugly envelope: 0-d to 3-d and zero-size shapes, non-tile-multiple sizes, transposed and strided views, misaligned slices, 2^31 logical-element expansion; per-case repro recipes; isolate=True survives context-killing candidates
sweep(factory, param_grid, check, budget_s, race) launch-config search with a correctness gate per row so a fast wrong config cannot win; successive halving with pruned points visible
alloc_stamps / read_stamps / stamped_demo clock64 phase stamps around agent-marked kernel phases
dossier_drift(current, pinned, rel_tol) / pinned_dossiers() drift of fresh measurements against the release-pinned dossier per architecture
mma_tile_check(trials) random integer tiles through the wmma fragment path, compared exactly against the reference GEMM
measure_clock() / ulp_diff(a, b) the throttle primitive; elementwise ULP distance
ops.mk_* the raw probe kernels

Method

Each dossier entry is produced by a dedicated probe kernel: streaming triads and read/write/gather kernels for bandwidth, Sattolo-cycle pointer chases for latency, register-resident FMA chains and register-fragment mma_sync loops for compute, an atomic hammer swept over distinct-address counts, empty-launch trains, and a cooperative-groups kernel for grid-barrier cost. Every number carries the SM clock it was measured under; bench probes the clock and rejects samples from sagged windows, reporting the count. The comparison ops treat the reference as a distribution: an order shuffle over contractual-noise dimensions bounds the legitimate numerical band, applied pointwise so a uniformly wrong candidate cannot hide behind one order-sensitive element.

Measured

Release-pinned dossier for RTX 6000 Ada (sm89, 142 SMs, 96 MB L2, ECC on):

entry value
clock under sustained FMA load 2.71 GHz max, 0.3% sag over 24 windows
device-memory bandwidth triad 638-705 GB/s; read-only 771; write-only 767; random 16 B gather 218
L2 / shared-memory bandwidth 9,325 GB/s / 21,250 GB/s aggregate
host transfers H2D 24.0 pinned; D2H 26.2; D2D 331 GB/s
dependent-load latency smem 30.0 cy; L1 45.8; L2 292; DRAM 649
FMA pipes f32 90.9 TFLOP/s; f64 1.28 TFLOP/s
MMA sustained (accumulator-verified) fp16 376, bf16 366, tf32 177, int8 770 TOP/s, 2:4-sparse fp16 713; fp8 194 as a flagged lower bound
atomics fp32 curve 0.72 Gop/s at 1 address to 484 at one per thread
launch overhead / grid barrier 5.6 us queued, 14.1 round trip / 0.54 us per grid.sync

A second pinned dossier covers GeForce RTX 3070 Ti Laptop (sm86). MMA flops derive from accumulator read-back with every fragment element observed; mma_tile_check reproduces the reference GEMM exactly through the same fragment path.

Certification

The test suite asserts invariants that must hold on any correctly functioning device: L2 exceeds DRAM bandwidth and random gather costs more than streaming; latencies order smem < L1 < L2 < DRAM; tensor-core rates beat the FMA pipe and hold the silicon's structural ratios; a large cuBLAS GEMM lands within [0.30, 1.05] of the dossier fp16 rate; compare passes reordered reductions and fails corrupted candidates through both routes; fuzz_shapes catches a contiguity-cheating op and reports the 2^31 case as run or skipped, never silently; the chase probe repeats within 20%.

Requirements and limits

  • No hardware counters: everything is inferred from timing, clocks, and the occupancy API, so entries carry a confidence field and timing-differential inference can misattribute.
  • Single device, single tenant: contention surfaces as rejected samples, not errors. launch_overhead includes torch dispatch.
  • MMA numbers are dense-tile register rates with no memory traffic, an upper bound. In-process fuzz_shapes cannot survive a context-killing candidate; isolate=True can, at a torch startup per child.
  • Published variants are Linux x86_64; on Windows load_local.py JIT-builds the same source.

References

Williams, Waterman, Patterson, "Roofline: An Insightful Visual Performance Model" (2009); Sattolo's algorithm for single-cycle pointer chases; the occupancy API and cudaEvent timing model.

License

Apache-2.0.

Downloads last month
-
apache-2.0
Supported hardwares new
CUDA
8.08.68.99.010.012.0
GPU
B300
288GB
NVIDIA SXM
B200
192GB
NVIDIA SXM
H200
141GB
NVIDIA SXM
H100
80GB
GPU
H800
80GB
GPU
H20
96GB
GPU
L40s
48GB
GPU
L40
48GB
GPU
L20
48GB
GPU
L4
24GB
DGX Spark
GB10
128GB
GPU
RTX PRO 6000 WS
96GB
GPU
RTX PRO 6000 Max-Q
96GB
GPU
RTX PRO 5000
48GB
GPU
RTX PRO 4500 WS
32GB
GPU
RTX PRO 4000
24GB
GPU
RTX PRO 4000 SFF
24GB
GPU
RTX PRO 2000
16GB
GPU
RTX 6000 Ada
48GB
GPU
RTX 5880 Ada
48GB
RTX
RTX 5000 Ada
32GB
GPU
RTX 4500 Ada
24GB
RTX
RTX 4000 Ada
20GB
RTX
RTX 4000 SFF Ada
20GB
GPU
RTX 3500 Ada Mobile
12GB
GPU
RTX 2000 Ada
16GB
GPU
RTX A6000
48GB
GPU
RTX A5000
8GB
GPU
RTX A5000 Max-Q
16GB
GPU
RTX A5000 Mobile
16GB
GPU
RTX A4000
16GB
GPU
RTX A4000 Max-Q
8GB
GPU
RTX A4000 Mobile
8GB
GPU
RTX A3000 Mobile
6GB
GPU
RTX A2000
6GB
GPU
RTX A2000 Embedded
4GB
GPU
RTX A2000 Max-Q
4GB
GPU
RTX A2000 Mobile
4GB
GPU
A800
40GB
GPU
A100
80GB
GPU
A40
48GB
GPU
A30
24GB
GPU
A10
24GB
GPU
A2
16GB
RTX
RTX 5090
32GB
RTX
RTX 5090 D
32GB
RTX
RTX 5090 Mobile
24GB
RTX
RTX 5080
16GB
RTX
RTX 5080 Mobile
16GB
RTX
RTX 5070
12GB
RTX
RTX 5070 Mobile
8GB
RTX
RTX 5070 Ti
16GB
RTX
RTX 5070 Ti Mobile
12GB
RTX
RTX 5060 Ti
16GB
RTX
RTX 5060
8GB
RTX
RTX 5060 Mobile
8GB
RTX
RTX 5050
8GB
RTX
RTX 5050 Mobile
8GB
RTX
RTX 4090
24GB
RTX
RTX 4090D
24GB
RTX
RTX 4090 Mobile
16GB
RTX
RTX 4080 SUPER
16GB
RTX
RTX 4080
16GB
RTX
RTX 4080 Mobile
12GB
RTX
RTX 4070
12GB
RTX
RTX 4070 Mobile
8GB
RTX
RTX 4070 Ti
12GB
RTX
RTX 4070 Super
12GB
RTX
RTX 4070 Ti Super
16GB
RTX
RTX 4060
8GB
RTX
RTX 4060 Ti
8GB
RTX
RTX 4090 Laptop
16GB
RTX
RTX 4080 Laptop
12GB
RTX
RTX 4070 Laptop
8GB
RTX
RTX 4060 Laptop
8GB
RTX
RTX 4050 Laptop
6GB
RTX
RTX 3090
24GB
RTX
RTX 3090 Ti
24GB
RTX
RTX 3080
12GB
RTX
RTX 3080 Ti
12GB
RTX
RTX 3080 Mobile
16GB
RTX
RTX 3070
8GB
RTX
RTX 3070 Ti
8GB
RTX
RTX 3070 Ti Mobile
8GB
RTX
RTX 3060 Ti
8GB
RTX
RTX 3060
12GB
RTX
RTX 3060 Mobile
6GB
RTX
RTX 3050 Mobile
4GB
GPU
RTX 2050 Mobile
4GB
Jetson
Jetson AGX Orin 64GB
64GB
Jetson
Jetson AGX Orin 32GB
32GB
Jetson
Jetson Orin NX 16GB
16GB
Jetson
Jetson Orin NX 8GB
8GB
Jetson
Jetson Orin Nano 8GB
8GB
Jetson
Jetson Orin Nano 4GB
4GB
OS
linux
Arch
x86_64
Kernel Builder
19aaa64